Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

143 results about "Speaker identification" patented technology

Simulation digital human real-time intelligent voice interaction system and method based on vision and large model

The invention relates to a simulation digital human real-time intelligent voice interaction system and a simulation digital human real-time intelligent voice interaction method based on vision and a large model, and aims to solve the problems of inaccurate target speaker recognition, high response delay and the like in digital human voice interaction in a complex scene. The system circles an effective recognition range through a camera, triggers audio collection in combination with face detection, locks a target speaker and reduces noise by using lip movement recognition and sound image fusion technologies, converts the target speaker into a text through voice wake-up, generates an answer by means of a large language model (LLM) and knowledge retrieval enhancement (RAG) technologies, generates low-delay voice through a voice synthesis technology accelerated by the vLLM, and performs voice recognition on the target speaker. And driving the preloaded digital human image to synthesize a video stream and pushing the video stream to a front end for rendering in real time. Accurate pickup, low-delay interaction and rapid digital human image switching in a complex environment are realized, the accuracy and real-time performance of intelligent voice question answering are improved, and the method is suitable for government affair halls, exhibition halls and other scenes.
Owner:UNICOM (HENAN) IND INTERNET CO LTD

Zero-configuration adaptive speaker recognition method and system

The invention discloses a zero-configuration adaptive speaker recognition method and system, and relates to the technical field of voice signal processing. The method comprises the following steps: receiving an audio stream, carrying out voice activity detection, obtaining a single-person voice segment, extracting a voiceprint embedding vector of the single-person voice segment, and carrying out online clustering to generate a speaker identity pool; and the identity pool is updated by calculating the multi-dimensional fusion similarity between the remaining single-person voice segments and the voice model, and temporary identity tags corresponding to the single-person voice segments are output. According to the speaker recognition method provided by the invention, a voiceprint template does not need to be registered in advance, and the flexibility and real-time performance of the system are improved.
Owner:北京文聿科技有限公司

System and methods for audio data analysis and tagging

A system for automated processing and analysis of audio files for large data sets in a cloud environment. A unified analytic environment can integrate audio machine learning models for processing and analysis with a knowledge management system, including graph presentations of tracked entities, linked to audio files and / or associated translations and transcripts. Entities within such data can be searched or filtered and proposed for tracking, or identified as tracked objects. These features can allow triage and prioritization of audio files for analysis. User interfaces can facilitate feedback on transcription and translation outputs, thereby improving present outputs and future inputs and outputs. Entities speaking or referred to can be found, tagged, and distinguished in audio files (e.g., using speaker identification in audio files, text searching in transcripts, etc.) Users can provide feedback and input on various aspects of a system, to enhance or adjust initial automated or other machine learning outputs.
Owner:PALANTIR TECHNOLOGIES INC

Multitask speech emotion recognition method based on parallel processing hybrid expert network

The invention belongs to the technical field of speech recognition, and particularly discloses a multi-task speech emotion recognition method based on a parallel processing hybrid expert network, and the method comprises the steps: carrying out the feature extraction of an input original speech, and obtaining a shared feature; constructing a parallel processing hybrid expert network, inputting the shared features into the parallel processing hybrid expert network, carrying out sentiment classification and speaker recognition on the shared features in parallel, and respectively outputting a predicted sentiment value and a predicted speaker; inputting the shared features into a full connection layer for voice recognition processing to obtain a text prediction value; and obtaining a loss function according to the predicted emotion value, the predicted speaker and the text predicted value in combination with the corresponding labels, and optimizing the model parameters according to the loss function to obtain the parallel processing hybrid expert network after optimization training. According to the invention, the speech emotion recognition performance and recognition accuracy can be improved.
Owner:HUAZHONG NORMAL UNIV

Speaker recognition method in medical real-time speech recognition scene

ActiveCN121938378Aconsistent with auditory logicReduce the risk of number fluctuationsSpeech recognitionAutomatic speechSpeech sound
The invention discloses a speaker recognition method in a medical real-time speech recognition scene, and the method specifically comprises the steps: carrying out the parallel distribution of a received PCM audio stream through an audio diverter in a gateway layer, and transmitting the audio stream to an automatic speech recognition link and a speaker separation link; in the automatic speech recognition link, outputting a recognition text and Token-level timestamps corresponding to each minimum semantic unit in the recognition text; in the speaker separation link, extracting a corresponding speaker embedding vector based on a preloaded speaker embedding model, and performing online clustering processing on the embedding vector to generate a speaker identity identifier; token-level boundary adsorption labeling processing is executed based on the Token-level timestamp and an online clustering result, so that a speaker switching position is determined; and carrying out one-to-one mapping on the speaker tags and the recognition text according to the Token-level timestamps, merging Tokens with the same speaker tags continuously, and outputting a transliteration text with the speaker tags.
Owner:ZOE SOFT CORP LTD

Automatic speech recognition result optimization system and method, medium and equipment

The invention provides an automatic speech recognition result optimization system and method, a medium and equipment, and the system comprises an ASR preprocessing module which carries out the initial ASR recognition of an original audio, obtains a preliminary recognition text, carries out the named entity recognition of the preliminary recognition text, enables a recognition result to be matched with a hot word with similar pronunciation in a hot word library, and carries out the recognition of the hot word; inputting the hot words into an ASR model for secondary identification; the ASR text post-processing module is used for carrying out semantic error correction on the text output by the ASR model by utilizing a large language model and verifying the reasonability of the text; and the ASR speaker recognition module is used for segmenting and numbering the audio by using a speaker separation technology, and mapping the number with a specific name. According to the invention, by adopting a mode of entity word extraction and hot word matching, the problem that a speech recognition system cannot screen related hot words in advance is solved, and the recognition effect of an ASR system is improved.
Owner:SHANGHAI SHENGHEKUN INFORMATION TECH CO LTD

Intelligent conference summary generation method and device, equipment and storage medium

The invention provides an intelligent conference summary generation method and device, equipment and a storage medium, and relates to the technical field of natural language processing. According to the method provided by the invention, clear capture and accurate identification of audios are realized by using the multi-channel microphone array; and inputting the audio with the speaker identifier into the speech recognition model, and converting the audio into a structured initial summary text with a timestamp. Performing semantic analysis on the initial summary text through a large language model, and extracting key text information; a rule engine and a large language model are combined for combined judgment, so that the information accuracy is ensured; the initial summary text is converted into a first summary text in a preset text style, so that the text specialty is improved; humanized feature injection enhances the language naturalness of the second summary text; and the second summary text is converted into the structured target summary text, so that the conference summary better meets the requirements of professional scenes such as financial risk control conferences and medical consultation conferences, and the accuracy and availability of intelligent generation of the conference summary are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Unlearnable speech data generation method against model fine-tuning attacks

This invention discloses a method for generating non-learnable speech data against model fine-tuning attacks, comprising: constructing a proxy speaker recognition model based on a pre-trained speech representation network and a differentiable classifier head; initializing perturbation variables and calculating a temporal saliency mask; determining target identity labels for source samples using a candidate buffer pool and cosine similarity rules; simulating an attacker in the inner layer optimization, and fine-tuning the model using protected speech and real speaker labels; in the outer layer optimization, using the model parameters obtained from the inner layer optimization as a temporary model, calculating the outer layer loss for the target identity labels, and solving the gradient of the perturbation according to a first-order approximation strategy; updating the perturbation and performing projection and spectral difference constraints; and outputting protected speech through repeated iterations. This invention improves the effectiveness against target identity attacks, enhances identity obfuscation capabilities, reduces optimization waste caused by invalid targets, and improves the stability and practicality of the speech identity protection process.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Audio processing method and device, electronic equipment and storage medium

The invention relates to the technical field of audio processing, in particular to an audio processing method and device, electronic equipment and a storage medium. The method comprises the following steps: a main audio acquisition device sets any device as a main speaking device in response to a main speaking mode starting operation triggered for any device, and sets a main speaking identifier of the main speaking device; if the audio data collected by at least one device is obtained, screening out a main speaking device from the devices based on the main speaking identifier and the device identifier carried by each audio data; if the main speaking equipment is screened out, preset operation is executed on the at least one piece of audio data to obtain audio data to be transmitted, and the preset operation comprises at least one of the following operations: enhancement processing is performed on the audio data collected by the main speaking equipment, and suppression processing is performed on other audio data; and transmitting the to-be-transmitted audio data to the terminal equipment for playing. According to the embodiment of the invention, in the speaker mode, the interference of other sounds on the speaker sound is reduced, and the sound collection effect of the cascade equipment is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

System and method for associated narrative based transcription speaker identification

Techniques for associated narrative based transcription speaker identification are provided. A narrative of an incident is received at a computing device. The narrative describes an incident. An identification of at least one person involved in the incident is extracted from the narrative. The identification includes a specific identifier for the at least one person. Semantic information is extracted from the narrative. A transcript of media capturing the incident is received at the computing device. The transcript includes a generic identifier for at least one speaker whose speech was transcribed. The generic identifier for the at least one speaker whose speech was transcribed is correlated with the identification based on the semantic information. The generic identifier for the at least one speaker in the transcript is replaced with the specific identifier included in the identification.
Owner:MOTOROLA SOLUTIONS INC

Electronic device and method for providing translation function by using same

This electronic device may display, on a display, a user interface for providing a translation function of a plurality of speakers, and when a first voice signal is received through a microphone and a first user interaction is detected, train, on the basis of the first voice signal, a speaker recognition model for a characteristic and language of a speaker corresponding to the first voice signal. The electronic device may determine whether a second voice signal is received through the microphone and whether a second user interaction is detected. When the second user interaction is not detected, the electronic device may determine whether there is a training-completed speaker recognition model corresponding to to the second voice signal. When there is the training-completed speaker recognition model corresponding to the second voice signal, the electronic device may translate the second voice signal into a target language on the basis of the training-completed speaker recognition model, and display same on the display.
Owner:SAMSUNG ELECTRONICS CO LTD

End-to-end speaker recognition using deep neural network

The present invention is directed to a deep neural network (DNN) having a triplet network architecture, which is suitable to perform speaker recognition. In particular, the DNN includes three feed-forward neural networks, which are trained according to a batch process utilizing a cohort set of negative training samples. After each batch of training samples is processed, the DNN may be trained according to a loss function, e.g., utilizing a cosine measure of similarity between respective samples, along with positive and negative margins, to provide a robust representation of voiceprints.
Owner:PINDROP SECURITY INC

Method and system for identifying a speaker of interest in an audio

The present method (300) identifies a speaker of interest in an audio file through a systematic approach. The process begins by receiving an input audio file via a processor (201). The audio file is then split into one or more chunks, followed by the extraction of relevant features from each chunk. Using a transformer encoder model, embeddings of the speaker of interest are generated based on these extracted features. The method identifies one or more nearest neighbours from various data structures corresponding to potential speakers, utilizing a classification model based on the generated embeddings. A set of nearest neighbours is then identified, ensuring that the count exceeds a predefined threshold and that the distance of each neighbour remains below a specified nearest-neighbour distance threshold. Finally, the method provides an identification of the speaker of interest as one of the recognized persons, enhancing speaker recognition capabilities in audio analysis.
Owner:ONIBER SOFTWARE PTE LTD

Intelligent sound equipment control method and system

The embodiment of the invention provides an intelligent sound equipment control method and system, and belongs to the technical field of data control. The method comprises the following steps: collecting first voice data when a target user controls a target sound box, and obtaining second voice data in a target environment where the target sound box is located; obtaining third voice data related to the first voice data from the second voice data according to the voice features of the first voice data; performing voice recognition on the first voice data to obtain a control text, and performing instruction recognition on the control text to obtain an initial instruction of the target user; speaker recognition is carried out on the third voice data to obtain a target speaker, and then a target scene and a target emotion of the target environment are determined according to the target speaker and a voice text of the third voice data; adjusting the initial instruction according to the target scene and the target emotion to obtain a target instruction of the target user; and controlling the target sound according to the target instruction to obtain a target control result.
Owner:GANZHOU DEHUIDA TECH CO LTD

Training methods, systems, media, and cross-channel and dialect voiceprint recognition models

The present application relates to the technical field of speaker recognition, in particular to a training method, system, medium and voiceprint recognition model across channels and dialects. The voiceprint recognition model has a first input layer, a first convolutional layer, a voiceprint feature extraction network layer, a feedforward network layer and a first output layer connected in turn. When training the voiceprint recognition model, the following steps are included: step S1, constructing a text feature extraction network, a second output layer and a third output layer; step S2, constructing a training set X; step S3, initializing the parameters of the voiceprint recognition model; step S4, training the voiceprint recognition model and constructing a loss function L; step S5, updating the parameters of the voiceprint recognition model; step S6, repeating steps S4 and S5 until the loss function is optimal, and completing the training of the voiceprint recognition model. The present application can preferably improve the generalization and robustness of the trained model.
Owner:NANJING PUBLIC SECURITY BUREAU

Systems and methods for generating video based on informational audio data

In one embodiment, a computer-implemented method may include receiving a media file, extracting, using an artificial intelligence engine including one or more trained machine learning models, one or more audio features from the media file. The one or more audio features include at least one of a time-synchronized transcript, speaker recognition data, mood data, index data, visual asset speaker data, color palette data, and written description speaker data. The method may include generating, based on the one or more audio features, a video, wherein the video is presented via a media player on a user interface of a computing device.
Owner:MUSIXMATCH SPA

Custom tone and vocal synthesis method and apparatus, electronic device, and storage medium

A custom tone and vocal synthesis method and apparatus, an electronic device, and a storage medium. The synthesis method comprises: training a first neural network by means of a speaker record sample to obtain a speaker recognition model, the output training result of the first neural network being a speaker vector sample (S102); training a second neural network by means of an unaccompanied vocal singing sample and the speaker vector sample to obtain an unaccompanied singing synthesis model (S104); inputting a speaker record to be synthesized into the speaker recognition model to obtain speaker information output by the intermediate hidden layer of the speaker recognition model (S106); and inputting unaccompanied singing music information to be synthesized and the speaker information into the unaccompanied singing synthesis model to obtain a synthesized custom tone and vocal (S108).
Owner:BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1

Audio processing method, training method of audio processing model and electronic equipment

The invention discloses an audio processing method, an audio processing model training method and an electronic device, and relates to the technical field of audio processing, the method is applied to a first electronic device, and the method comprises the following steps: the first electronic device performs feature extraction on an input audio, and obtains a voiceprint feature sequence corresponding to the input audio; and the first electronic equipment performs voiceprint extraction on the voiceprint feature sequence, and identifies the voiceprint of the speaker corresponding to the voiceprint feature sequence. And the first electronic equipment takes the voiceprint of the speaker as auxiliary query, performs voiceprint clustering on the voiceprint feature sequence, and obtains voice activity information of the speaker corresponding to the voiceprint feature sequence. In the application, the first electronic device can update the voiceprint library by using the voiceprint information output in real time, and does not need to actively register the voiceprint information of a speaker. Moreover, the voiceprint of the speaker is used as auxiliary query and is applied to the voiceprint clustering part, so that the whole audio processing performance and the speaker recognition accuracy can be improved.
Owner:HONOR DEVICE CO LTD +1

Learning method, speaker recognition method, and recording medium

The problem to be solved by the present invention is to identify a speaker with high accuracy. A learning method, a speaker identification method, and a recording medium are provided. The learning method is a learning method for a speaker identification model (20). When voice data is input, the speaker identification model (20) outputs speaker identification information identifying the speaker uttering the voice contained in the voice data. The speaker identification model (20) generates second voice data of a second speaker by performing voice feature transformation processing on first voice data of a first speaker. The first voice data and the second voice data are used as learning data for learning processing of the speaker identification model (20).
Owner:PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA

Device and method with target speaker identification

A processor-implemented method includes: extracting a target speaker voice feature based on an input voice of a target speaker; determining an utterance scenario of the input voice based on the target speaker voice feature; generating a final target speaker voice feature based on the determined utterance scenario; and determining whether the target speaker corresponds to a user based on the final target speaker voice feature and a final user voice feature, wherein the determined utterance scenario comprises either one of a single-speaker scenario and a multiple-speaker scenario.
Owner:SAMSUNG ELECTRONICS CO LTD

Speaker-specific speech extraction method and device based on speaker auxiliary information

This application relates to artificial intelligence technology and provides a method, apparatus, device, and medium for extracting the speech of a specific speaker based on speaker auxiliary information. The method first obtains a framing result corresponding to the audio and video data to be identified, then obtains the sub-speaker activity information corresponding to each frame of sub-audio and video data in the framing result to form the speaker activity information. Input data for the corresponding frame of sub-audio and video data is then generated based on the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data in the framing result. Finally, the input data is input into the speaker classification model for classification. The obtained target speaker recognition features are then multiplied by the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data to obtain the target speaker spectrum corresponding to each frame of sub-audio and video data. This method achieves the separation of the speech spectrum of a specific speaker from mixed speech without inputting the target speaker's speech, simplifying the extraction process.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speaker recognition model frequency modulation trigger injection method

InactiveCN121459823ASpeech analysisFrequency spectrumSpeaker recognition system
The invention discloses a speaker recognition model frequency modulation trigger injection method, particularly relates to the technical field of voice signal processing and artificial intelligence safety, and is used for solving the problems that an existing speaker recognition model back door injection method is insufficient in concealment, poor in black box environment adaptability, prone to distortion of trigger characteristics in the physical transmission process and poor in safety. And mainstream defense strategies such as model fine tuning, pruning and spectrum detection are difficult to resist. The method comprises the following steps: constructing a piecewise linear periodic low-frequency modulation curve for original voice, generating a smooth phase track according to a frequency and phase relationship, constructing a sinusoidal modulator to generate a hidden sample, mixing normal samples in proportion and uniformly marking, training a model by a black box, testing reproduction modulation and inputting the model; according to the method, the defects of insufficient concealment, black box adaptability, physical robustness and defensive resistance of an existing method are overcome, high-concealment, high-robustness and wide-application backdoor injection is realized, and reliable support is provided for security evaluation of a speaker recognition system.
Owner:SUZHOU VOCATIONAL INSTITUTE OF INDUSTRIAL TECHNOLOGY

Speaker recognition method based on bipartite graph matching and electronic device

The application discloses a speaker recognition method based on bipartite graph matching and an electronic device, and belongs to the technical field of audio recognition. The method comprises the following steps: acquiring an audio stream, and splitting the audio stream into continuous target time length audio segments in real time; determining a voiceprint embedding vector corresponding to each target time length audio segment; determining an embedding matrix based on the voiceprint embedding vector; determining a target voiceprint feature corresponding to each target time length audio segment according to a bipartite graph method; calculating the similarity between each voiceprint embedding vector; grouping each voiceprint embedding vector based on the similarity between each voiceprint embedding vector, clustering the voiceprint embedding vectors corresponding to the same speaker, obtaining a clustering result, and generating a target recognition report. The application can reduce the calculation complexity while improving the recognition accuracy.
Owner:BEIJING TONGXIANG QIANFANG TECHNOLOGY CO LTD

Channel-compensated low-level features for speaker recognition

A system for generating channel-compensated features of a speech signal includes a channel noise simulator that degrades the speech signal, a feed forward convolutional neural network (CNN) that generates channel-compensated features of the degraded speech signal, and a loss function that computes a difference between the channel-compensated features and handcrafted features for the same raw speech signal. Each loss result may be used to update connection weights of the CNN until a predetermined threshold loss is satisfied, and the CNN may be used as a front-end for a deep neural network (DNN) for speaker recognition / verification. The DNN may include convolutional layers, a bottleneck features layer, multiple fully-connected layers, and an output layer. The bottleneck features may be used to update connection weights of the convolutional layers, and dropout may be applied to the convolutional layers.
Owner:PINDROP SECURITY INC

Unsupervised construction method and device for speaker recognition model, equipment and medium

The invention discloses an unsupervised construction method and device for a speaker recognition model, equipment and a medium, and relates to the technical field of information processing, and the method comprises the steps: extracting the features of a speaker of an original audio data set through a voice pre-training model; performing clustering analysis based on the speaker feature set through a target clustering algorithm, and screening clustering results to obtain a first pseudo-tag set and a first pseudo-tag data set composed of the first pseudo-tag set and the original audio data set; taking the first pseudo tag set as a supervision signal, training the initial speaker recognition model by using the first pseudo tag data set to obtain a first trained speaker recognition model, predicting the original audio data set, and correcting a first prediction result; and training the first trained speaker recognition model according to the first corrected pseudo tag set to obtain a first optimized speaker recognition model, and performing iterative optimization based on the original audio data set until a preset stop condition is met to obtain a target speaker recognition model.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Speaker recognition method and device, electronic equipment and storage medium

The invention provides a speaker recognition method and device, electronic equipment and a storage medium. The speaker recognition method comprises the following steps: acquiring a first communication record; according to the first communication record, determining a candidate participant set of the first video conference, the candidate participant set comprising a plurality of candidate participants; and speaker recognition is carried out according to the audio data of the first video conference and the preset voiceprint information of each candidate participant in the candidate participant set, and a speaker recognition result is determined. According to the speaker recognition method, the voiceprint matching range is more accurate, and the accuracy and efficiency of speaker recognition in a video conference are improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

End-to-end voiceprint recognition method and voiceprint recognition device

Disclosed are an end-to-end voiceprint recognition method and a voiceprint recognition device. The voiceprint recognition method comprises: based on received input speech, using a speaker voice extraction module of an end-to-end deep learning network to perform a speaker voice extraction task to extract speech features of a target speaker; and based on the speech features of the target speaker, using a speaker recognition module of the end-to-end deep learning network to perform a speaker recognition task to recognize the target speaker in the received input speech.
Owner:SAMSUNG (CHINA) SEMICONDUCTOR CO LTD +1

Speaker recognition methods, devices, equipment, and media based on pre-trained speech models

This application discloses a speaker recognition method, apparatus, device, and medium based on a pre-trained speech model, relating to the field of audio recognition technology. The method includes: inserting a target adapter after the target layer of a pre-trained speech model, and processing the target hidden representation output by the target layer using the target adapter; fusing the adapted representations to obtain a fused feature representation, and obtaining a main task loss based on the fused feature representation; acquiring several speech segments, acquiring target speech features corresponding to each speech segment, and adjusting the distance between the target speech features in the target vector space to obtain a content-identity decoupling contrast loss; adjusting the parameters of the target adapter according to the main task loss and the content-identity decoupling contrast loss, and using the adjusted adapter for speaker recognition. By introducing the content-identity decoupling contrast loss, interference from changes in speech content on speaker embedding is suppressed.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY