Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

318 results about "Automatic speech" patented technology

Formulaic language (previously known as automatic speech or embolalia) is a linguistic term for verbal expressions that are fixed in form, often non-literal in meaning with attitudinal nuances, and closely related to communicative-pragmatic context. Along with idioms, expletives and proverbs, formulaic language includes pause fillers (e.g., "Like", "Er" or "Uhm") and conversational speech formulas (e.g., "You've got to be kidding," "Excuse me?" or "Hang on a minute").

Automatic speech recognition using language model-generated context

Techniques for ASR processing using language model (LM)-generated context are described. A LM is prompted to generate words that are relevant for / may be included in a future user input. The prompt to the LM can include words from user interaction history, dialog history, dialog topic, user preferences, etc. The information included in the prompt may focus on rare or unique words rather than words that the ASR model is already confident in recognizing. The techniques can be plugged into an existing / pretrained ASR model and can be used with any existing / pretrained LM, thus saving resources needed to implement and maintain the components.
Owner:AMAZON TECH INC

Personalized Russian spoken language practice recommendation method and system based on artificial intelligence

The invention relates to the technical field of artificial intelligence education, in particular to a Russian spoken language practice personalized recommendation method and system based on artificial intelligence, and the method comprises the steps: 1, outputting a phoneme sequence with a timestamp through Russian automatic voice recognition; 2, collecting an exercise interruption position and repeated read-after behavior data; 3, generating a dynamic learner portrait; 4, mapping high-frequency errors in the learner portrait into abnormal path weights of map nodes; 5, a lattice tail error option and a non-matching body verb interference item are injected; 6, when the voice fluency attenuation of the learner exceeds a dynamic threshold value, the sentence complexity is reduced; and 7, calculating an error rate descent gradient based on the exercise completion data, and dynamically adjusting the abnormal path weight of the knowledge graph. Through audio stream analysis and syntax tree construction, the system can accurately identify errors of the learner in grammar, pronunciation and other aspects, and the learning efficiency is improved.
Owner:HARBIN UNIV

Speech recognition model training method and device, equipment and medium

The invention relates to a speech recognition technology, discloses a speech recognition model training method and device, equipment and a medium, and aims to improve the speech recognition rate in a noise environment. The method comprises the following steps: firstly, training an automatic speech recognition network to obtain a pre-trained network; a noise reduction module is introduced in front of an output classification layer, the noise reduction module is trained by taking embedded features output by the pre-trained automatic speech recognition network as a reference, and parameters of the pre-trained automatic speech recognition network are fixed during training to form an initial speech recognition model; and finally, noise reduction module parameters are fixed, the pre-trained automatic speech recognition network is retrained, and a final model is obtained. Through staged training and a parameter fixing strategy, training target conflicts among modules are avoided, and the training stability and the convergence speed are improved; the noise reduction module focuses on feature denoising required by recognition, is high in adaptability, has the characteristics of light weight and low delay, and can be widely applied to low-resource real-time scenes such as embedded equipment and edge computing.
Owner:WUXUE GUANGJI DATA TECHNOLOGY CO LTD

System for latency-aware orchestration and performance optimization in artificial intelligence telephone communication

A system for latency-aware orchestration and performance optimization in AI-driven telephone communication, consisting of: a speech capture unit configured to capture an analog audio signal from a telephone interface and convert the analog audio signal into a digital audio signal stream; a feature extraction unit that is operationally coupled with the speech acquisition unit and is configured to generate a feature representation of the digital audio signal stream through spectral decomposition, noise reduction, and temporal segmentation; an AI inference processor communicatively connected to the feature extraction unit, configured to run one or more AI models for automatic speech recognition, natural language understanding, and emotion recognition on the feature representation to generate intermediate results for inference; a latency orchestration controller coupled to the AI ​​inference processor, wherein the latency orchestration controller is configured to monitor latency across multiple processing stages, predict cumulative delay propagation using a hybrid latency estimation model, and orchestrate the execution scheduling of the AI ​​inference processor based on the predicted latency deviation; a performance optimization unit coupled with the latency orchestration controller and configured to dynamically adjust computational accuracy, inference batch size, and feature processing resolution based on latency thresholds and quality constraints set by the latency orchestration controller; and a transmission synchronization array configured to time-align the processed output generated by the AI ​​inference processor and transmit it to a remote communication node, with the transmission synchronization array maintaining deterministic time coordination between successive packets and the orchestrated inference results.
Owner:CHEEKURI KARTHIK CHAKRAVARTHY DULUTH

Customizable latency for automatic speech recognition

Techniques for customizable latency, from the customer's side, for automatic speech recognition (ASR) are described. In particular, the customer may specify a parameter that controls how fast or how slow the customer's media content will be streamed or processed. Slower processing means higher accuracy, with near real-time latency, while faster processing means lower accuracy, but offers much lower latency (e.g., less than 600 ms). Enabling tuning of the latency-versus-accuracy tradeoff of the ASR system offers customers the flexibility to meet varying needs for different ASR applications.
Owner:AMAZON TECH INC

Neural network based conversation-aware automatic speech recognition

A system uses a machine learning based model such as a neural network for transcribing audio inputs. The system receives a set of audio inputs representing utterances of a conversation. For each conversation, the system determines a dialogue state for each utterance. The system uses a hierarchical language model for transcribing audio inputs of an online conversation using the received conversations. The hierarchical language model includes a top-level language model and a plurality of lower-level language model. The training is performed by (1) training the top-level language model using sequences of corresponding dialogue state, each sequence of dialogue states for a conversation, and (2) for each dialogue state, training a lower-level language model using utterances having that dialogue state. The system executes the hierarchical language model to transcribe audio input of new conversations.
Owner:INTERACTIONS LLC (US)

Injecting short-term spectro-temporal knowledge into automatic speech recognition models

Mechanisms are provided for training an Automatic Speech Recognition (ASR) model. An automatic speech recognition (ASR) computer model is trained based on full utterance training data, the ASR computer model having an audio encoder, text predictor, and a joint network which combines outputs of both the audio encoder and the text predictor. Fine-tuning training of the ASR computer model is performed by a knowledge distillation framework at least by: executing a chunking operation on full utterance data to generate a plurality of data chunks corresponding to full utterances in the full utterance data; and executing a knowledge distillation operation with two encoder embeddings. A first encoder embedding is obtained from the full utterance data and a second encoder embedding is obtained from the data chunks. Operational parameters of the trained ASR model are updated based on a loss determined from the first encoder embedding and second encoder embedding.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Automatic voice response fault diagnosis method and system based on multi-source information fusion

The invention discloses an automatic voice response fault diagnosis method and system based on multi-source information fusion. The method comprises the following steps: receiving user voice input, synchronously obtaining user side intelligent electric meter data, meteorological environment information and historical service records, and extracting multi-modal features; semantic association is realized through an electric power service knowledge graph, and voice features, equipment data and environmental parameters are fused through space-time alignment; a machine learning dynamic decision tree engine is combined with a power consumption behavior analysis model to form a fault reasoning model, and a diagnosis result is effectively verified; and outputting the grading disposal scheme and triggering intelligent chemical order distribution. According to the invention, various information sources of the calling platform are fully utilized, a fault reasoning model with high accuracy is formed, the technical problems of inaccurate fault positioning and insufficient multi-source data collaboration of traditional voice response in power customer service are effectively solved, the diagnosis accuracy of customer power consumption problems is effectively improved, and the average processing time is shortened.
Owner:国家电网有限公司客户服务中心

Semiautomated relay method and apparatus

ActiveUS12482458B2Special service for subscribersTelephone sets with user guidance/featuresElectrophonic hearingAutomatic speech
A captioning method for presenting captions to an assisted user (AU) during communication with a hearing user (HU) where the assisted user uses a captioned device and the hearing user uses a hearing user's device to facilitate the communication, the captioned device including a display screen and a speaker for presenting captions and broadcasting the hearing user's voice signals, respectively, the method comprising the steps of during an ongoing call between the AU and the HU, using an automated speech recognition (ASR) engine to generate initial ASR captions associated with the HU's voice signal, assessing at least one caption quality factor associated with prior initial ASR captions generated during the ongoing call, delaying broadcast of HU voice signal to the AU and based on the at least one caption quality factor, adjusting a duration of the HU voice signal broadcast delay.
Owner:ULTRATEC INC

Language model arbitration for natural language processing

Devices and techniques are generally described for arbitration between LLM-based and intent-based natural language processing flows. In various examples, an automatic speech recognition (ASR) component may generate first ASR output data representing the first natural language input. A first machine learning model may select an LLM-based natural language processing flow. The first machine learning model may be trained to select between at least the LLM-based and non-LLM-based natural language processing flows. The first ASR output data may be processed using the LLM-based processing flow. The LLM-based processing flow may generate first executable data.
Owner:AMAZON TECH INC

Automatic De-identification of Sensitive Conversational Audio Data

Techniques for automatically de-identifying sensitive information in audio conversations by combining un-transcribed voice activity detection (VAD) with large language model (LLM) analysis are disclosed. An audio de-identification system processes speech-to-text transcriptions while identifying segments where automatic speech recognition (ASR) failed to transcribe spoken content. These un-transcribed segments are represented as placeholders in prompts sent to an LLM, which analyzes the surrounding textual context to determine if sensitive information (such as PII or PHI) was likely spoken during these gaps. When sensitive content is identified, the system modifies the corresponding audio segments through an audio identification tactic. This approach addresses the technical challenge of incomplete de-identification in automated audio processing by leveraging LLMs' contextual understanding to detect sensitive information in segments that traditional ASR systems miss, particularly in scenarios involving poor audio quality or diverse accents. The result is a more comprehensive and reliable audio de-identification system.
Owner:ORACLE INT CORP

Speaker recognition method in medical real-time speech recognition scene

ActiveCN121938378Aconsistent with auditory logicReduce the risk of number fluctuationsSpeech recognitionAutomatic speechSpeech sound
The invention discloses a speaker recognition method in a medical real-time speech recognition scene, and the method specifically comprises the steps: carrying out the parallel distribution of a received PCM audio stream through an audio diverter in a gateway layer, and transmitting the audio stream to an automatic speech recognition link and a speaker separation link; in the automatic speech recognition link, outputting a recognition text and Token-level timestamps corresponding to each minimum semantic unit in the recognition text; in the speaker separation link, extracting a corresponding speaker embedding vector based on a preloaded speaker embedding model, and performing online clustering processing on the embedding vector to generate a speaker identity identifier; token-level boundary adsorption labeling processing is executed based on the Token-level timestamp and an online clustering result, so that a speaker switching position is determined; and carrying out one-to-one mapping on the speaker tags and the recognition text according to the Token-level timestamps, merging Tokens with the same speaker tags continuously, and outputting a transliteration text with the speaker tags.
Owner:ZOE SOFT CORP LTD

Predictor-corrector method for including speech hints in automatic speech recognition

A method comprises: receiving an automatic speech recognition (ASR) text transcript generated by an ASR process that encoded input audio into audio encodings and converted the audio encodings to ASR words of the ASR text transcript that correspond to the audio encodings; receiving speech hints for non-standard words, and generating alternative words for an ASR word of the ASR words based on the speech hints; correlating an audio encoding of the audio encodings that corresponds to the ASR word against the ASR word and each of the alternative words, to produce correspondence scores; selecting an output word among the ASR word and the alternative words based on the correspondence scores; and providing the output word to a corrected transcript.
Owner:CISCO TECHNOLOGY INC

Training of speech recognition systems

A method may include obtaining first audio data of a first communication session between a first and second device and during the first communication session, obtaining a first text string that is a transcription of the first audio data and training a model of an automatic speech recognition system using the first text string and the first audio data. The method may further include in response to completion of the training, deleting the first audio data and the first text string and after deleting the first audio data and the first text string, obtaining second audio data of a second communication session between a third and fourth device and during the second communication session obtaining a second text string that is a transcription of the second audio data and further training the model of the automatic speech recognition system using the second text string and the second audio data.
Owner:SORENSON IP HOLDINGS LLC

Captioning videos with multiple cross-modality teachers

Automatic captioning pipelines and methods for automatically annotating video data with subtitles, which can be obtained using automatic speech recognition (ASR). An automatic captioning pipeline with inputs of multimodal data scales up the dataset of high-quality video-caption pairs. The automatic captioning pipeline generates video-caption pairs by establishing and using a large video-language dataset along with an automatic captioning approach leveraging multimodal inputs, such as textual video description, subtitles, and individual video frames.
Owner:SNAP INC

End-cloud collaborative three-level decision-making full duplex voice dialogue dynamic management and control method

The invention discloses an end-cloud collaborative three-level decision-making full duplex voice dialogue dynamic management and control method, and relates to the technical field of voice interaction control, and the method comprises the steps: an end side collects microphone audio, obtains loudspeaker reference audio, judges a broadcast state, and switches a voice activity detection strategy; end-side framing detection is carried out, voice segment events are output through a three-state state machine, characteristics such as duration, relative volume and broadcast overlapping degree are extracted, and candidate interruption events are generated in combination with an automatic voice recognition intermediate text; the end side coarse screening reports a gray area event to the cloud side semantic research and judgment, and pre-suppression, buffer foresight contraction and block buffer are executed during waiting; and the cloud side issues an instruction containing a period of validity, a deduplication key and an interruption type, and the execution side stops generating, synthesizing and playing and empties buffer so as to play a completion receipt, submit a dialogue history and roll back unplayed content, thereby reducing false triggering and missed triggering, shortening interruption delay, reducing residual broadcast and inhibiting context drift.
Owner:SUZHOU MENGWU INTELLIGENT TECHNOLOGY CO LTD

Vehicle type robot voice navigation control system and control method thereof

The invention relates to the technical field of voice recognition and natural language processing, in particular to a vehicle type robot voice navigation control system and a control method thereof.The method takes voice as a man-machine interaction entrance and comprises the following steps that S1, environment audio is continuously monitored and preprocessed, including noise reduction, reverberation removal, pre-emphasis and framing windowing; when a preset wake-up word is detected, sending a wake-up signal to a main control unit through a serial port, and entering a command identification window with a fixed duration; s2, in the command recognition window, performing automatic voice recognition and natural language processing on the input voice, analyzing user intentions and key entities, and classifying analysis results into a mapping instruction, a motion control instruction or a navigation instruction; according to the invention, an integrated system with off-line voice as an entrance, a modular architecture as a core and closed-loop control and multi-mode feedback as guarantee is constructed, so that the limitation of the prior art is effectively overcome.
Owner:SHENZHEN ZHENSHI INTELLIGENT CONTROL CO LTD

Automated speech recognition to support context-aware intent recognition

A computing system for determining a user intent from a speech input to effect a user intended action is provided. The computer system comprises a set of processing nodes and a controller module. Each processing node is capable of understanding only a subset of words directly relevant to a particular context. The processing nodes of the set are arranged to receive a same speech input, and each processing node attempts to interpret the input, based on its subset of words, to extract therefrom an output indicative of user intent. Each node is unable to interpret any portion of the input containing a word outside of its subset. The controller module receives the outputs from the set of processing nodes and determine a most likely user intent based on the outputs.
Owner:OAKSPIRE LTD

Automatic movie and television video script extraction method based on multi-modal large model

The invention relates to an automatic movie and television video script extraction method based on a multi-modal large model, and belongs to the field of artificial intelligence and video content analysis. Aiming at the problem that an existing automatic speech recognition tool cannot generate a structured script, the method comprises the following steps: firstly, carrying out multi-mode decomposition on an input video, extracting frames by adopting a scene self-adaptive strategy, and establishing sound and picture timestamp alignment mapping; then extracting features through a CLIP-ViT visual feature encoder and a Whisper audio feature encoder, and associating semantics by using a cross-modal attention mechanism; role identity recognition, scene type recognition and action character description are achieved, and finally a script file conforming to the standardization specification is generated and comprises a scene title, a time code mark and a special narrative mark. Compared with a mode of manually dictating marks and an automatic voice recognition tool, the method effectively improves the accuracy and efficiency of movie and television video script extraction.
Owner:BEIJING INST OF COMP TECH & APPL

Enrollment-free automated speech recognition in multi-speaker environments quality metrics

The present disclosure relates to systems and methods for enrollment-free automated speech recognition (ASR) in multi-speaker environments. A system can process mixed audio signals containing speech from a target speaker and one or more interfering speakers. By applying acoustic characteristics such as room impulse responses (RIRs) and / or speech-to-interference energy ratios, the system can simulate environments to improve speech separation and recognition accuracy. A neural network model can be trained to identify and transcribe the speech of target speakers and / or secondary speakers, while filtering out interference. The system can update the model using ground-truth text and performance feedback, thereby facilitating real-time or near real-time ASR in multi-speaker environments without requiring predefined speaker data.
Owner:NVIDIA CORP

Systems for and methods of speech diarization using artificial intelligence models with sorting functionality

In various examples, multi-speaker audio is diarized using artificial intelligence models including a sorting functionality. Sorting is performed based on the first time a speaker is indicated as speaking and / or based on the variance of a dimension of a speech embedding. Sorting speech sequences has the advantage of requiring fewer computations of cross-entropy loss during training and / or allowing diarization models to focus on the difference between speakers. Diarized speech may be used to create a transcript in conjunction with automatic speech recognition models.
Owner:NVIDIA CORP

Resume generation and optimization method and system based on multiple rounds of natural language interaction

The invention discloses a resume generation and optimization method and system based on multi-round natural language interaction, and belongs to the technical field of natural language processing, the resume generation and optimization method and system based on multi-round natural language interaction comprises the following specific steps: 1, receiving initial information input by a user through a natural language mode; the natural language mode comprises voice input, text input and multi-mode input containing images, and if the input is voice input, the input is converted into initial text information through an automatic voice recognition engine. Through multiple rounds of natural language interaction, the resume input and updating threshold is remarkably reduced, and a user can generate a first version of resume and match posts in real time only by short voices. The system actively excavates the potential advantages of the user, guides and complements key information, supports multi-modal input and dynamic updating, and effectively improves the resume quality, the matching efficiency and the user experience.
Owner:THORSON (XIONGAN) ENTERPRISE MANAGEMENT CONSULTING CO LTD

Multi-modal behavior data processing system

The invention discloses a multi-modal behavior data processing system, which comprises a data acquisition module, a feature extraction layer, a dynamic attention weight layer, a multi-modal fusion layer and a downstream task decision-making layer, the data acquisition module guides human-computer interaction through an international neurological and mental interview tool matched with a DSM-5 standard and acquires audio and video stream data; the feature extraction layer extracts a video feature vector (including facial action unit activation intensity and the like), an audio feature vector (including Mel frequency cepstrum coefficient and the like) and a text feature vector (generated by a deep language model after automatic speech recognition transcription) in parallel; the dynamic attention weight layer is combined with data quality, symptomatic priori knowledge and cross-modal correlation to generate a dynamic fusion weight; weighting, splicing and dimensionality reduction are carried out on the multi-modal fusion layer to obtain a fusion feature vector; and the downstream task decision-making layer completes evaluation and generates a multi-modal behavioral index evaluation report. The system is deployed in a non-intrusive manner, the risk is controllable, and the evaluation robustness and accuracy can be improved.
Owner:NEW MAYO HEALTH MANAGEMENT RESEARCH INSTITUTE (CHONGQING) CO LTD +1

Real-time use of multiple parallel automatic speech recognition (ASR) modules in a conversational artificial intelligence (AI) architecture

As an example, a conversational artificial intelligence (AI) that is in a conversation receives a human response from a human, augments the human response to create an augmented response, determines a context of the conversation, and provides the context and the augmented response to a plurality of automatic speech recognition (ASR) modules that individually process the augmented response in parallel. The conversational AI receives a plurality of intermediate text outputs from the plurality of ASR modules, wherein individual intermediate text outputs of the plurality of intermediate text outputs are received from individual ones of the ASR modules. A reconciliation AI performs a contextual reconciliation of the plurality of intermediate text outputs based at least in part on the context of the conversation to create a final text output. The conversational AI provides, in real-time, an artificial intelligence response to the human based on the final text output.
Owner:HEALTHGPT INC DBA HIPPOCRATIC AI

Generating suggested modifications for configuring and training an automatic speech recognition model

Methods and systems for receiving a trained machine learning model, receiving a test dataset, wherein the test dataset is used to evaluate the trained machine learning model, generating, based on the test dataset and the trained machine learning model, one or more suggested modifications to at least one aspect of configuring and training of the trained machine learning model, and applying the one or more suggested modifications to at least one aspect of configuring and training of the trained machine learning model.
Owner:INFINEON TECHNOLOGIES AMERICAS CORP

Teaching interaction quality evaluation method and system based on large language model

The invention relates to the field of teaching interaction quality evaluation, in particular to a teaching interaction quality evaluation method and system based on a large language model. The method comprises the following steps: audio transcription: converting classroom audio into an original transcription text through voice activity detection, speaker classification, automatic voice recognition and punctuation recovery; transcriptional refining: performing context-based text error correction on the original transcriptional text by using a large language model in combination with a preschool education field knowledge base to generate a refined transcriptional text; a quality evaluation step: based on a preschool education quality evaluation scale, using few sample example guidance and thinking chain reasoning for each scoring point, judging whether a voice segment conforming to the scoring point exists in the refined transcriptional text, performing binary scoring, and determining whether the voice segment conforms to the scoring point; and generating an interactive quality evaluation report containing the standard-reaching rate of each evaluation dimension, teaching bright spot analysis and staged teaching optimization suggestions. The evaluation efficiency is remarkably improved. The method is suitable for teaching interaction quality evaluation.
Owner:THE CHINESE UNIV OF HONG KONG (SHENZHEN)

Family storm risk identification method, system and device based on artificial intelligence and medium

The invention relates to a home violation risk identification method, system and device based on artificial intelligence and a medium. The method comprises the following steps: preprocessing audio data to obtain a voice segment; performing automatic voice recognition on the voice segments to obtain a dialogue text sequence, and performing acoustic feature extraction to obtain a time sequence acoustic feature sequence; performing key feature extraction based on context semantics and risk knowledge on the dialogue text sequence to generate a semantic feature vector; performing deep emotion mode learning on the time sequence acoustic feature sequence to generate an acoustic feature vector; performing multi-modal fusion on the semantic feature vector and the acoustic feature vector to obtain a fusion feature vector; and carrying out collaborative risk judgment on the fused feature vector to generate a result containing high, medium and low risk levels and corresponding judgment confidence coefficients. By adopting the method, the limitation of single modal analysis can be overcome, and the home violence risk can be identified more comprehensively and accurately.
Owner:天津仁爱学院

Voice call emotion and intention joint recognition system based on deep learning

The invention discloses a voice call emotion and intention combined recognition system and method based on deep learning, and relates to the technical field of artificial intelligence voice processing. The system comprises a voice data receiving module, an audio feature extraction module, a joint recognition module and a result output module. The audio feature extraction module uses a timestamp generated by automatic speech recognition (ASR) to forcibly align acoustic features and linguistic features in a time dimension through a time alignment unit. And the joint recognition module adopts a two-way encoder to encode the aligned multi-modal features, captures interaction information of acoustics and texts through a multi-modal attention fusion layer, establishes connection between an emotion layer and an intention layer by utilizing a gating interaction unit, and injects the emotion features as auxiliary information into an intention recognition task. Through feature-level deep fusion and gating interaction between tasks, the problems that in the prior art, modals are not aligned, and the intrinsic strong relevance between emotions and intentions is ignored are solved, voice intonation features are effectively utilized to correct the intention recognition result, and the recognition accuracy and robustness in a voice call scene are remarkably improved.
Owner:杭州智慧沟通智能科技有限公司

Speech recognition method and device, electronic equipment and storage medium

PendingCN121281502ASpeech recognitionSemantic representationSpeech code
The invention provides a speech recognition method and device, electronic equipment and a storage medium, and relates to the technical field of natural language processing, an adopted target automatic speech recognition model generates context semantic representation through a current text prefix sequence, and combines the context semantic representation with the current text prefix sequence and acoustic features to obtain a speech recognition result. A target text sequence is obtained through step-by-step prediction in an autoregression mode. Context semantic representation is introduced into the target automatic speech recognition model, the language switching moment can be accurately judged when the input speech signal is speech code conversion speech, speech recognition is carried out in time according to a new language during language switching, the recognition precision and robustness of a language switching boundary can be effectively improved, and the speech recognition efficiency is improved. And the speech code conversion speech recognition effect is improved. Moreover, context semantic representation is introduced during prediction, the challenge of ambiguity or ambiguity of acoustic signals can be overcome, the accuracy of the target text sequence is improved, and errors caused by untimely language model switching are reduced.
Owner:IFLYTEK CO LTD

Context-based automatic speech recognition processing

Techniques for biasing for entities during automatic speech recognition (ASR) processing are described. In some embodiments, a system implements a gating component that is configured to switch on and off entity biasing on an audio frame basis when processing a spoken input. The gating component processes an audio frame to determine whether the audio frame likely includes a representation of a custom entity. Based on the determination, a biasing component, which is configured to generate entity embeddings, may be turned on or off. In this manner, entity biasing does not run on every audio frame, but only on the audio frames where it can be helpful in increasing ASR accuracy.
Owner:AMAZON TECH INC