Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

472 results about "Automatic speech" patented technology

Formulaic language (previously known as automatic speech or embolalia) is a linguistic term for verbal expressions that are fixed in form, often non-literal in meaning with attitudinal nuances, and closely related to communicative-pragmatic context. Along with idioms, expletives and proverbs, formulaic language includes pause fillers (e.g., "Like", "Er" or "Uhm") and conversational speech formulas (e.g., "You've got to be kidding," "Excuse me?" or "Hang on a minute").

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Natural language processing system

Techniques for processing with respect to a user input as contextual information is available are described. A system generates a first task prediction using first context data that is available when a user input is received. The system generates a second task prediction (e.g., updated first task prediction) when second context data is received, and then further generates a third task prediction when third context data is received. Example first context data may include device type information, time information, location, etc. Example second context data may include automatic speech recognition (ASR) data. Example third context data may include natural language understanding (NLU) data. Using the third task prediction, the system generates an output responsive to the user input.
Owner:AMAZON TECH INC

Automatic speech recognition using language model-generated context

Techniques for ASR processing using language model (LM)-generated context are described. A LM is prompted to generate words that are relevant for / may be included in a future user input. The prompt to the LM can include words from user interaction history, dialog history, dialog topic, user preferences, etc. The information included in the prompt may focus on rare or unique words rather than words that the ASR model is already confident in recognizing. The techniques can be plugged into an existing / pretrained ASR model and can be used with any existing / pretrained LM, thus saving resources needed to implement and maintain the components.
Owner:AMAZON TECH INC

Use Of Modulation Spectrums In Automatic Speech Recognition Models

Techniques for speech recognition models using modulation spectrum are disclosed herein. A modulation spectrum is generated from time series data output of an encoder layer of a speech recognition model and used as input into a decoder layer of the speech recognition model to improve accuracy of the model such as for recognizing subword units. The modulation spectrum is determined by applying a convolution filter to the output of the encoder layer of the speech recognition model. The time series data and / or the modulation spectrum can be normalized. A rectified linear unit activation function can be applied to the output of the convolution filter. The output of the encoder layer may be residually connected to the output of the rectified linear unit activation function prior to being input into the decoder layer.
Owner:ORACLE INT CORP

Fragmented semantic understanding order building system and method based on knowledge graph

The invention belongs to the technical field of semantic understanding order establishment systems, and discloses a fragmented semantic understanding order establishment system and method based on a knowledge graph. The method comprises the following steps: collecting multi-source data of the same user in the same preset time window to obtain a multi-modal data set with a timestamp; performing automatic voice recognition and screenshot character recognition on the multi-modal data set to obtain text segments; standardizing the text fragments to obtain a text fragment set with timestamps; identifying entities in the text fragment set, and merging the identified entities to obtain a merged entity set; positioning and merging nodes corresponding to entities in the entity set in a pre-constructed domain knowledge graph, and generating a temporary knowledge sub-graph in combination with a timestamp sequence and a context relationship in the text fragment set; according to the invention, the manual information comparison cost of operation and maintenance personnel is reduced, and the fault diagnosis and solution efficiency is improved.
Owner:SHANGHAI SUQING SOFTWARE CO LTD

Natural language question generation

Techniques for generating a natural language prompt to further a goal of a dialog, are described. During a dialog, the system receives one or more user inputs including a user question, a user response to the question, and a request to generate a further question following the response. The system determines ASR output data corresponding to the user inputs, and determines dialog history data of the dialog. Using the ASR output data and the dialog history data, the system determines a category and an explanation of relevance corresponding to the category. Using the ASR output data, the dialog history, the category, and the explanation, the system determines the further question to be output to the user.
Owner:AMAZON TECH INC

Personalized Russian spoken language practice recommendation method and system based on artificial intelligence

The invention relates to the technical field of artificial intelligence education, in particular to a Russian spoken language practice personalized recommendation method and system based on artificial intelligence, and the method comprises the steps: 1, outputting a phoneme sequence with a timestamp through Russian automatic voice recognition; 2, collecting an exercise interruption position and repeated read-after behavior data; 3, generating a dynamic learner portrait; 4, mapping high-frequency errors in the learner portrait into abnormal path weights of map nodes; 5, a lattice tail error option and a non-matching body verb interference item are injected; 6, when the voice fluency attenuation of the learner exceeds a dynamic threshold value, the sentence complexity is reduced; and 7, calculating an error rate descent gradient based on the exercise completion data, and dynamically adjusting the abnormal path weight of the knowledge graph. Through audio stream analysis and syntax tree construction, the system can accurately identify errors of the learner in grammar, pronunciation and other aspects, and the learning efficiency is improved.
Owner:HARBIN UNIV

Unified speech recognition models for diacriticized languages

Disclosed are apparatuses, systems, and techniques that leverage one or more artificial intelligence models for efficient automatic speech recognition (ASR) of speech in a diacritized language. The techniques include processing, using an ASR model, audio frame(s) encoding a speech in the diacritized language to generate, for a transcription token (TT) of the speech, likelihoods that the TT corresponds to various vocabulary tokens that include both non-diacritized and diacritized tokens of the language, and generating, using the likelihoods, a transcription of the speech.
Owner:NVIDIA CORP

Speech recognition model training method and device, equipment and medium

The invention relates to a speech recognition technology, discloses a speech recognition model training method and device, equipment and a medium, and aims to improve the speech recognition rate in a noise environment. The method comprises the following steps: firstly, training an automatic speech recognition network to obtain a pre-trained network; a noise reduction module is introduced in front of an output classification layer, the noise reduction module is trained by taking embedded features output by the pre-trained automatic speech recognition network as a reference, and parameters of the pre-trained automatic speech recognition network are fixed during training to form an initial speech recognition model; and finally, noise reduction module parameters are fixed, the pre-trained automatic speech recognition network is retrained, and a final model is obtained. Through staged training and a parameter fixing strategy, training target conflicts among modules are avoided, and the training stability and the convergence speed are improved; the noise reduction module focuses on feature denoising required by recognition, is high in adaptability, has the characteristics of light weight and low delay, and can be widely applied to low-resource real-time scenes such as embedded equipment and edge computing.
Owner:WUXUE GUANGJI DATA TECHNOLOGY CO LTD

System for latency-aware orchestration and performance optimization in artificial intelligence telephone communication

A system for latency-aware orchestration and performance optimization in AI-driven telephone communication, consisting of: a speech capture unit configured to capture an analog audio signal from a telephone interface and convert the analog audio signal into a digital audio signal stream; a feature extraction unit that is operationally coupled with the speech acquisition unit and is configured to generate a feature representation of the digital audio signal stream through spectral decomposition, noise reduction, and temporal segmentation; an AI inference processor communicatively connected to the feature extraction unit, configured to run one or more AI models for automatic speech recognition, natural language understanding, and emotion recognition on the feature representation to generate intermediate results for inference; a latency orchestration controller coupled to the AI ​​inference processor, wherein the latency orchestration controller is configured to monitor latency across multiple processing stages, predict cumulative delay propagation using a hybrid latency estimation model, and orchestrate the execution scheduling of the AI ​​inference processor based on the predicted latency deviation; a performance optimization unit coupled with the latency orchestration controller and configured to dynamically adjust computational accuracy, inference batch size, and feature processing resolution based on latency thresholds and quality constraints set by the latency orchestration controller; and a transmission synchronization array configured to time-align the processed output generated by the AI ​​inference processor and transmit it to a remote communication node, with the transmission synchronization array maintaining deterministic time coordination between successive packets and the orchestrated inference results.
Owner:CHEEKURI KARTHIK CHAKRAVARTHY DULUTH

Customizable latency for automatic speech recognition

Techniques for customizable latency, from the customer's side, for automatic speech recognition (ASR) are described. In particular, the customer may specify a parameter that controls how fast or how slow the customer's media content will be streamed or processed. Slower processing means higher accuracy, with near real-time latency, while faster processing means lower accuracy, but offers much lower latency (e.g., less than 600 ms). Enabling tuning of the latency-versus-accuracy tradeoff of the ASR system offers customers the flexibility to meet varying needs for different ASR applications.
Owner:AMAZON TECH INC

Neural network based conversation-aware automatic speech recognition

A system uses a machine learning based model such as a neural network for transcribing audio inputs. The system receives a set of audio inputs representing utterances of a conversation. For each conversation, the system determines a dialogue state for each utterance. The system uses a hierarchical language model for transcribing audio inputs of an online conversation using the received conversations. The hierarchical language model includes a top-level language model and a plurality of lower-level language model. The training is performed by (1) training the top-level language model using sequences of corresponding dialogue state, each sequence of dialogue states for a conversation, and (2) for each dialogue state, training a lower-level language model using utterances having that dialogue state. The system executes the hierarchical language model to transcribe audio input of new conversations.
Owner:INTERACTIONS LLC (US)

Intelligent call dynamic response method integrating ASR and emotion recognition

The invention discloses an intelligent call dynamic response method and system fusing ASR (automatic voice recognition) and emotion recognition, and is applied to an interaction scene of a voice robot and a call center. The robot response strategy is dynamically adjusted by synchronously analyzing the text content (ASR) and emotional characteristics (such as intonation, speech speed and energy) of the user voice in real time. When the negative emotion is detected, a manual seat is automatically triggered to switch over or switch the pacifying verbal skill; and for the positive emotion of the high-value customer, a precision marketing module is started. The method solves the problems that an existing call center is single in response mode and cannot sense the emotion of a user, and the customer satisfaction and the service conversion rate are remarkably improved.
Owner:SHANGHAI ZHAOKUN INFORMATION TECHNOLOGY CO LTD

Injecting short-term spectro-temporal knowledge into automatic speech recognition models

Mechanisms are provided for training an Automatic Speech Recognition (ASR) model. An automatic speech recognition (ASR) computer model is trained based on full utterance training data, the ASR computer model having an audio encoder, text predictor, and a joint network which combines outputs of both the audio encoder and the text predictor. Fine-tuning training of the ASR computer model is performed by a knowledge distillation framework at least by: executing a chunking operation on full utterance data to generate a plurality of data chunks corresponding to full utterances in the full utterance data; and executing a knowledge distillation operation with two encoder embeddings. A first encoder embedding is obtained from the full utterance data and a second encoder embedding is obtained from the data chunks. Operational parameters of the trained ASR model are updated based on a loss determined from the first encoder embedding and second encoder embedding.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Automatic voice response fault diagnosis method and system based on multi-source information fusion

The invention discloses an automatic voice response fault diagnosis method and system based on multi-source information fusion. The method comprises the following steps: receiving user voice input, synchronously obtaining user side intelligent electric meter data, meteorological environment information and historical service records, and extracting multi-modal features; semantic association is realized through an electric power service knowledge graph, and voice features, equipment data and environmental parameters are fused through space-time alignment; a machine learning dynamic decision tree engine is combined with a power consumption behavior analysis model to form a fault reasoning model, and a diagnosis result is effectively verified; and outputting the grading disposal scheme and triggering intelligent chemical order distribution. According to the invention, various information sources of the calling platform are fully utilized, a fault reasoning model with high accuracy is formed, the technical problems of inaccurate fault positioning and insufficient multi-source data collaboration of traditional voice response in power customer service are effectively solved, the diagnosis accuracy of customer power consumption problems is effectively improved, and the average processing time is shortened.
Owner:国家电网有限公司客户服务中心

Semiautomated relay method and apparatus

A captioning method for presenting captions to an assisted user (AU) during communication with a hearing user (HU) where the assisted user uses a captioned device and the hearing user uses a hearing user's device to facilitate the communication, the captioned device including a display screen and a speaker for presenting captions and broadcasting the hearing user's voice signals, respectively, the method comprising the steps of during an ongoing call between the AU and the HU, using an automated speech recognition (ASR) engine to generate initial ASR captions associated with the HU's voice signal, assessing at least one caption quality factor associated with prior initial ASR captions generated during the ongoing call, delaying broadcast of HU voice signal to the AU and based on the at least one caption quality factor, adjusting a duration of the HU voice signal broadcast delay.
Owner:ULTRATEC INC

Language model arbitration for natural language processing

Devices and techniques are generally described for arbitration between LLM-based and intent-based natural language processing flows. In various examples, an automatic speech recognition (ASR) component may generate first ASR output data representing the first natural language input. A first machine learning model may select an LLM-based natural language processing flow. The first machine learning model may be trained to select between at least the LLM-based and non-LLM-based natural language processing flows. The first ASR output data may be processed using the LLM-based processing flow. The LLM-based processing flow may generate first executable data.
Owner:AMAZON TECH INC

Automatic De-identification of Sensitive Conversational Audio Data

Techniques for automatically de-identifying sensitive information in audio conversations by combining un-transcribed voice activity detection (VAD) with large language model (LLM) analysis are disclosed. An audio de-identification system processes speech-to-text transcriptions while identifying segments where automatic speech recognition (ASR) failed to transcribe spoken content. These un-transcribed segments are represented as placeholders in prompts sent to an LLM, which analyzes the surrounding textual context to determine if sensitive information (such as PII or PHI) was likely spoken during these gaps. When sensitive content is identified, the system modifies the corresponding audio segments through an audio identification tactic. This approach addresses the technical challenge of incomplete de-identification in automated audio processing by leveraging LLMs' contextual understanding to detect sensitive information in segments that traditional ASR systems miss, particularly in scenarios involving poor audio quality or diverse accents. The result is a more comprehensive and reliable audio de-identification system.
Owner:ORACLE INT CORP

Speaker recognition method in medical real-time speech recognition scene

ActiveCN121938378Aconsistent with auditory logicReduce the risk of number fluctuationsSpeech recognitionAutomatic speechSpeech sound
The invention discloses a speaker recognition method in a medical real-time speech recognition scene, and the method specifically comprises the steps: carrying out the parallel distribution of a received PCM audio stream through an audio diverter in a gateway layer, and transmitting the audio stream to an automatic speech recognition link and a speaker separation link; in the automatic speech recognition link, outputting a recognition text and Token-level timestamps corresponding to each minimum semantic unit in the recognition text; in the speaker separation link, extracting a corresponding speaker embedding vector based on a preloaded speaker embedding model, and performing online clustering processing on the embedding vector to generate a speaker identity identifier; token-level boundary adsorption labeling processing is executed based on the Token-level timestamp and an online clustering result, so that a speaker switching position is determined; and carrying out one-to-one mapping on the speaker tags and the recognition text according to the Token-level timestamps, merging Tokens with the same speaker tags continuously, and outputting a transliteration text with the speaker tags.
Owner:ZOE SOFT CORP LTD

Predictor-corrector method for including speech hints in automatic speech recognition

A method comprises: receiving an automatic speech recognition (ASR) text transcript generated by an ASR process that encoded input audio into audio encodings and converted the audio encodings to ASR words of the ASR text transcript that correspond to the audio encodings; receiving speech hints for non-standard words, and generating alternative words for an ASR word of the ASR words based on the speech hints; correlating an audio encoding of the audio encodings that corresponds to the ASR word against the ASR word and each of the alternative words, to produce correspondence scores; selecting an output word among the ASR word and the alternative words based on the correspondence scores; and providing the output word to a corrected transcript.
Owner:CISCO TECHNOLOGY INC

Classification method and system for data enhancement and hybrid expert mechanism feature selection

The invention discloses a classification method and system for data enhancement and hybrid expert mechanism feature selection. The system comprises an automatic speech recognition (ASR) module, a text-to-speech synthesis (TTS) module, a multi-modal feature extraction module, a hybrid expert mechanism (MoE) module, a common attention mechanism module, a feature fusion module and a classification module. According to the invention, a voice data enhancement module based on a voice-to-text (TTS) technology is utilized to improve data diversity and model generalization ability; by means of multi-level acoustic and text feature extraction, language changes are represented more comprehensively; a hybrid expert mechanism (MoE) is utilized to realize dynamic selection of multi-modal features, and the feature utilization efficiency is improved; according to the method, the fusion mode between different modal features is optimized by using a co-attention mechanism, the interaction expression ability between the features is enhanced, the recognition precision and the robustness of the system in a multi-modal environment are remarkably improved, and the defects are overcome.
Owner:SHANGHAI JIAOTONG UNIV

Query response interface with server side generative model(s)

Various implementations include processing, at a client device, an instance of audio data capturing a user voice query using an automatic speech recognition model to generate a sequence of instances of tokenizable query text. In many implementations, one or more instances of the sequence can be transmitted to a remote computing system prior to generating the entire sequence. In a variety of implementations, each instance in the sequence can be processed using a generative model which includes a streaming multi-head attention portion. Responsive output can be transmitted from the remote computing system to the client device, where the client device renders the responsive output to the user. In many implementations, the time between the user speaking the user query and the client device rendering the responsive output is reduced, thus decreasing latency in the system.
Owner:GOOGLE LLC

Training of speech recognition systems

A method may include obtaining first audio data of a first communication session between a first and second device and during the first communication session, obtaining a first text string that is a transcription of the first audio data and training a model of an automatic speech recognition system using the first text string and the first audio data. The method may further include in response to completion of the training, deleting the first audio data and the first text string and after deleting the first audio data and the first text string, obtaining second audio data of a second communication session between a third and fourth device and during the second communication session obtaining a second text string that is a transcription of the second audio data and further training the model of the automatic speech recognition system using the second text string and the second audio data.
Owner:SORENSON IP HOLDINGS LLC

Captioning videos with multiple cross-modality teachers

Automatic captioning pipelines and methods for automatically annotating video data with subtitles, which can be obtained using automatic speech recognition (ASR). An automatic captioning pipeline with inputs of multimodal data scales up the dataset of high-quality video-caption pairs. The automatic captioning pipeline generates video-caption pairs by establishing and using a large video-language dataset along with an automatic captioning approach leveraging multimodal inputs, such as textual video description, subtitles, and individual video frames.
Owner:SNAP INC

Automatic voice work order generation system and method based on multi-modal processing

ActiveCN120766676ASpeech recognitionTransliterationEngineering
The invention provides an automatic voice work order generation system and method based on multi-modal processing, and relates to the technical field of voice work order generation, and the system comprises a multi-modal input layer which is used for collecting voice data, equipment metadata and auxiliary modal data; the voice processing layer is used for carrying out audio preprocessing on the voice data to extract audio feature information and translating the audio feature information into text information; the semantic analysis layer is used for receiving equipment metadata, auxiliary modal data, audio feature information and transliteration text information, and performing multi-modal field adaptive analysis in combination with historical work order data; and the work order output layer is used for receiving the structural work order elements and then converting the structural work order elements into standard work orders. Through the method and the device, the technical problem that the work order generation efficiency is further influenced due to low voice recognition accuracy caused by high manual intervention dependence and multi-modal data processing splitting in the prior art can be solved, and the work order generation efficiency is improved by realizing end-to-end automatic conversion from the voice data to the structured work order.
Owner:BEIJING WANXUN BOTONG TECH DEV CO LTD

End-cloud collaborative three-level decision-making full duplex voice dialogue dynamic management and control method

The invention discloses an end-cloud collaborative three-level decision-making full duplex voice dialogue dynamic management and control method, and relates to the technical field of voice interaction control, and the method comprises the steps: an end side collects microphone audio, obtains loudspeaker reference audio, judges a broadcast state, and switches a voice activity detection strategy; end-side framing detection is carried out, voice segment events are output through a three-state state machine, characteristics such as duration, relative volume and broadcast overlapping degree are extracted, and candidate interruption events are generated in combination with an automatic voice recognition intermediate text; the end side coarse screening reports a gray area event to the cloud side semantic research and judgment, and pre-suppression, buffer foresight contraction and block buffer are executed during waiting; and the cloud side issues an instruction containing a period of validity, a deduplication key and an interruption type, and the execution side stops generating, synthesizing and playing and empties buffer so as to play a completion receipt, submit a dialogue history and roll back unplayed content, thereby reducing false triggering and missed triggering, shortening interruption delay, reducing residual broadcast and inhibiting context drift.
Owner:SUZHOU MENGWU INTELLIGENT TECHNOLOGY CO LTD

Vehicle type robot voice navigation control system and control method thereof

The invention relates to the technical field of voice recognition and natural language processing, in particular to a vehicle type robot voice navigation control system and a control method thereof.The method takes voice as a man-machine interaction entrance and comprises the following steps that S1, environment audio is continuously monitored and preprocessed, including noise reduction, reverberation removal, pre-emphasis and framing windowing; when a preset wake-up word is detected, sending a wake-up signal to a main control unit through a serial port, and entering a command identification window with a fixed duration; s2, in the command recognition window, performing automatic voice recognition and natural language processing on the input voice, analyzing user intentions and key entities, and classifying analysis results into a mapping instruction, a motion control instruction or a navigation instruction; according to the invention, an integrated system with off-line voice as an entrance, a modular architecture as a core and closed-loop control and multi-mode feedback as guarantee is constructed, so that the limitation of the prior art is effectively overcome.
Owner:SHENZHEN ZHENSHI INTELLIGENT CONTROL CO LTD

Artificial intelligence and machine learning for transcription and translation for media editing

Dialog in a language unfamiliar to an editor poses obvious challenges during the media editing process. It is nearly impossible to edit media containing spoken dialog without a clear comprehension of the underlying language. The methods described here use a combination of artificial intelligence and machine learning models to generate a language proxy in which the dialog is translated into a language that is familiar to the media editor. The editor is then able to edit the media composition in their own language. To generate an edited media composition with spoken dialog in the original language, the edited language proxy is synchronized with and linked back to the original media. The methods combine automatic speech recognition, translation, speech to text, and voice cloning together with existing non-AI technologies such as captioning and media relinking.
Owner:AVID TECHNOLOGY INC

Automated speech recognition to support context-aware intent recognition

A computing system for determining a user intent from a speech input to effect a user intended action is provided. The computer system comprises a set of processing nodes and a controller module. Each processing node is capable of understanding only a subset of words directly relevant to a particular context. The processing nodes of the set are arranged to receive a same speech input, and each processing node attempts to interpret the input, based on its subset of words, to extract therefrom an output indicative of user intent. Each node is unable to interpret any portion of the input containing a word outside of its subset. The controller module receives the outputs from the set of processing nodes and determine a most likely user intent based on the outputs.
Owner:OAKSPIRE LTD

All deep learning minimum variance distortionless response beamformer for speech separation and enhancement

A method, computer program, and computer system is provided for automated speech recognition. Audio data corresponding to one or more speakers is received. Covariance matrices of target speech and noise associated with the received audio data are estimated based on a gated recurrent unit-based network. A predicted target waveform corresponding to a target speaker from among the one or more speakers is generated by a minimum variance distortionless response function based on the estimated covariance matrices.
Owner:TENCENT AMERICA LLC