Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1422 results about "Utterance" patented technology

In spoken language analysis, an utterance is the smallest unit of speech. It is a continuous piece of speech beginning and ending with a clear pause. In the case of oral languages, it is generally, but not always, bounded by silence. Utterances do not exist in written language, however, only their representations do. They can be represented and delineated in written language in many ways.

Detecting and using non-textual information in human speech

PCT designated stage expiredWO2025141559A1Speech recognitionMachine learningEncoder decoderAcoustic model
Automatic recognition of non-verbal messages in speech, and in particular to detection or analysis of prosodic multilayered analysis of intonation units such as prosodic unit prototypes and their multi-labeled variations, may form a hierarchical classification for the analysis of non¬ verbal information or cues in speech. A speech captured by a microphone is fed to a weakly- supervised deep learning acoustic model for speech recognition and transcription, that may be based on encoder-decoder Transformer architecture, such as Whisper by OpenAI. The model is trained to output multiple words form the text in the captured speech, to identify Intonation Units (IUs) that include one or more words, and associate non-verbal labels to each of the IUs. The labels may indicate a prototype, a discourse function (such as a conversation action), an emotion, an emphasis, or an attitude, as well as a genre of a part of, or whole of, the entire captured speech.
Owner:YEDA RES & DEV CO LTD

System and method for augmenting training data for natural language to meaning representation language systems

Techniques for augmenting training data include accessing training data comprising a plurality of training examples comprising a first training example comprising a first natural language utterance and a first logical form for the first natural language utterance. A second natural language utterance is generated by adding or replacing one or more values in the first natural language utterance. A logical form for the second natural language utterance is generated. A second training example is generated, comprising the second natural language utterance and the logical form for the second natural language utterance. The training data is augmented by adding the second training example to the plurality of training examples to generate an augmented training data set. A machine learning model is trained to generate logical forms for utterances using the augmented training data set.
Owner:ORACLE INT CORP

Logical text passage generation and retrieval for retrieval-augmented generation

Techniques for logical text passage generation and retrieval for retrieval-augmented generation. The techniques involve processing markup language documents to generate logical text passages and their corresponding embeddings. These embeddings are indexed for efficient retrieval. Upon receiving a user utterance, a user query is formed and transformed into an embedding to query the index. Relevant text passages are identified and used to prompt a large language model (LLM), which generates a completion. This completion is then sent as a response to the user. The process effectively bridges user queries with relevant information through advanced embedding and natural language processing techniques, enabling accurate and contextually appropriate interactions within a user-agent dialogue framework.
Owner:AMAZON TECH INC

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Detecting corrupted speech in voice-based computer interfaces

Approaches are generally described for corrupted speech detection in voice-based computer interfaces. First input data including first audio data representing a user utterance may be received. First data representing the first audio data may be generated using a first encoder. First text data representing a transcription of the user utterance may be generated. Second data representing the first text data may be generated using a second encoder different from the first encoder. Third data may be generated by combining the first data and the second data. The third data may be sent to a classifier network trained to predict a relevant corruption state for speech processing inputs. The classifier network may determine that the first input data corresponds to a first corruption state.
Owner:AMAZON TECH INC

Audio and video recording-based ASR identification enhancement method

The invention discloses an ASR identification enhancement method based on audio and video recording. According to the method, the accuracy and compliance of voice recognition in the financial service interaction process are improved by fusing the audio and environment feature information in the banking business double-recording scene. The method comprises the following steps: firstly, constructing an acoustic model for a bank outlet environment, and extracting audio features and interaction scene information of conversation between a client and a worker; and then, designing a vocabulary recognition module special for the financial field, dynamically adjusting language model parameters according to professional term libraries and utterance modes of different business types, and effectively coping with key links such as financial product introduction, risk prompt and customer confirmation. Compared with a traditional ASR system, the voice recognition accuracy in the banking business handling process is remarkably improved, particularly, key term recognition and important information extraction are prominent, and more reliable technical support is provided for financial service standardized management and double-recording quality inspection.
Owner:GUANGZHOU BAIRUI NETWORK TECH CO LTD

Deepfake detection

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Owner:PINDROP SECURITY INC

Intent discovery using large language models

Systems and methods to cause an intent discovery system to identify new user intents without additional training. The system may comprise of two neural networks. The first neural network generates a prompt tailored to a particular domain (e.g., travel), and may include known intents pertinent to the domain selected examples from a training dataset to provide context to the prompt. The second neural network may use this prompt to identify intents from new utterances in the prompt. The identified intents that are not in the list of known intents are then used to update the database.
Owner:SERVICENOW INC

System and method for contextual analysis and metadata database generation for user-specific speech patterns

A system for contextual analysis and metadata database generation for user-specific speech patterns is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system splits the speech signal into a first set of audio frames, where each audio frame comprises an utterance of one or more words. The system determines a context associated with each word. In response, the system detects a context change between a first text and a second text. The system generates a contextually split set of frames by splitting the speech signal into a second set of audio frames according to the detected context changes.
Owner:BANK OF AMERICA CORP

Automated generation of targeted feedback using speech characteristics extracted from audio samples to address speech defects

Provided herein are systems and methods for providing instructions for speech based on speech classifications of verbal communications from users. Ae computing system can generate speech characteristics for a first verbal communication using a first audio sample from a user. The computing system can determine, from a plurality of speech classifications, a first speech classification for the first verbal communication based on the speech characteristics. The computing system can select, from a plurality of actions, an action comprising modifying one or more of the speech characteristics to define an utterance for the user based on the first speech classification. The computing system can provide an instruction presenting a message to prompt the user to perform the utterance defined by the action selected from the plurality of actions. The efficacy of the medication that the user is taking to address the condition may be increased.
Owner:CLICK THERAPEUTICS INC

Low-resource task-oriented semantic parsing via intrinsic modeling for assistant systems

ActiveUS12443797B1Semantic analysisSpeech recognitionLanguage understandingStructural representation
In one embodiment, a method includes receiving training utterances associated with a domain, receiving ontology labels for the domain, wherein the ontology labels comprise one or more of an intent or a slot, generating an inventory for the domain, wherein the inventory comprises at least a respective index and respective span for each intent or slot, wherein the respective span comprises a respective descriptive label associated with the intent or slot, and wherein the respective descriptive label comprises a natural-language description of the intent or slot, generating frames for training utterances based on the training utterances and the inventory by a natural-language understanding (NLU) model, wherein each frame comprises a structural representation of the respective training utterance, wherein the structural representation is generated based on a comparison between the corresponding training utterance and the inventory, and updating the NLU model based on the frames.
Owner:META PLATFORMS INC

Data augmentation for intent classification

The present disclosure relates to a data augmentation system and method that uses a large pre-trained encoder language model to generate new, useful intent samples from existing intent samples without fine-tuning. In certain embodiments, for a given class (intent), a limited number of sample utterances of a seed intent classification dataset may be concatenated and provided as input to the encoder language model, which may generate new sample utterances for the given class (intent). Additionally, when the augmented dataset is used to fine-tune an encoder language model of an intent classifier, this technique improves the performance of the intent classifier.
Owner:SERVICENOW INC

System and method for scalable generation of synthetic data for semantic parsers

A system and method for scalable generations of synthetic <logical form, utterance> pairs for training a semantic parser is disclosed. A semantic parser is trained on pairs of <logical form, utterance>. An ontology graph is constructed and derived from a plurality of enterprise documents and provides relationships among the concepts or classes of an organization. One or more paths are traversed among a plurality of source and destination node pairs, facilitating comprehensive semantic representation. Attributed query subgraphs are generated of source nodes, destination nodes, and hidden nodes. Each path is recognized among a variety of possible paths between source and destination nodes. Each path in the ontology query subgraph is validated by considering a plurality of predicates and a knowledge graph generates a natural language utterance. The utterances are refined and rephrased using a large language model, enhancing their coherence and linguistic quality.
Owner:OPENSTREAM INC

Real-time natural language processing and fulfillment

A system and method of real-time feedback confirmation to solicit a virtual assistant response from an evolving semantic state of at least a portion of an utterance. A user accesses a virtual assistant on an electronic device having the system and / or method configured to capture a command, a question, and / or a fulfillment request from audio such as, the speech emitted from the speaking user. The speech may be intercepted by a speech engine configured to transcribe the speech into text that is matched with the fragment pattern's regular expression to generate a fragment and / or the speech may be processed with a machine learning model to identify fragments. The fragments are identified by a domain handler configured to update a data structure of the current semantic state of the utterance in real-time on an interface of an electronic device.
Owner:SOUNDHOUND AI IP LLC

Multi-level emotional enhancement of dialogue

A system for emotionally enhancing dialogue includes a computing platform having processing hardware and a system memory storing a software code including a predictive model. The processing hardware is configured to execute the software code to receive dialogue data identifying an utterance for use by a digital character in a conversation, analyze, using the dialogue data, an emotionality of the utterance at multiple structural levels of the utterance, and supplement the utterance with one or more emotional attributions, using the predictive model and the emotionality of the utterance at the multiple structural levels, to provide one or more candidate emotionally enhanced utterance(s). The processing hardware further executes the software code to perform an audio validation of the candidate emotionally enhanced utterance(s) to provide a validated emotionally enhanced utterance, and output an emotionally attributed dialogue data providing the validated emotionally enhanced utterance for use by the digital character in the conversation.
Owner:DISNEY ENTERPRISES INC

Conversation dialogue orchestration in virtual assistant communication sessions

Methods and apparatuses for conversation dialogue orchestration in virtual assistant communication sessions include a server that establishes a chat session between a virtual assistant (VA) application and a client device. The VA application captures an utterance generated by a user and processes the utterance to instantiate a dialogue behavior tree comprising workflow agents each associated with executable code for completing a corresponding workflow action. The VA application traverses the behavior tree to generate a response to the utterance, including evaluating one or more conditions associated with a workflow agent to determine whether to execute the code in the workflow agent, and when the conditions associated with the workflow agent are met, executing the code to complete the workflow action and storing a sub-response in a dialogue memory. The VA application coalesces the sub-responses to generate a final response and transmits the final response to the client device.
Owner:FMR CORP

Next best agent selection in an adaptive workflow

An embodiment extracts, from an utterance, a goal. An embodiment prompts a large language model (LLM) to select, using metadata describing a plurality of agents, a set of candidate agents from the plurality of agents, each candidate agent in the set of candidate agents corresponding to the goal. An embodiment scores, using metadata of the set of candidate agents, each candidate agent in the set of candidate agents, the scoring resulting in a set of scored candidate agents. An embodiment prompts the LLM to select, using a set of business policy constraints, a next agent from the set of scored candidate agents. An embodiment invokes the next agent, the invoking causing the next agent to perform an action furthering the goal.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Techniques for efficient encoding in neural semantic parsing systems

Techniques for natural language processing include accessing an input string comprising a natural language utterance and a database schema representation for a database; providing the natural language utterance to a first encoder to generate one or more embeddings of the natural language utterance; providing the database schema representation to the first encoder to generate one or more embeddings of the database schema representation; encoding, by a second encoder, relations between elements in the database schema representation and words in the natural language utterance based on the one or more embeddings of the natural language utterance and the one or more embeddings of the database schema representation; and generating a logical form for the natural language utterance based on the encoded relations, the one or more embeddings of the natural language utterance, and the one or more embeddings of the database schema representation.
Owner:ORACLE INT CORP

Wearable device and operation method thereof, wearable intelligent system and storage medium

The invention relates to the technical field of artificial intelligence, particularly provides wearable equipment and an operation method thereof, a wearable intelligent system and a storage medium, and aims to solve the problem of how to timely and accurately obtain sentences which are not talked out by a user due to jamming and prompt the sentences. The method provided by the invention comprises the steps of collecting an electroencephalogram signal of a user and a target voice signal in an environment where the user is located, wherein the target voice signal comprises a first voice signal of the user; analyzing the speaking state of the user according to the first voice signal; when the speaking state is a pause state, according to the electroencephalogram signal and the target voice signal, predicting a missing utterance of the user during pause, and outputting prompt information; or, the electroencephalogram signal and the target voice signal are sent to the server, prompt information sent by the server is received and output, and the prompt information is generated by the server according to the utterance missed by the user during the jamming. By means of the method, the sentences which are not taught by the user due to jamming can be timely and accurately obtained and prompted.
Owner:BEIJING BOE TECH DEV CO LTD +1

Modifying Software Functionality based on Implicit Input

An implementation may involve: receiving audio input that contains utterances; determining, by a speech-to-text engine that receives the audio input, a textual representation of the utterances; providing, to a natural language model, a request to determine an intent of the textual representation of the utterances, wherein the request indicates that the intent is to be selected from a plurality of predefined intents; receiving, from the natural language model, the intent; determining, based on the intent, an action; and based on the action, modifying operation of a software application.
Owner:GAMES GLOBAL OPERATIONS LTD

Transforming natural language to structured query language based on multi- task learning and joint training

Techniques are disclosed for training a model, using multi-task learning, to transform natural language to a logical form. In one particular aspect, a method includes accessing a first set of utterances that have non-follow-up utterances and a second set of utterances that have initial utterances and associated one or more follow-up utterances and training a model for translating an utterance to a logical form. The training is a joint training process that includes calculating a first loss for a first semantic parsing task based on one or more non-follow-up utterances from the first set of utterances, calculating a second loss for a second semantic parsing task based on one or more initial utterances and associated one or more follow-up utterances from the second set of utterances, combining the first and second losses to obtain a final loss, and updating model parameters of the model based on the final loss.
Owner:ORACLE INT CORP

Reference resolution during natural language processing

Systems and processes for operating a digital assistant are provided. An example method includes, at an electronic device having one or more processors and memory, detecting invocation of a digital assistant; determining, using a reference resolution service, a set of possible entities; receiving a user utterance including an ambiguous reference; determining based on the user utterance and the list of possible entities, a candidate interpretation including a preliminary set of entities corresponding to the ambiguous reference; determining, with the reference resolution service and based on the candidate interpretation including the preliminary set of entities corresponding to the ambiguous reference, an entity corresponding to the ambiguous reference; and performing, based on the candidate interpretation and the entity corresponding to the ambiguous reference, a task associated with the user utterance.
Owner:APPLE INC