Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

935 results about "Utterance" patented technology

In spoken language analysis, an utterance is the smallest unit of speech. It is a continuous piece of speech beginning and ending with a clear pause. In the case of oral languages, it is generally, but not always, bounded by silence. Utterances do not exist in written language, however, only their representations do. They can be represented and delineated in written language in many ways.

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Deepfake detection

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Owner:PINDROP SECURITY INC

System and method for contextual analysis and metadata database generation for user-specific speech patterns

A system for contextual analysis and metadata database generation for user-specific speech patterns is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system splits the speech signal into a first set of audio frames, where each audio frame comprises an utterance of one or more words. The system determines a context associated with each word. In response, the system detects a context change between a first text and a second text. The system generates a contextually split set of frames by splitting the speech signal into a second set of audio frames according to the detected context changes.
Owner:BANK OF AMERICA CORP

Next best agent selection in an adaptive workflow

An embodiment extracts, from an utterance, a goal. An embodiment prompts a large language model (LLM) to select, using metadata describing a plurality of agents, a set of candidate agents from the plurality of agents, each candidate agent in the set of candidate agents corresponding to the goal. An embodiment scores, using metadata of the set of candidate agents, each candidate agent in the set of candidate agents, the scoring resulting in a set of scored candidate agents. An embodiment prompts the LLM to select, using a set of business policy constraints, a next agent from the set of scored candidate agents. An embodiment invokes the next agent, the invoking causing the next agent to perform an action furthering the goal.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Information processing system, information processing method and program

An information processing system, method, and program for conducting interviews with applicants is provided. [Solution] The method includes starting a conference session in an interview preparation mode, and switching the conference session to the interview mode when an instruction to switch to the interview mode is received from a user, starting an interview with an avatar selected in the interview preparation mode as the interviewer, displaying an avatar in the interview preparation mode, outputting a first utterance for breaking the ice as the avatar's speech, outputting a second utterance for setting the avatar as the avatar's speech, and displaying selectable options for multiple setting items of the avatar, and accepting an avatar selection including a first response to the first utterance and a second response or selection of an option to the second utterance, and a switching instruction, a first display control step of displaying an avatar changed in accordance with the avatar selection, and a second display control step of displaying the changed avatar.
Owner:BIZREACH INC

Smart dispatcher in a composite artificial intelligence (AI) system

Certain aspects of the disclosure provide methods and systems for implementing a composite artificial intelligence system. The method may include generating a user request from a user utterance submitted by a user. The method may also include classifying a user intent from the user request and a context of the user utterance. The method may furthermore include determining to send the user request to one of a first AI model or a second AI model based on a determination that the user intent is fullfillable by one of the first AI model or the second AI model. The method may in addition include generating a first response by one of the first AI model or the second AI model based on the determination. Method may moreover include transmitting the first response to the user.
Owner:INTUIT INC

Virtual Meeting Coaching

In one embodiment, a system receives a set of coaching items including a number of questions each associated with an expected answer; connects to a coaching session including one or more participants and a virtual coaching agent; for each question and for at least a subset of the participants: transmitting the question, by the virtual coaching agent, to the client device used by the participant; receiving an answer to the question by the participant, the answer including media of the participant; receiving text of utterances spoken by the participant during the answer; generating one or more evaluation scores for the answer based on evaluating at least the content of the answer to the question; and transmitting an overall evaluation score for each of the subset of participants based on the generated evaluation scores for that participant.
Owner:ZOOM COMMUNICATIONS INC

Performance optimization with reflection tokens in a self-reflective retrieval-augmented generation framework

A method includes obtaining data source tokens that identify corresponding data sources, and adding the data source tokens to a vocabulary list of a response tokenizer. An internal mapping of the response tokenizer is updated with the data source tokens, mapping each data source token to a corresponding token identifier (ID). Vector representations corresponding to the data source tokens are added to a response embedding model. The response tokenizer and the response embedding model are included in a response foundation model. An annotated dataset including multiple instances of an input utterances, retrieved data, and a data source token is created. The retrieved data is retrieved from a data source corresponding to the data source token. The response foundation model is trained with the annotated dataset to output data source tokens for a new input utterance based on a context of the new input utterance.
Owner:INTUIT INC

Passive and continuous multi-speaker voice biometrics

Embodiments described herein provide for a voice biometrics system execute machine-learning architectures capable of passive, active, continuous, or static operations, or a combination thereof. Systems passively and / or continuously, in some cases in addition to actively and / or statically, enrolling speakers as the speakers speak into or around an edge device (e.g., car, television, radio, phone). The system identifies users on the fly without requiring a new speaker to mirror prompted utterances for reconfiguring operations. The system manages speaker profiles as speakers provide utterances to the system. Machine-learning architectures implement a passive and continuous voice biometrics system, possibly without knowledge of speaker identities. The system creates identities in an unsupervised manner, sometimes passively enrolling and recognizing known or unknown speakers. The system offers personalization and security across a wide range of applications, including media content for over-the-top services and IoT devices (e.g., personal assistants, vehicles), and call centers.
Owner:PINDROP SECURITY INC

Conversation content generation method and apparatus, and storage medium and terminal

A conversation content generation method and apparatus, a storage medium and a terminal are provided. The method includes: acquiring a current utterance entered by a user; reading a preset topic transfer graph and target topic, wherein the topic transfer graph includes nodes and connecting lines between the nodes, the nodes correspond to topics in one-to-one correspondence, each connecting line points from a first node to a second node, a weight of the connecting line indicates probability of transferring from a topic corresponding to the first node to a topic corresponding to the second node, and the topic transfer graph includes a node corresponding to the target topic; determining a topic of reply content of the current utterance at least based on the current utterance, the topic transfer graph and the target topic, and recording it as a reply topic; generating the reply content at least based on the reply topic.
Owner:UNIDT (SHANGHAI) CO LTD

Wake-Word Processing in an Electronic Device

Wake-word processing by a wearable electronic device could be carried out when the device is worn by a user and is in a device sleep state, the device including a linear microphone array having at least two microphones vertically spaced from each other, and the device also including a processor. And the example method could involve (i) the at least two microphones of the linear microphone array receiving an audio waveform representing a wake-word utterance, (ii) the processor making a determination, based at least on an angle of arrival of the audio waveform at the at least two microphones of the linear microphone array and / or an energy level of the audio waveform received at the at least two microphones of the linear array, of whether to accept the wake-word utterance or rather to reject the wake-word utterance, and (iii) the processor controlling operation of the device based on the determination.
Owner:STRYKER CORP

Large language models for nl2SQL with long context finetuning

The present disclosure relates to manufacturing training and testing data by leveraging data augmentation techniques to generate examples of long context database schemas. Aspects are directed towards accessing a training dataset comprising training examples where each training example may include i) a prompt including a natural language utterance and a database schema having one or more tables, and ii) a gold logical form corresponding to the natural language utterance, combining the tables from the database schemas in the training examples may generate a combined database schema set, generating a set of long context training examples based on the training dataset and the combined database schema set, and incorporating the long context database schema into the selected training example to generate a long context training example to train a generative artificial intelligence model with at least the set of long context training examples to generate a trained generative artificial intelligence model.
Owner:ORACLE INT CORP

Neural network based conversation-aware automatic speech recognition

A system uses a machine learning based model such as a neural network for transcribing audio inputs. The system receives a set of audio inputs representing utterances of a conversation. For each conversation, the system determines a dialogue state for each utterance. The system uses a hierarchical language model for transcribing audio inputs of an online conversation using the received conversations. The hierarchical language model includes a top-level language model and a plurality of lower-level language model. The training is performed by (1) training the top-level language model using sequences of corresponding dialogue state, each sequence of dialogue states for a conversation, and (2) for each dialogue state, training a lower-level language model using utterances having that dialogue state. The system executes the hierarchical language model to transcribe audio input of new conversations.
Owner:INTERACTIONS LLC (US)

Voice assistance system and method for holding a conversation with a person

A voice assistance system for holding a spoken conversation with a person. The system can include at least one microphone configured for detecting a voice utterance of the person, at least one speaker configured for outputting a sound to the person, at least one processor configured for executing computer instructions, and at least one memory. The at least one memory stores computer instructions configured for operating the system to perform steps including: providing at least one machine learning (ML) model configured for generating contextually relevant and varied responses in natural language conversations, detecting a voice utterance using the microphone, providing the voice utterance as an input to the ML model, prompting the ML model to generate an output based on the input, and providing the output to the speaker to be output to the person.
Owner:FRIENDLYBUZZ CO PBC

Injecting short-term spectro-temporal knowledge into automatic speech recognition models

Mechanisms are provided for training an Automatic Speech Recognition (ASR) model. An automatic speech recognition (ASR) computer model is trained based on full utterance training data, the ASR computer model having an audio encoder, text predictor, and a joint network which combines outputs of both the audio encoder and the text predictor. Fine-tuning training of the ASR computer model is performed by a knowledge distillation framework at least by: executing a chunking operation on full utterance data to generate a plurality of data chunks corresponding to full utterances in the full utterance data; and executing a knowledge distillation operation with two encoder embeddings. A first encoder embedding is obtained from the full utterance data and a second encoder embedding is obtained from the data chunks. Operational parameters of the trained ASR model are updated based on a loss determined from the first encoder embedding and second encoder embedding.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Using optimal articulatory event-types for computer analysis of speech

A method includes obtaining, by a processor, a score quantifying an estimated degree to which an instance of an articulatory event-type indicates a state, with respect to a disease, in which the instance was produced. The method further includes, using the score, facilitating, by the processor, a computer-implemented procedure for evaluating the state of a subject based on a test utterance produced by the subject. Other embodiments are also described.
Owner:CORDIO MEDICAL LTD

Lightweight member marketing method and device for real-time conversation, terminal and storage medium

The invention discloses a lightweight member marketing method and device for real-time conversation, a terminal and a storage medium, and the method comprises the steps: responding to a received user voice stream, continuously generating an incomplete intermediate text stream through streaming voice recognition, and inputting the intermediate text stream into a lightweight marketing response model; the lightweight marketing response model predicts the potential intention of the user in real time based on the intermediate text stream, and pre-assembles one or more candidate response frameworks from the fragmented marketing corpus; when it is detected that the utterance of the user ends, the real intention of the user is determined based on a complete final text stream generated by streaming speech recognition; and arbitrating the candidate reply framework according to the weighted decision factor set to obtain an arbitration result, determining final reply content according to the arbitration result, converting the final reply content into an instant response call, and outputting the instant response call. According to the method, the time delay of the real-time call marketing process can be reduced, and the concurrent processing capability is improved.
Owner:SUZHOU XIYUAN DIGITAL TECH CO LTD

Output interpretation for a meaning representation language system

The present disclosure is related to techniques for converting a natural language utterance to a logical form query and deriving a natural language interpretation of the logical form query. The techniques include accessing a Meaning Resource Language (MRL) query and converting the MRL query into a MRL structure including logical form statements. The converting includes extracting operations and associated attributes from the MRL query and generating the logical form statements from the operations and associated attributes. The techniques further include translating each of the logical form statements into a natural language expression based on a grammar data structure that includes a set of rules for translating logical form statements into corresponding natural language expressions, combining the natural language expressions into a single natural language expression, and providing the single natural language expression as an interpretation of the natural language utterance.
Owner:ORACLE INT CORP

Method and apparatus to provide comprehensive smart assistant services

An apparatus supports smart assistant services with a plurality of smart service providers. The apparatus includes an audio device that receives a speech signal having a user utterance, captures the user utterance when the user utterance includes a user wake word, and sends the captured utterance to a backend computing device. The backend computing device replaces the user wake word with specific wake words associated with different smart service providers. The processed utterances are then sent to selected smart service providers. The backend computing device subsequently constructs feedback to the user utterance based on voice responses from the different smart service providers. The backend computing device then passes a digital representation of the feedback to the audio device, and the audio device converts the digital representation to an audio reply to the user utterance.
Owner:COMPUTIME LTD

Method and system for training a virtual agent using optimal utterances

A method and system for training a virtual agent is provided herein. The method and system comprises storing conversations between the virtual agent and a user in logs. The method and system further comprises mining the logs to retrieve utterances. The method and system further comprises computing regression for each of the plurality of the charging time segments. The method and system further comprises providing a score to the utterances. Further, the method ranking the utterances based on the score.
Owner:QUANTIPHI INC

Analyzing underspecified natural language utterances in a data visualization user interface

A computing device parses a user-specified natural language command to form a first expression. The computing device determines that the first expression is ambiguous or underspecified. The computing device, in accordance with the determination, infers first information using one or more inferencing rules, where at least one of the inferencing rules is based on an attribute of data fields and / or data values in a data source. The computing device forms a second expression based on the first expression using and the first information. The computing device retrieves one or more data sets from the data source using according to the second expression. The computing device generates and displays a data visualization of the retrieved one or more data sets.
Owner:TABLEAU SOFTWARE INC

Language-agnostic multilingual modeling using effective script normalization

A method includes obtaining a plurality of training data sets each associated with a respective native language and includes a plurality of respective training data samples. For each respective training data sample of each training data set in the respective native language, the method includes transliterating the corresponding transcription in the respective native script into corresponding transliterated text representing the respective native language of the corresponding audio in a target script and associating the corresponding transliterated text in the target script with the corresponding audio in the respective native language to generate a respective normalized training data sample. The method also includes training, using the normalized training data samples, a multilingual end-to-end speech recognition model to predict speech recognition results in the target script for corresponding speech utterances spoken in any of the different native languages associated with the plurality of training data sets.
Owner:GOOGLE LLC

Application Programming Interfaces For On-Device Speech Services

A method (500) includes receiving, from an application (50) executing on a client device (110), at a speech service interface (200), configuration parameters (211) for integrating a speech service (250) into the application. The configuration parameters include a language pack directory (225) that maps a primary language code (235) to an on-device path of a primary language pack (110) of the speech service for use in recognizing speech in a primary language and each of one or more codeswitch language codes to an on-device path. The method also includes receiving audio data (102) characterizing an utterance (106) and processing, using a language ID predictor model (230), the audio data to determine that the audio data is associated with the primary language code. The method also includes processing, using the primary language pack, the audio data to determine a transcription (120) that includes one or more words in the primary language.
Owner:GOOGLE LLC

Reliable and interpretable drift detection in streams of short texts

Various systems and methods are presented regarding detecting data drift. The data of interest can be batches of utterances received at an interface (e.g., a chatbot). The batches of utterances can be compared with topics present in training data utilized to train a data classifier (e.g., an autoencoder), wherein topics identified in the batches of utterances that are not present in the training data can be considered to be novel topics. The greater the presence of novel topics in a batch of utterances, the greater the divergence of the batch of utterances from the content of the training data. The novel topics can be identified and subsequently applied to the training data such that the data classifier can be re-trained with the novel topics, thereby causing the data classifier to be contemporaneous with the novel topics. In an embodiment, the utterances can be short streams of text, symbols, and suchlike.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Interactive ai toy capable of holding a conversation with a person, and method of interacting with same

An interactive AI toy capable of holding a spoken conversation with a person. The toy can include any of a microphone configured for detecting a voice utterance of the person, a speaker configured for outputting a sound to the person, at least one processor configured for executing computer instructions, and at least one memory. The at least one memory can store computer instructions configured for operating the toy to perform steps comprising providing at least one machine learning (ML) model configured for generating contextually relevant and varied responses in natural language conversations. The steps can also include detecting a voice utterance of the person using the microphone, providing the voice utterance as an input to the ML model, prompting the ML model to generate an output based on the input, and providing the output to the speaker to be output to the person.
Owner:FRIENDLYBUZZ CO PBC

Multi-round dialogue robot intention hit optimization method

The invention relates to the technical field of conversation type artificial intelligence, in particular to a multi-round conversation robot intention hit optimization method. The method comprises the following steps: after a user utterance is received, a system generates candidate intentions through a double-path mechanism; then, each candidate intention is evaluated through a multi-factor scoring algorithm, and the algorithm fuses four key signals: a graph transition probability, semantic correlation, a context coherence score calculated through a cross attention mechanism, and an out-of-domain penalty; and finally, making a decision according to the final scores of the candidate intentions: if the confidence coefficient of the candidate intention with the highest score is high enough and the candidate intention is obviously distinguished from other candidates, confirming the intention; otherwise, a targeted clarification problem is generated to solve ambiguity. According to the method, the accuracy and robustness of intention detection are remarkably improved through collaborative integration of dialogue streams, semantic matching and deep context analysis, and meanwhile, the user experience is improved through intelligent ambiguity processing.
Owner:MINIMALIST INTERNET (BEIJING) INFORMATION TECHNOLOGY CO LTD

Model robustness on operators and triggering keywords in natural language to a meaning representation language system

ActiveUS12475325B2Natural language translationSpeech analysisRepresentation languageData set
Techniques are disclosed herein for improving model robustness on operators and triggering keywords in natural language to a meaning representation language system. The techniques include augmenting an original set of training data for a target robustness bucket by leveraging a combination of two training data generation techniques: (1) modification of existing training examples and (2) synthetic template-based example generation. The resulting set of augmented data examples from the two training data generation techniques are appended to the original set of training data to generate an augmented training data set and the augmented training data set is used to train a machine learning model to generate logical forms for utterances.
Owner:ORACLE INT CORP

Localizing and verifying utterances by audio fingerprinting

Methods and systems are disclosed for enhancing the security of a user device such as a voice command device. A computing device associated with the user device may be configured to receive an indication of a trigger, such as a predetermined word or passcode. In response to receiving the indication of the trigger, the computing device may be configured to determine a verification signal marker and to cause transmission of the verification signal marker. The computing device may receive an audio input comprising a voice command and a detected signal marker and verify the voice command based on a comparison of the detected signal marker and the verification signal marker. In response to the verifying the voice command, the computing device may be configured to cause execution of an operation associated with the voice command such as tuning to a specific channel on a nearby set-top box.
Owner:COMCAST CABLE COMM LLC

Systems and methods for emotion-based call summarization

Embodiments of the present disclosure provide systems and methods for emotion-based call summarization. One method may include receiving an emotion prediction vector for an utterance text segment from a transcript data object, the emotion prediction vector comprising a plurality of emotion prediction scores respectively corresponding to a plurality of emotion identifiers; generating a domain-specific relevancy prediction for the utterance text segment based on a category-relevant subset of the plurality of emotion prediction scores that correspond to one or more category-specific emotion identifiers of the plurality of emotion identifiers associated with a domain-specific summarization category; identifying the utterance text segment as a relevant utterance from the transcript data object based on a comparison between the domain-specific relevancy prediction and a relevancy threshold; and initiating a performance of a machine learning summarization operation based on the utterance text segment.
Owner:OPTUM INC