Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

157 results about "Utterance" patented technology

In spoken language analysis, an utterance is the smallest unit of speech. It is a continuous piece of speech beginning and ending with a clear pause. In the case of oral languages, it is generally, but not always, bounded by silence. Utterances do not exist in written language, however, only their representations do. They can be represented and delineated in written language in many ways.

Signal generation system, computing device, and signal generation method

A signal generation system includes a first computing device configured to generate a trajectory signal representing a planned trajectory of a vehicle; and a second computing device configured to generate an output word signal representing word information to be given to an occupant of the vehicle, by inputting an input word signal representing an utterance of the occupant in natural language format into a trained language model. At least one of the first computing device and the second computing device converts at least one of the sensor signal and a signal generated based on the sensor signal to a trajectory data signal represented in a format that can be inputted into the trained language model. The second computing device generates a trajectory word signal related to the planned trajectory as the output word signal, based on the trajectory data signal.
Owner:TOYOTA JIDOSHA KK

Spoken language understanding method based on multi-view expert fusion and interest word selection

This invention discloses a spoken language understanding method based on multi-view expert fusion and interest lexical selection, belonging to the field of natural language understanding and semantic parsing technology. The method includes: extracting hidden state sequences from the input utterance using a shared encoder; constructing a multi-view expert fusion module containing utterance view experts, block view experts, and lexical view experts to generate multi-granularity expert features; generating intent aggregation features and slot aggregation features through a task decoupling gating mechanism; performing interest lexical selection on the intent aggregation features, calculating lexical importance scores, generating weighted intent representations, and predicting multi-intent labels; and inputting the slot aggregation features into a decoder with diagonal mask constraints to generate a position-aligned slot label sequence. This invention enhances the model's dynamic focusing ability on key intent signals and can be applied to intelligent dialogue systems, virtual assistants, and vertical domain semantic parsing scenarios.
Owner:JIANGNAN UNIV

Information processing device, information processing method, and information processing program

The present invention provides an information processing device that enables a more natural improvement in the accuracy of identifying past related utterances made by a user or identifying their relationship with other users. [Solution] The system includes: a speech acquisition unit 1 that acquires user utterances; a prompt generation unit 2 that generates prompts for a large-scale language model to generate tag information that includes at least one of the following based on the acquired utterances: the relationships between the constituent elements of the utterance, the attributes to which the utterance belongs, or a summary of the utterance; an interface 3 that transmits the generated prompts to the large-scale language model and receives response information that includes tag information generated by the large-scale language model in response to the transmitted prompts; and a recording unit 5 that records the acquired utterances and the received tag information in association.
Owner:PIONEER IP

Techniques for transforming natural language conversation into a visualization representation

Techniques are disclosed herein for transforming natural language conversations into a visual output. In one aspect, a computer-implement method includes generating an input string by concatenating a natural language utterance with a schema representation comprising a set of entities for visualization actions, generating, by a first encoder of a machine learning model, one or more embeddings of the input string, encoding, by a second encoder of the machine learning model, relations between elements in the schema representation and words in the natural language utterance based on the one or more embeddings, generating, by a grammar-based decoder of the machine learning model and based on the encoded relations and the one or more embeddings, an intermediate logical form that represents at least the query, the one or more visualization actions, or the combination thereof, and generating, based on the intermediate logical form, a command for a computing system.
Owner:ORACLE INT CORP

System and method for contextual analysis and metadata database generation for user-specific speech patterns

A system for contextual analysis and metadata database generation for user-specific speech patterns is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system splits the speech signal into a first set of audio frames, where each audio frame comprises an utterance of one or more words. The system determines a context associated with each word. In response, the system detects a context change between a first text and a second text. The system generates a contextually split set of frames by splitting the speech signal into a second set of audio frames according to the detected context changes.
Owner:BANK OF AMERICA CORP

Information processing equipment, communication support systems, programs

Evaluating content based on video or audio data acquired during communication. [Solution] The present invention relates to an information processing device 20 for evaluating content displayed on a user's terminal device in online communication, wherein the audio data recorded in the communication includes the user's utterances in response to the explanation of the content, and the video data recorded in the communication includes the video of the user receiving the explanation of the content, and comprises an acquisition unit 23 for acquiring at least one of the audio data or the video data, and an evaluation unit 27 for analyzing at least one of the user's utterances included in the audio data or the user's facial expressions included in the video data acquired by the acquisition unit to determine an evaluation score for the content.
Owner:RICOH CO LTD

Hotword suppression

PendingUS20260188318A1Audio watermarkSpeech sound
A method includes adding, by a first computing device, a first audio watermark to first speech data corresponding to playback of a first utterance including a hotword used to invoke an attention of a second computing device. The method includes outputting, by the first computing device, the playback of the first utterance corresponding to the watermarked first speech data. The second computing device is configured to receive the watermarked first speech data and determine to cease processing of the watermarked first speech data.
Owner:GOOGLE LLC

Information processing device and information processing method

A user is assisted in performing a voice operation appropriately.A situation determination section determines a situation. A state control section controls a voice command appropriate for the determined situation to put the voice command into a receivable state. For example, the user is informed of what the voice command in a receivable state is, by means of display or voice output. The user can utter a voice command without performing a user action to prevent false recognition, such as the utterance of a wake word. This reduces the troublesomeness and burden of the user.
Owner:SATURN LICENSING LLC

Domain knowledge guided fine-tuning of neural conversational models

Examples of the present disclosure describe systems and methods that leverage domain knowledge to influence selection of candidate action templates in a neural conversational model. More specifically, natural language rules can be provided to a natural language rule reasoner to bias selection of candidate action templates. In some instances, the natural language rules can include user input and system actions. In other instances, the natural language rules can include a previous system action and a next system action. A bias vector can then influence selection of a candidate action template from a set of candidate action templates to determine the most relevant candidate action template based on the natural language rules, the candidate action templates, and a user utterance or other system input.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Streaming speech-to-speech model with automatic speaker turn detection

A method for turn detection in a speech-to-speech model includes receiving, as input to the speech-to-speech (S2S) model, a sequence of acoustic frames corresponding to an utterance. The method further includes, at each of a plurality of output steps, generating, by an audio encoder of the S2S model, a higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames, and determining, by a turn detector of the S2S model, based on the higher order feature representation generated by the audio encoder at the corresponding output step, whether the utterance is at a breakpoint at the corresponding output step. When the turn detector determines that the utterance is at the breakpoint, the method includes synthesizing a sequence of output audio frames output by a speech decoder of the S2S model into a time-domain audio waveform of synthesized speech representing the utterance spoken by the user.
Owner:GOOGLE LLC

The word of god (WOG): the 1,197,000 letter string of encoded hebrew letters underlying the original bible

PendingUS20260141175A1Natural language translationSemantic analysisAlgorithmPattern detection
A data structure and associated methods for analysis of a continuous 1,197,000-letter unvocalized Hebrew string referred to as the Word of God (WOG). The data structure contains only the twenty-two classical Hebrew letters and their five final forms, with no spacing, punctuation, vowelization, or editorial symbols. Intrinsic placement of the final letters enables deterministic segmentation of the string into 305,490 lexical units and 23,206 verses without external conventions. Fixed letter-number assignments provide a numeric architecture for evaluating substrings, detecting alterations, identifying encoded mathematical correspondences, and performing pattern analysis. The system preserves full semantic range by supporting multiple morphologically valid interpretations of unvocalized Hebrew strings. Methods for segmentation, numeric evaluation, reconstruction, integrity verification, semantic analysis, and mathematical pattern detection are provided thereby providing a reproducible foundation for computational and linguistic research.
Owner:JURAVIN DON KARL

Electronic device, method, and non-transitory computer-readable storage medium for generating animation of character

PCT designated stageWO2026141709A1AnimationProcessing
This electronic device comprises: a memory storing instructions and including one or more storage media; and at least one processor including processing circuitry. When the instructions are individually or collectively executed by the at least one processor, the instructions may cause the electronic device to: acquire a speech signal for utterance of a character and text information about the speech signal; identify a word having a designated meaning among words included in the text information by providing the text information to an artificial intelligence model; acquire data, which indicates a time point when the word is to be uttered, on the basis of identifying the time point by using the speech signal; generate an animation representing the face of the character uttering the speech signal by providing the speech signal and the data to a trained model; and control the speech signal to be output while displaying the animation.
Owner:NCSOFT CORP

Voice interaction services

The disclosed embodiments include computerized methods, systems, and devices, including computer programs encoded on a computer storage medium, for integrating voice-based interaction and control into a native graphical user interface (GUI) of an executed application. For example, a communications device may receive audio data corresponding to an utterance spoken by a user, and may obtain structured data representative of the received audio data. The communications device may provide structured data to the executed application through a programmatic interface, and the executed application may perform the one or more operations in accordance with the structured data. The communications device may generate data indicative of an output of the one or more operations performed by the executed application, and may present at least a portion of the generated output data to a user through a corresponding interface.
Owner:GOOGLE LLC

Application execution method and wearable device supporting same

A wearable device disclosed in the present document may comprise a display, a memory, and a processor. The memory may store instructions that cause the wearable device to: execute a speech recognition application for processing a speech input of a user; receive a first utterance input of the user through the speech recognition application; determine a first virtual object executed in an application different from the speech recognition application and corresponding to the first utterance input; determine a position on the display at which the first virtual object is to be displayed or a shape of the first virtual object on the basis of at least one of information on an external device around the wearable device, information on an object recognized in images that are being output through the display or have a history of being output through the display, and information related to the user; and display the first virtual object on the display according to the position or the shape.
Owner:SAMSUNG ELECTRONICS CO LTD

Adaptive speech recognition system with user feedback analysis and dynamic retriggering

A computing system providing an enhanced user experience with a voice assistant (VA) that may be part of an automotive or vehicle infotainment system. A speech recognition module 10 converts spoken in
Owner:MERCEDES BENZ GROUP AG

Electronic device, method, and non-transitory computer-readable storage medium for determining uplink transmission power

This electronic device may comprise at least one processor. The at least one processor is configured to: while an application for providing an interpretation service for a call, which is performed by using a communication circuit, is executed, transmit uplink voice data of a first language generated on the basis of a user's utterance; determine a transmission time interval for transmitting uplink voice data of a second language on the basis of generating the uplink voice data of the second language from the utterance of the first language; determine transmission power for transmitting the uplink voice data of the second language on the basis of the transmission time interval for transmitting the uplink voice data of the second language; and transmit the uplink voice data of the second language on the basis of the transmission power.
Owner:SAMSUNG ELECTRONICS CO LTD

Selective amplification of voice and interactive language simulator

ActiveUS12670639B2EngineeringSpeech sound
Systems and processes for operating a digital assistant are provided. An example method includes, at an electronic device having one or more processors and memory, receiving an audio input including an utterance, determining, based on a speaker profile, an identity of a speaker of the utterance, determining whether the identity of the speaker matches a predetermined identity, and in accordance with a determination that the identity of the speaker matches the predetermined identity selectively adjusting a volume of the utterance relative to a volume of other sound of the audio input and providing an output of the adjusted utterance.
Owner:APPLE INC

Electronic device for vehicle, and method of operating same

A method of operating an electronic device for a vehicle is disclosed. The method comprises the steps of: obtaining, on the basis of a user's utterance, first features corresponding to a plurality of items of utterance analysis information; obtaining second features corresponding to a plurality of items of context analysis information indicating a context inside the vehicle on the basis of one or more images capturing the inside of the vehicle; obtaining, on the basis of sensor data from one or more sensors in the vehicle, third features corresponding to a plurality of items of vehicle state information indicating a vehicle state related to driving; mapping at least one of the first features, the second features, or the third features with model analysis data to select a model to respond to the utterance from among a plurality of models; and processing the utterance through the selected model to provide a response.
Owner:SAMSUNG ELECTRONICS CO LTD

Systems, methods, and storage media for performing actions based on utterance of a command

ActiveUS12670905B1Speech soundAudio frequency
Systems and methods for recognizing and executing spoken commands using speech recognition. Exemplary implementations may: store actionable phrases; obtain audio information representing sound captured by a mobile client computing platform associated with a user; detect any spoken instances of a predetermined keyword present in the sound represented by the audio information; perform speech recognition on the sound represented by the audio information; identify an utterance of an individual actionable phrase in speech temporally adjacent to the spoken instance of the predetermined keyword that is present in the sound represented by the audio information; perform natural language processing to identify an individual command uttered temporally adjacent to the spoken instance of the predetermined keyword that is present in the sound represented by the audio information; and effectuate performance of instructions corresponding to the command.
Owner:SUKI AI INC

Context-based device arbitration

ActiveUS12670907B2EngineeringDevice status
This disclosure describes, in part, context-based device arbitration techniques to select a voice-enabled device from multiple voice-enabled devices to provide a response to a command included in a speech utterance of a user. In some examples, the context-driven arbitration techniques may include determining a ranked list of voice-enabled devices that are ranked based on audio signal metric values for audio signals generated by each voice-enabled device, and iteratively moving through the list to determine, based on device states of the voice-enabled devices, whether one of the voice-enabled devices can perform an action responsive to the command. If the voice-enabled devices that detected the speech utterance are unable to perform the action responsive to the command, all other voice-enabled devices associated with an account may be analyzed to determine whether one of the other voice-enabled devices can perform the action responsive to the command in the speech utterance.
Owner:AMAZON TECH INC

Electronic device, method, and storage medium for providing information of object

PCT designated stageWO2026141890A1Computer hardwareEngineering
An electronic device, a method, and a storage medium for providing information of an object are provided. The electronic device comprises: a microphone; at least one camera; an output interface; at least one processor including a processing circuit; and a memory storing instructions. The instructions, when executed individually or collectively by the at least one processor, instruct the electronic device to acquire an image including at least one object through the at least one camera. The instructions instruct the electronic device to identify a user's utterance received through the microphone by using an artificial intelligence model. The instructions instruct the electronic device to determine whether multi-modal input is required according to whether a target object is specified by the user's utterance. The instructions instruct the electronic device to determine the target object included in the image on the basis of the user's utterance when multi-modal input is unnecessary. The instructions instruct the electronic device to obtain a predetermined area including the target object. The instructions instruct the electronic device to obtain information of the target object included in the predetermined area. The instructions instruct the electronic device to provide information of the target object through the output interface.
Owner:SAMSUNG ELECTRONICS CO LTD

Automated assistant adapted for multiple age groups and / or vocabulary levels

This application relates to automated assistants that accommodate multiple age groups and / or vocabulary levels. Described herein are techniques for enabling an automated assistant to adjust its behavior depending on a detected age range and / or "vocabulary level" of a user that is interfacing with the automated assistant. Data indicative of utterances of the user can be used to estimate one or more of an age range and / or vocabulary level of the user. The estimated age range / vocabulary level can be used to influence aspects of a data processing pipeline employed by the automated assistant. Aspects of the data processing pipeline that can be influenced by the age range / vocabulary level of the user can include one or more of automated assistant invocation, speech-to-text processing, intent matching, intent resolution, natural language generation, and / or text-to-speech processing. In some implementations, one or more tolerance thresholds associated with one or more of these aspects such as grammatical tolerance, vocabulary tolerance, etc. can be adjusted.
Owner:GOOGLE LLC

Reducing response time in conversational ai systems and applications

This disclosure relates to reducing response times in conversational AI systems and applications. In various examples, response times (e.g., latency) associated with conversational artificial intelligence (AI) systems can be shortened by sharing partial results between modules or components of the conversational AI system. For example, rather than waiting for an ASR system to convert a user utterance into textual data, a system described in this disclosure can obtain a candidate prefix for the utterance from the ASR system and use a language model to predict the complete utterance based on the candidate prefix and generate a response to the predicted utterance. As more information is obtained (e.g., the remaining portion of the utterance), the language model can update the predicted utterance and / or the response. Additionally, in some cases, a system of this disclosure can begin forwarding a response to a text-to-speech (TTS) system before the language model completes generating the response.
Owner:NVIDIA CORP

Generative suggestion chips for enhanced chatbot conversations

Implementations disclosed herein relate to dynamically generating suggestion chip(s) during a conversation between a chatbot, that is engaged in the conversation on behalf of a user, and an additional user. During the conversation, processor(s) of a system can: receive an initial portion of a spoken utterance of the additional user; generate, based on processing at least the initial portion of the spoken utterance, initial suggestion chip(s) that are each associated with a corresponding initial suggestion to respond to the spoken utterance; receive a subsequent portion of the spoken utterance; and determine, based on processing the subsequent portion of the spoken utterance, whether to cause the initial suggestion chip(s) to be rendered at a client device of the user. If rendered, the initial suggestion chip(s) can guide the chatbot's response. Otherwise, the system can generate subsequent suggestion chip(s) that are based on the entirety of the spoken utterance.
Owner:GOOGLE LLC

Speech recognition method and speech recognition apparatus

A speech recognition method of the disclosure includes a speech data accumulation step of acquiring speech data and accumulating the acquired speech data as accumulated speech data, an utterance determination step of determining whether there is an utterance for the acquired speech data or not, a decision step of deciding whether to end the accumulation of the speech data or continue the accumulation of the speech data based on a result of the determination as to whether there is an utterance or not that is determined in the utterance determination step, and an accumulation time of the accumulated speech data, and a speech recognition step of performing speech recognition based on the accumulated speech data in a case where the accumulation of the speech data has ended and the accumulated speech data exists.
Owner:SHARP KK

Multimodal human-machine interface for a vehicle

The invention relates to a method for operating a human-machine interface of a vehicle, wherein a speech converter of a human-machine interface of a vehicle converts a spoken utterance of an occupant of the vehicle, the spoken utterance being detected by a microphone of the human-machine interface, into a written utterance; an LLM module of the human-machine interface determines, on the basis of the written utterance, a written response to the detected spoken utterance; the voice converter converts the determined written response into a spoken response; and a loudspeaker of the human-machine interface outputs the spoken response. The invention also relates to a human-machine interface for a vehicle and to a vehicle.
Owner:AUDI AG

Arbitration between automated assistant devices based on interaction cues

Techniques are described herein for arbitration between automated assistant devices based on interaction cues. A method includes: receiving, via one or more microphones of a first computing device, first audio data that captures a spoken utterance of a user; determining that each of one or more additional computing devices has detected the spoken utterance of the user; determining that hotword arbitration is to be initiated between the first computing device and the one or more additional computing devices; for each of the first computing device and the one or more additional computing devices, identifying a similarity score for the computing device; selecting a target computing device, from the first computing device and the one or more additional computing devices, based on the similarity scores; and causing the target computing device to respond to a query that is included in the spoken utterance of the user.
Owner:GOOGLE LLC