Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

39 results about "Speech interface" patented technology

A speech interface is a software application that enables interaction between humans and voice enabled-applications, such as virtual assistants and voice assistants. Speech interfaces use and mimic human speech via speech recognition technology. But designing an effective speech interface requires more than writing a script for your voice assistant.

Voice Interface Integration System and Method for Mobile, Computer and Web Applications

A method for voice-based interaction between a user (e.g., a shopper) and online stores includes capturing an initial voice input from a user and converting the voice input to text. The text is analyzed and initial search terms are created such as product name, product description, product category, product brand name, product manufacturer and retailer. One or more searches are conducted using the initial search terms to identify one or more retailers meeting the search terms. One or more identified online retailers are accessed and searches are executed at the one or more identified retailers, generating search results. The identified online retailers have a user interface allowing purchases of the identified product or item. The search results are returned to the user.
Owner:OMNIBEK IP HLDG LLC

Account association with device

Systems and methods for account data association with voice interface devices are disclosed. For example, when a host user / primary user and guest user have consented for account data to be associated with the primary user's devices, a request to associate the account data may be received. Voice and device-based authentication may be performed to confirm the identity of the guest user and the guest user's account data may be associated with the primary user's devices. During a guest session, voice recognition may be utilized to determine if a given user utterance is from the guest user or the primary user, and actions may be performed by the voice interface device accordingly.
Owner:AMAZON TECH INC

Intelligent voice interface data calling method

The invention relates to the technical field of intelligent voice data processing, and discloses an intelligent voice interface data calling method. The method comprises the following steps: receiving a voice instruction stream input by a user, and converting the voice instruction stream into a structured semantic unit set through a multi-level semantic analysis engine; and dynamically layering the instruction intention according to the semantic association degree in the structured semantic unit set to generate an intention sequence with priority marks. And based on the priority mark of the intention sequence, extracting a candidate data block set matched with the intention of each level from a distributed data storage cluster. And performing cross-modal feature alignment on the candidate data block set, and calculating the semantic coverage degree of each data block and the corresponding intention hierarchy. And screening out target data blocks according to a semantic coverage threshold, and recombining the data blocks according to a priority sequence of the intention sequence. And loading the recombined target data block set to a real-time processing channel of the voice interface to generate an interactive voice response data stream.
Owner:NAT ENERGY CHANGYUAN HANCHUAN POWER GENERATION CO LTD

A low cost voice protection circuit

The utility model relates to interface EMC protection technical field, concretely is a low cost voice protection circuit, a low cost voice protection circuit, including speech interface RJ11, voice protection circuit, speech matching circuit and speech SLIC chip that electrically connect gradually, voice protection circuit includes the primary protection circuit and secondary protection circuit that set gradually along signal transmission direction, the input end of primary protection circuit connects speech interface RJ11's tip end and ring, and the input end of secondary protection circuit is connected to the output end of primary protection circuit. The present application sets up the hierarchical protection framework containing primary protection circuit and secondary protection circuit between speech interface and SLIC chip, and combines the hierarchical configuration of matching resistance group, and forms the low cost, high reliability speech signal protection system.
Owner:TAICANG T&W ELECTRONICS CO LTD

Use of Context to Disambiguate Automation-Configuration Command

A method and system for use of context information to disambiguate an automation-configuration command. In an example method, a computing system receives a voice command uttered by a user into a voice-interface device, the voice command describing an Internet-of-Things (IoT) automation. Further, in response to receiving the voice command, the computing system determines, based on context information not specified by the voice command, which of multiple IoT devices should be a subject of an IoT rule that implements the described IoT automation, and provisions the IoT rule with the determined IoT device as the subject of the IoT rule. In example implementations, the context information could be based on network signaling between devices, ambient audio in the user's environment, and / or one or more other factors.
Owner:ROKU INC

Visual and voice interface for a dialysis machine

ActiveUS12673144B2Home dialysisKidney machines
A dialysis machine includes a user interface for providing visual information and / or spoken information to a user. For example, in some implementations, the user interface may be configured to provide visual information related to an action, such as showing the action being partially or fully completed, and a speaker can provide spoken instructions to assist the user in machine set-up, calibration and / or operation. Such instructions can be particularly useful in a home dialysis setting. In some implementations, the speaker can provide spoken alarms that are related to alarm conditions. The spoken alarms may include patient and / or dialysis machine identifying information. The verbosity of the spoken instructions and / or the spoken alarms may be adjustable, and both may be accompanied by visual information displayed by the dialysis machine (e.g. visual alarms, images and / or video).
Owner:FRESENIUS MEDICAL CARE HOLDINGS INC

Dependency-inverted speech conversion methods, devices, systems, and media

This invention discloses a speech conversion method, apparatus, system, and medium based on dependency inversion. The method includes: receiving a user's speech conversion request; invoking a preset abstract speech interface according to the speech conversion request; obtaining the currently injected speech SDK in the abstract speech interface; processing the speech conversion request according to the function of the currently injected speech SDK; and returning the corresponding speech conversion result. By pre-setting an abstract speech interface on which both the application's main program and the underlying speech SDK modules depend, the required speech SDK is flexibly injected into the abstract speech interface based on dependency injection to fulfill the user's speech conversion request. This makes the application's main program no longer dependent on the underlying speech SDK, reducing the dependencies between components and eliminating the impact of changes in the underlying modules on the application's operation. This improves the flexibility of speech conversion functionality while also enhancing the stability of mobile application operation.
Owner:PING AN BANK CO LTD

Method and system for converting plant procedures to voice interfacing smart procedures

The present invention proposes a method and system that converts a procedure document into a smart procedure system capable of voice interfacing. The system of the present invention may include a conversion module that converts the BPMN process model of a procedure expressed in XML, automatically generated via NLP-based technology, into a smart procedure system capable of voice interfacing. The conversion module may be configured as a client-server architecture composed of a mobile application client on a smart device and a server system providing backend services.
Owner:INJE UNIVERSITY INDUSTRY ACADEMIC COOPERATION FOUNDATION

Systems and methods for generating a dynamic list of hint words for automated speech recognition

Systems and methods are provided for determining hint words that improve the accuracy of automated speech recognition (ASR) systems. Hint words are typically determined in the context of a user issuing voice commands in connection with a voice interface system, however, a voice interface system may capture terms from overheard content and / or conversations. A system may determine a sliding window of hint words using set of qualifier rules. The system may capture audio, e.g., from a conversation or played back content, as a first input and decipher a plurality of words including a qualifying first term added to the hint words. The voice interface system may capture more audio as a second input and decipher a second plurality of words including a qualifying second term. The first term may be removed from the set of hint words, e.g., when the second term is added or after an expiration time.
Owner:ADEIA GUIDES INC

An AI outbound calling system and its control method and medium

This invention relates to the field of voice interface control technology, specifically to an AI outbound calling system and its control method and medium, comprising modules for physical link construction, interaction state recognition, flow path control, and intervention and collaborative response. By smoothly connecting analog signal gaps through impedance matching and gain compensation, a synchronous and continuous uplink and downlink voice signal sequence is generated; audio energy jumps are identified to pinpoint call state transition periods; frequency distribution characteristics of AI and human mode switching are analyzed to extract abnormal frequency segments of routing switching; and response synchronization flags are located by combining human intervention trend positioning to ultimately determine the core interactive control result. This invention improves the time consistency and compatibility of audio transmission across hardware terminals, and achieves efficient collaborative response and closed-loop control of interaction states between AI and human modes by accurately identifying flow paths and intervention response locations.
Owner:SHENZHEN DAZHI SOFTWARE TECH CO LTD

Context driven device arbitration

This disclosure describes, in part, context-driven device arbitration techniques to select a speech interface device from multiple speech interface devices to provide a response to a command included in a speech utterance of a user. In some examples, the context-driven arbitration techniques may include executing multiple pipeline instances to analyze audio signals and device metadata received from each of the multiple speech interface devices which detected the speech utterance. A remote speech processing service may execute the multiple pipeline instances and analyze the audio signals and / or metadata, at various stages of the pipeline instances, to determine which speech interface device is to respond to the speech utterance.
Owner:AMAZON TECH INC

A method and system for testing low-noise power supplies using waveforms

This invention belongs to the field of communication noise detection and processing technology, and relates to a method and system for testing low-noise power supplies using waveforms. The method connects the power supply under test (DUT) and a voice gateway, and powers them on in real time. The voice gateway is connected to an external noise-shielded environment via its voice interface. Within this environment, the voice gateway is modulated to a periodic state of comfortable noise or a dial tone, and the sound is played. The playing sound is acquired and processed, and based on the processed sound and the sound signal waveform, the noise level of the DUT is detected in real time. This transforms uncontrollable human-induced noise into an electrical signal measurable by an oscilloscope, accurately determining whether the power supply meets requirements. In other words, by utilizing a clean environment, the unmeasurable power frequency noise on the telephone line is converted into a differential-mode signal measurable by an oscilloscope through a physical sound isolation process, providing testers with an objective experimental result.
Owner:GUANGZHOU V-SOLUTION TELECOMM TECH CO LTD

Multi-participant voice ordering

A voice interface recognizes spoken utterances from a plurality of users. The voice interface responds to these utterances by modifying item instance attributes, etc. The voice interface computes a voice vector for each utterance and associates it with the modified item instance. For subsequent utterances with highly matched speech vectors, the speech interface will modify the same instance; for subsequent utterances for which the speech vector does not match the speech vector stored for any item instance, the speech interface will modify different item instances.
Owner:SOUNDHOUND AI IP LLC

Integration of speech processing functionality with organization systems

Systems and methods for integration of speech processing functionality with organization systems are disclosed. For example, a voice interface application may be created to enable a voice interface functionality for devices associated with an organization. Space identifiers of spaces of the organization may be created and associated with the voice interface application. Devices associated with the space identifiers may be enabled for utilizing the voice interface application and may be set up utilizing wireless network identifiers associated with the spaces and / or the organization.
Owner:AMAZON TECH INC

Method for human speech processing

In a method for human speech processing in an automatic speech recognition (ASR) system, human speech is received at a speech interface of the ASR system, wherein the ASR system comprises embedded componentry for onboard processing of the human speech and cloud-based componentry for remote processing of the human speech. A keyword is identified at the speech interface within a first portion of the human speech. Responsive to identifying the keyword, a second portion of the human speech is analyzed to identify at least one command, the second portion following the first portion. The at least one command is identified within the second portion of the human speech. The at least one command is selectively processed within at least one of the embedded componentry and the cloud-based componentry.
Owner:TDK CORP

Voice deception attack detection system and method based on microphone array

The present invention belongs to the technical field of voice command liveness detection, and discloses a method for performing passive voice command liveness detection based on a circular microphone array of a smart speaker in a smart home environment to resist the threat of voice replay attacks from the device. In the process of liveness detection, through fine-grained analysis of the audio frequency domains of different channels and extraction of multi-channel features, the present invention can efficiently determine whether the voice command is generated by a real user or forged by an electronic device. The present invention can achieve fast and flexible defense against voice deception attacks, highly guarantee the security of voice interfaces in smart homes, and meet the needs of the industry. Relying only on the voice signal collected by the microphone of the smart speaker, it is possible to identify whether the voice command is generated by a real user or by a deceptive device.
Owner:SHANGHAI JIAOTONG UNIV

A soft, flexible and communicative device for patients with laryngectomy and its method thereof

PCT designated stageWO2025220035A1SensorsDiagnostic recording/measuringLaryngectomyAcoustics
The present invention relates to a soft, flexible, and communicative device for patients with laryngectomy and its method thereof The present invention discloses the fabrication of soft, lightweight, skin-compatible, and reusable device for silent speech interfaces (SSIs). It further comprises the unique feature to recognize speech-related information using surface electromyography (sEMG). sEMG-based speech recognition operates on signals recorded from a set of sEMG sensors that are strategically located on the face (with the help of desired form factor) and measure muscle activity associated with the phonation, resonation, and articulation of speech. sEMG sensors are printed onto strategical locations on soft, stretchable, and biocompatible membrane shaped into a form-factor of a face mask. The silent speech interfaces (SSIs) can also be deployed in acoustically challenging environment or where privacy / confidentially is a desirable such as in defense or military application and it is cost effective with reliable speech recognition accuracy of 96.2%(>95 %).
Owner:SHITASHII INNOVATIONS PTE LTD

Testing device and system

The utility model provides a testing device and system. The testing device comprises a testing host, a control panel, a telephone set and an off-hook and on-hook assembly, a first communication interface of the test host is in wired connection with a communication interface of the control panel, an input / output interface of the control panel is in wired connection with a control end of the off-hook and on-hook assembly, and an operation end of the off-hook and on-hook assembly is arranged opposite to an insertion spring of the telephone so as to press or release the insertion spring of the telephone under the driving of the off-hook and on-hook assembly; the telephone connection interface of the telephone is connected with the voice interface of the to-be-tested call gateway module through a telephone line, and the second communication interface of the test host is in wired connection with the communication interface of the to-be-tested call gateway module. Therefore, a tester does not need to carry out operations such as connection and hanging up, the test efficiency and the test precision are improved, and the test labor cost is reduced.
Owner:ANYSMART TECH CO LTD

Detecting accidental activation of speech interface

A replay detector executing in a vehicle prevents a speech interface from acting on audio inputs received from a microphone when those audio inputs are from a first class of audio inputs and permits the speech interface to acting on audio inputs from a second class of audio inputs. Audio inputs from the first class result from technically- produced acoustic signals that have been produced in an environment of the microphone. Audio inputs from the second class result from acoustic signals that have been produced by at least one person in that environment.
Owner:CERENCE OPERATING CO

Control device for controlling voice interface and thermal device equipped with such control device

To provide a technique related to a voice interface (I / F) control for operating a thermal appliance by voice.SOLUTION: A control device that controls a voice I / F includes an input I / F that receives an input signal from the outside, a computing unit that determines, when the input signal is a voice signal, whether or not the voice signal represents a prepared operation word, and an output I / F that outputs the control signal to the outside. The computing unit receives a voice input signal from the input I / F, and outputs, through the output I / F, a control signal requesting confirmation of execution of the voice command corresponding to a predetermined operation word when the computing unit determines that the signal represents a predetermined operation word. After transmission of the control signal, a second input signal is received via the input I / F, and when it is determined that the second input signal is an acknowledgment signal confirming the execution of the voice command, the computing unit outputs the control signal to instruct the execution of the voice command through the output I / F.SELECTED DRAWING: Figure 9
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Transitioning voice interactions

Techniques for processing a voice initiated request by a web server are presented. The techniques may include receiving, by a web server, request data representing a voice command to a user device, the request data including an identification of a requested webpage; determining, by the web server, that a response to the request data will continue a voice interaction; and providing, by the web server and to the user device, data for a voice enabled webpage associated with the requested webpage, where the data for the voice enabled webpage is configured to invoke a voice interface for the user device.
Owner:VERISIGN INC

Conversational rag pipeline for an LLM-based automotive assistant

An apparatus for interacting with an occupant in a vehicle includes an infotainment system that has been integrated into the vehicle, an automotive assistant that is configured to execute in the infotainment system, a speech interface that is configured to receive, from the occupant, an original utterance that is to be processed by the automotive assistant, and a classifier that determines that the original utterance has either anaphora or ellipsis. A model that provides a response to the occupant based on a prompt that includes the original utterance and that has been augmented by a function specification selected based at least in part on a rewritten utterance that has been derived from the original utterance.
Owner:CERENCE OPERATING CO

Memory degradation detection and improvement

Memory degradation detection and assessment includes capturing human utterances using a voice interface and generating a corpus of human utterances for a user, selecting human utterances from the corpus of human utterances based on a meaning of the human utterances determined by a computer processor through natural language processing. Contextual information for one or more of the corpus of human utterances is determined based on data generated in response to signals sensed by one or more sensing devices operatively coupled with the computer processor. Patterns in the corpus of human utterances are identified based on pattern recognition performed by the computer processor using one or more machine learning models. Changes in memory functioning of the user are identified based on the pattern recognition. The changes identified based on the contextual information are classified as to whether the changes are likely due to a memory impairment of the user.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Digitally assisted mobile assistance system with a rolling or flying carrier platform, speech AI, camera system and socio-functional support logic for interaction with people in private, assisted or medical environments.

Digitally assisted mobile assistance system with at least one rolling or flying mobile carrier platform, at least one computing unit, at least one voice interface, at least one audio unit, at least one camera unit and at least one artificial intelligence for interaction with at least one user.
Owner:WIRKUNGSDRIVE UG (HAFTUNGSBESCHRÄNKT)

Automotive assistant with hierarchy having a backbone and domain-specific delegees

An automotive assistant in an infotainment system of a vehicle includes a hierarchy that receives a top-level query via a speech interface and that provides a top-level response to the query. The hierarchy includes a top-level agent and a first and second domain with a related domain specific query. Both domains are queried using prompts comprising natural language. The top level response is formulated based on first and second domain specific activities.
Owner:CERENCE OPERATING CO

Speech interface device with caching component

A speech interface device is configured to receive response data from a remote speech processing system for responding to user speech. This response data may be enhanced with information such as a remote ASR result(s) and a remote NLU result(s). The response data from the remote speech processing system may include one or more cacheable status indicators associated with the NLU result(s) and / or remote directive data, which indicate whether the remote NLU result(s) and / or the remote directive data are individually cacheable. A caching component of the speech interface device allows for caching at least some of this cacheable remote speech processing information, and using the cached information locally on the speech interface device when responding to user speech in the future. This allows for responding to user speech, even when the speech interface device is unable to communicate with a remote speech processing system over a wide area network.
Owner:AMAZON TECH INC

Systems and methods for a text-to-speech interface

A computing system and related techniques for selecting content to be automatically converted to speech and provided as an audio signal are provided. A text-to-speech request associated with a first document can be received that includes data associated with a playback position of a selector associated with a text-to-speech interface overlaid on the first document. First content associated with the first document can be determined based at least in part on the playback position, the first content including content that is displayed in the user interface at the playback position. The first document can be analyzed to identify one or more structural features associated with the first content. Speech data can be generated based on the first content and the one or more structural features.
Owner:GOOGLE LLC