Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30 results about "Voice engine" patented technology

A voice engine is a software subsystem for bidirectional audio communication, typically used as part of a telecommunications system to simulate a telephone. It functions like a data pump for audio data, specifically voice data. The voice engine is typically used in an embedded system.

Comprehensive AI-enabled systems for immersive voice, companion, and augmented / virtual reality interaction solutions

A computer-implemented method for operating an artificial intelligence voice agent system includes receiving voice input through communication channels; analyzing converted text through natural language processing (NLP) pipelines implementing intent recognition and sentiment analysis detecting emotional cues using a multimodal large language model (LLM); generating response content using machine learning models trained on domain-specific corpora; converting generated responses to synthetic speech through text-to-speech (TTS) engines; integrating with a customer relationship management (CRM) platforms or an enterprise resource planning (ERP) database; and implementing continuous learning by updating language understanding models using conversation logs, voice recognition parameters based on user feedback, and response generation patterns. One implementation is a computer-implemented system and method that operates a suite of intelligent interactive devices and platforms including an artificial intelligence voice agent, enhanced communication platforms, an intimacy companion system, and augmented / virtual reality eyeglasses. Further, one implementation includes AR / VR eyeglasses that project visual content onto interchangeable lenses or directly onto the user's retina via laser-based retinal projection, provide prescription adjustments, incorporate ear-mounted sensors for monitoring physiological parameters like heart rate, oxygen saturation, and blood pressure, and utilize wireless data transmission, onboard environmental sensing, and remote calibration, all designed to offer dynamically adaptive, secure, and context-aware interactions across communication, personal assistance, health monitoring, and immersive augmented or virtual reality environments.
Owner:TRAN BAO

Voice interaction processing method, device and system, intelligent door lock and cloud server

The invention is suitable for the technical field of intelligent door locks, and provides a voice interaction processing method, device and system, an intelligent door lock and a cloud server, and the method comprises the steps: receiving a voice instruction sent by a user, and carrying out the local recognition, so as to obtain a local recognition result and the confidence of the local recognition result; when the local recognition result is matched with the local instruction set and the confidence is high, executing a local operation corresponding to the local recognition result; otherwise, establishing a secure communication session with the cloud server based on voiceprint biological characteristics extracted from the voice instruction; sending the voice instruction to a cloud server through the secure communication session, and receiving a cloud operation instruction and / or a dynamic response text returned by the cloud server; and executing corresponding operation according to the cloud operation instruction, and / or converting the dynamic response text into voice for broadcasting through a text-to-voice engine. The intelligent door lock voice interaction system solves the problem that the real-time performance, the reliability, the safety and the interaction intellectualization are difficult to balance in the existing intelligent door lock voice interaction technology.
Owner:SHANGHAI ZHENGZHI INTELLIGENT TECH CO LTD

Voice-AI warning system for predicted events

Embodiments include a computing system, computer-implemented method and non-transitory computer readable medium for predicting events and voice-AI warnings. According to embodiments, data corresponding to a predicted event is received, and users that are predicted to be affected by the predicted event are identified. A voice-AI engine is initiated to perform a voice-AI call to the identified users, where the voice-AI call provides a warning to each of the users.
Owner:ASSURED INSURANCE TECH INC

Voice call real-time transcription system and method

The invention provides a voice call real-time transcription system and method, and relates to the technical field of computers, and the system comprises a network element module which is used for obtaining corresponding audio data when a call request of a user side is detected; the voice streaming engine is used for carrying out hierarchical compression on the audio data based on a preset perceptual weighted vector quantization algorithm to obtain audio compressed data, and carrying out format conversion processing on the audio compressed data to obtain temporary audio data; the voice engine is used for performing feature extraction on the temporary audio data to obtain multi-modal feature data, and processing the multi-modal feature data based on a preset voice recognition model to obtain text information; and the analysis and optimization module is used for obtaining corresponding real-time transliteration text data according to the text information and a preset vocabulary library based on a preset large model. According to the method, the voice information is comprehensively represented by using the multi-modal feature data, so that the voice recognition model can more accurately convert the voice into the text.
Owner:CHINA UNICOM WO MUSIC & CULTURE CO LTD

A method and system for virtual standardized patient avatar generation and dialogue

The application discloses a kind of virtual standardization patient image generation and dialogue method and system, comprising: using large language model to extract the multidimensional information of patient from original medical data and construct structured case data, infer the visual features of patient, and integrate visual features as portrait description prompt word;Using text encoder, portrait description prompt word is converted into high-dimensional semantic vector, guide latent diffusion model to generate patient portrait image in line with medical data;Convert user voice signal into natural language text as role prompt word;According to the structured case data of patient and preset emotional factor, system prompt word is constructed;Using large language model and based on role and system prompt word, the reply text of user voice signal is generated;Start voice engine to convert reply text into reply voice, and output reply voice through loudspeaker.The application can generate virtual patient entity with visual, audible and intelligent interaction capability according to static medical text.
Owner:THE SECOND XIANGYA HOSPITAL OF CENT SOUTH UNIV

Manufacturing industry ERP system and method based on dynamic module loading and anti-noise voice control

The invention relates to an industrial intelligent ERP (Enterprise Resource Planning) production management system and a key technology, belongs to the technical field of industrial informatization software, artificial intelligence voice interaction and equipment predictive maintenance, and aims at solving the problems that a traditional industrial ERP system is insufficient in flexibility, workshop voice interaction is difficult, and equipment maintenance is passive and low in efficiency. The core innovation of the invention lies in that: a module hot plug architecture is adopted, so that a user can switch a full-function ERP system / a lightweight chemical single system / a purchase-sale-stock combined system within 5 seconds; according to an anti-noise voice engine, the instruction recognition rate gt in the 90dB workshop environment is determined; 92%, response delay lt; 200 ms; a closed loop is predictively maintained, an equipment health model is called in real time in work order execution, and shutdown loss is reduced by 40%. The problems that a traditional ERP system cannot be flexibly split, the workshop voice interaction rate is low, and equipment maintenance is passively responded are solved.
Owner:GUIZHOU QINGJUN TECHNOLOGY CO LTD

Text-to-speech from media content item snippets

A text-to-speech engine creates audio output that includes synthesized speech and one or more media content item snippets. The input text is obtained and partitioned into text sets. A track having lyrics that match a part of one of the text sets is identified. The location of the track's audio that contains the lyric is extracted based on forced alignment data. The extracted audio is combined with synthesized speech corresponding to the remainder of the input text to form audio output.
Owner:SPOTIFY

Video-generation system with structured data-based video generation feature

In one aspect, an example method includes (i) obtaining, by a computing system, structured data; (ii) generating, by the computing system using a natural language generator, a textual description of the structured data; (iii) transforming, by the computing system using a text-to-speech engine, the textual description of the structured data into synthesized speech; and (iv) generating, by the computing system using the synthesized speech, a synthetic video comprising the synthesized speech.
Owner:ROKU INC

Peer to peer conversation captioning system

A peer-to-peer conversation captioning (“PPCC”) system includes a first electronic device for a first user, a second electronic device for a second user, and a PPCC server. The PPCC system further includes a virtual relay entity, which includes a first virtual router for the first language, a second virtual router for the second language, and a global switch. The virtual routers include a speech-to-text engine, a translation engine, a text-to-speech engine, and a multiplexer. The speech-to-text engine converts a user's spoken language to texts and sends the texts to the global switch. Then, the translation engine receives and translates the texts where the text-to-speech engine converts the translated texts to voice. The multiplexer combines the translated texts and the voice to be sent to another user.
Owner:MEZMO CORP

Efficient human-to-machine and machine-to-human voice transmission

A human-to-machine transmission system and a machine-to-human transmission system includes a non-linear encoder that converts data representing an utterance into a frequency spectrum representing an aural range. An encoder in the human-to-machine transmission system that converts the output of the non-linear decoder into a bitstream and a decoder converts the bitstream into the frequency spectrum. An automatic speech recognition engine of the human-to-machine transmission system processes an output of the decoder to deliver an output. The machine-to-human transmission system includes a language model that replies to inquiries and a neural network that converts the text-to-speech received from a text-to-speech engine. It also includes a decoder that converts an input into the frequency spectrum, a vocoder that converts the frequency spectrum into audio frames, and a loudspeaker that converts the audio frame into audible sound.
Owner:SYNTHETIC MEDIA PROCESSING LABORATORY PTE LTD

Natural language understanding based meta-voice system using assistant system improves voice recognition accuracy

In one embodiment, a method includes receiving, from a client system associated with a first user, a first audio input. The method includes generating, based on a plurality of automatic speech recognition (ASR) engines, a plurality of transcriptions corresponding to the first audio input. Each ASR engine is associated with a respective domain of a plurality of domains. The method includes determining, for each transcription, a combination of one or more intents and one or more slots associated with the transcription. The method includes selecting, by a meta speech engine, one or more combinations of intents and slots associated with the first user input from the plurality of combinations. The method includes generating a response to the first audio input based on the selected combination and sending, to the client system, instructions for presenting the response to the first audio input.
Owner:CTRL-LABS CORP

Real-time natural language processing and fulfillment

A real-time feedback confirmation system and method for requesting a virtual assistant response from the evolving semantic state of at least a portion of an utterance. A user accesses a virtual assistant on an electronic device having a system and / or method configured to capture commands, questions, and / or requests from audio, such as speech, uttered by a speaking user. The speech may be captured by a speech engine configured to transcribe the speech into text that is matched with regular expressions of fragment patterns to generate fragments, and / or the speech may be processed by a machine learning model to identify fragments. The fragments are identified by a domain handler configured to update a data structure of the current semantic state of the utterance in real time on the interface of the electronic device.
Owner:SOUNDHOUND AI IP LLC

Multi-mode fusion interaction system of self-service terminal

The invention discloses a multi-mode fusion interaction system of a self-service terminal, and the system comprises a voice interaction module which is used for recognizing a voice instruction based on an AI voice engine and converting the voice instruction into a first operation command; the gesture recognition module is used for capturing a hand action by using a 3D camera and converting the hand action into a second operation instruction; the touch screen interaction module is used for processing touch input on terminal touch screen equipment as a third operation instruction; the fusion processing unit is used for fusing, analyzing and sorting the first operation command, the second operation command and the third operation command to determine a final operation command; the output module is used for executing corresponding operation according to the final operation instruction; according to the method, multiple interaction modes can be integrated, the interaction flexibility is greatly improved, the user can freely switch among different interaction modes according to own habits, the current environment and the physical condition, and the convenience and ability of special groups to use various devices are improved.
Owner:INSPUR FINANCIAL INFORMATION TECHNOLOGY CO LTD

Language-independent dictionary-trained grapheme-to-phoneme converter and text-to-speech engine for improved speech recognition

A language-independent dictionary-trained grapheme-to-phoneme converter and text-to-speech engine for improved speech recognition is disclosed. Techniques are described for recognizing a spoken wakeup word (WW) or command using a speech recognition system that does not need to be trained with any speech data that matches the WW / command for a human machine interface. Techniques train a word segmentation device and include decomposing a word into a plurality of combinations by splitting the word from a database at a plurality of different points for each of the plurality of combinations of unique sub-words. The word includes one or more writing units and one or more corresponding acoustic units. Techniques map acoustic units constituting a word to writing units constituting the word to generate an acoustic unit-to-writing unit mapping. The technique includes assigning a subset of the acoustic units to each of the unique sub-words based on the acoustic unit-to-write unit map to generate an acoustic unit-to-sub-word assignment of the word. Techniques accumulate acoustic units to sub-word assignments for a plurality of words from a database to create a sub-word likelihood dictionary.
Owner:INFINEON TECHNOLOGIES AMERICAS CORP

AI-based power supply and energy efficiency service interpretation system and method

The application discloses a power supply and energy efficiency service interpretation system based on AI artificial intelligence, a background management module of which selects corresponding voice broadcast services and energy use suggestion services according to user demand information; when the background management module selects the voice broadcast services, a voice broadcast service module encodes power supply and energy efficiency service voice messages to be executed in a queue to obtain corresponding audio stream data, and a text-to-speech engine performs voice synthesis output; when the background management module selects the energy use suggestion services, an energy use suggestion module queries corresponding cases in a power supply and energy efficiency service case database through an energy use recommendation label, adds a name value corresponding to case information to a content header sent to a browser according to an http-equiv attribute of the case information, and then jumps to a specified address. The application realizes the intelligentization of customer service and improves customer service experience.
Owner:STATE GRID NINGXIA ELECTRIC POWER CO LTD MARKETING SERVICE CENT STATE GRID NINGXIA ELECTRIC POWER CO LTD METERING CENT +1

System

A system is provided.SOLUTION: A system comprising: means for receiving a voice input of a user; means for converting the voice input into digital data; means for transmitting the digital data to a server; means including a speech recognition engine for converting the digital data into text data; means including a text-to-speech engine for analyzing the text data, understanding a user's intention, and acquiring information from an external service based on the user's intention; means for generating a response message in a natural language based on the acquired information; means for converting the response message into voice data; and means for transmitting the voice data to a terminal.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Universal voice command distribution method, electronic device, and computer-readable storage medium

The application belongs to the technical field of voice instructions, and specifically discloses a general voice instruction distribution method, an electronic device and a storage medium. The method of the application realizes the API by accessing a voice engine, converts returned semantic callback instructions into unified entity classes, defines two parameter variables, inputs the unified parameters into a semantic executor interface, instantiates the semantic executor interface through dynamic proxy, and the semantic executor interface obtains the callback interface modified by annotation, the semantic callback instruction method and the unified parameters of the semantic callback instruction method through reflection. The parameter variables in the unified parameters are matched with the callback interface modified by annotation and the parameters of the annotation, and it is judged whether to execute. If it is judged to be callback, the callback semantic instruction is executed. Otherwise, it is prompted that the command is not supported. When the voice engine is replaced, the new engine API can be flexibly accessed, the voice instructions of different voice engines are distributed and processed, and the research and development cost is reduced.
Owner:BEIJING MINGTUO HENGXIN TECH DEV CO LTD

Ai-voice call escalation and empathy system

PendingUS20250372076A1Speech recognitionSpeech synthesisEngineeringHuman agent
Disclosed are systems and methods for enabling collaborative AI-human voice interactions in an outbound call environment. An AI voice engine synthesizes real-time speech responses, optionally using a voice model representing a specific human agent. A sentiment analysis engine monitors user speech to classify emotional state, and an emotion modulation engine dynamically adjusts tone, pitch, and prosody of AI output based on inferred sentiment or operator commands. An operator console allows live human oversight, including editing AI-generated text, selecting emotional tone presets, or overriding responses entirely. The system supports real-time disclosure of AI identity and records call metadata for compliance. Training modules log operator interventions, sentiment patterns, and interaction outcomes to improve future AI behavior, enabling a scalable voice platform that balances automation with human empathy.
Owner:CELLIGENCE INTERNATIONAL LLC

Multiple modality human performance improvement systems

A system for improving human performance includes a software application with an onboarding engine, an initialization engine, an optimization engine, and a voice engine to populate a user profile, and use the profile to select and optimize a training protocol. The application is executed on a computer with processor, memory, user interface, and storage. The application exchanges information with the user and accesses a file containing the user profile. A method for improving human performance includes onboarding a new user by dynamically generating questions to create a profile; selecting an initial training protocol for the user based on the profile and verified training pathways; modifying the initial protocol based on user feedback to create an optimized training protocol, and guiding the user's training session with voice instructions that are generated by a voice engine.
Owner:MINDRIGHT HLDG INC

A voice call real-time transcription system and method

The application provides a voice call real-time transcription system and method, and relates to the technical field of computers.The system comprises a network element module for acquiring corresponding audio data when detecting a call request of a user terminal; a voice stream sending engine for performing hierarchical compression on the audio data based on a preset perceptual weighting vector quantization algorithm to obtain audio compression data, and performing format conversion processing on the audio compression data to obtain temporary audio data; a voice engine for performing feature extraction on the temporary audio data to obtain multimodal feature data, and processing the multimodal feature data based on a preset voice recognition model to obtain text information; and an analysis optimization module for obtaining corresponding real-time transcription text data based on a preset large model and according to the text information and a preset vocabulary library.The application comprehensively represents voice information by utilizing multimodal feature data, so that the voice recognition model can more accurately convert voice into text.
Owner:CHINA UNICOM WO MUSIC & CULTURE CO LTD

A method and system for establishing and evaluating an intelligent speech engine capability sharing model

The application discloses a kind of intelligent voice engine capability sharing model establishment and evaluation method and system include, obtain the audio data generated by dispatcher in the dispatching process when power system operation, and the data is preprocessed;The data after pre-processing is divided into first data set and second data set, first data set is used to establish voice engine sharing model, and second data set is used to verify voice engine sharing model;According to the verification result, the performance of voice engine sharing model is graded evaluation.The audio data of dispatching command platform is converted into structured index information, through data intelligent analysis, help accident analysis and disposal, combined with the actual scene of dispatching production intelligent learning and optimization are carried out, improve the accuracy of intelligent voice analysis, realize the scientific evaluation of dispatching language specification, improve the quality of dispatching command.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD +1

Video-Generation System with Structured Data-Based Video Generation Feature

In one aspect, an example method includes (i) obtaining, by a computing system, structured data; (ii) generating, by the computing system using a natural language generator, a textual description of the structured data; (iii) transforming, by the computing system using a text-to-speech engine, the textual description of the structured data into synthesized speech; and (iv) generating, by the computing system using the synthesized speech, a synthetic video comprising the synthesized speech.
Owner:ROKU INC

Autonomous mobile interaction and multi-mode control display system based on edge AI

The invention discloses an autonomous mobile interaction and multi-mode control display system based on edge AI. The system comprises an all-terrain adaptive mobile module used for driving a display device to move on different types of terrains; the three-dimensional space interaction module is used for identifying the posture of a human body lying station and detecting to form environmental map information; the edge intelligent voice module is internally provided with an end-side large model platform or a mixed architecture voice engine and a special acceleration chip to realize an edge AI function and provide offline voice instruction recognition and continuous dialogue; the data sovereignty protection module carries out localization processing and encrypted storage on the user data; the multi-mode cooperative control module is used for fusing multi-source sensing data and executing a cooperative decision; the end-side large model module is used for providing localized AI reasoning capability; the ecological interconnection module provides a cross-platform protocol conversion gateway and an open API interface so as to establish application circulation and seamless projection with different systems of external equipment. According to the invention, autonomous movement, intelligent interaction, privacy protection and cross-platform seamless cooperation of equipment are realized.
Owner:TPV ELECTRONICS (FUJIAN) CO LTD

Real-time system for spoken natural stylistic conversations with large language models

The techniques disclosed herein enable systems for spoken natural stylistic conversations with large language models. In contrast to many existing modalities for interacting with large language models that are limited to text, the techniques presented herein enable users to carry a fully spoken conversation with a large language model. This is accomplished by converting a user speech audio input to text and utilizing a prompt engine to analyze a sentiment expressed by the user. A large language model, having been trained on example conversations, by generating a text response as well as a style cue to express emotion in response to the sentiment expressed by speech audio input. A text-to-speech engine can subsequently interpret the text response and style cue to generate an audio output which emulates the sensation of human conversation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Virtual standardized patient image generation and dialogue method and system

The invention discloses a virtual standardized patient image generation and dialogue method and system, and the method comprises the steps: extracting multi-dimensional information of a patient from original medical data through a large language model, constructing structured case data, reasoning the visual features of the patient, and integrating the visual features into portrait description cues; using a text encoder to convert the portrait description cue word into a high-dimensional semantic vector, and guiding the potential diffusion model to generate a patient portrait image conforming to the medical data; converting the user voice signal into a natural language text as a role prompt word; according to the structured case data of the patient and a preset emotion factor, constructing a system prompt word; generating a reply text of the user voice signal by using a large language model based on the role and the system prompt word; and starting a voice engine to convert the reply text into reply voice, and outputting the reply voice through a loudspeaker. According to the method, the visual and audible virtual patient entity with the intelligent interaction capability can be generated according to the static medical text.
Owner:THE SECOND XIANGYA HOSPITAL OF CENT SOUTH UNIV

Exploder with voice engine, method, computer equipment and storage medium

PendingCN120656441ABlasting cartridgesAlarmsEngineeringVoice engine
The invention relates to the technical field of exploder interaction, in particular to an exploder with a voice engine, and the exploder comprises a TTS voice engine which is integrated in an exploder control module and is used for converting text information into voice signals in real time; the resource scheduling module is in control connection with the TTS voice engine and is used for providing a running environment for an external application program and controlling communication of the TTS voice engine; the voice interface protocol library is connected with the TTS voice engine and is used for defining a standardized interface for calling the TTS voice engine by an application program; and the real-time voice feedback channel is connected with the TTS voice engine and the resource scheduling system and is used for ensuring the instantaneity of user operation and system state voice broadcast through a priority queue mechanism. According to the application, the TTS is applied to the exploder and is very friendly to interaction, the TTS serves as a text-to-voice technology, operation items and operation results can be broadcasted through operation, the importance of visual interaction in the using process of the detonation software is reduced, and the method is compatible with more users of different ages and different culture levels.
Owner:RONGGUI SICHUANG BEIJING TECH

Real-time natural language processing and fulfillment

A system and method of real-time feedback confirmation to solicit a virtual assistant response from an evolving semantic state of at least a portion of an utterance. A user accesses a virtual assistant on an electronic device having the system and / or method configured to capture a command, a question, and / or a fulfillment request from audio such as, the speech emitted from the speaking user. The speech may be intercepted by a speech engine configured to transcribe the speech into text that is matched with the fragment pattern's regular expression to generate a fragment and / or the speech may be processed with a machine learning model to identify fragments. The fragments are identified by a domain handler configured to update a data structure of the current semantic state of the utterance in real-time on an interface of an electronic device.
Owner:SOUNDHOUND AI IP LLC

Real-time natural language processing and implementation

A system and method for real-time feedback acknowledgement for requesting a virtual assistant response from an evolving semantic state of at least a portion of an utterance. A user accesses a virtual assistant on an electronic device having systems and / or methods configured to capture commands, questions, and / or fulfillment requests from audio (e.g., speech made by a speaking user). The speech may be intercepted by a speech engine configured to transcribe the speech into text that matches a regular expression of a schema of segments to generate segments, and / or the speech may be processed with a machine learning model to identify segments. The segment is identified by a domain handler configured to update a data structure of a current semantic state of the utterance in real-time on an interface of the electronic device.
Owner:SOUNDHOUND AI IP LLC

Voice interaction method and device, and storage medium

The invention discloses a voice interaction method and device and a storage medium, and relates to the technical field of vehicles, and the voice interaction method comprises the steps: collecting vehicle scene features under the condition that voice information input by a user is obtained, the vehicle scene features at least comprise one of network features, voice information features, vehicle state features and user personalized features; based on a preset multi-factor decision model, determining a target routing decision according to the vehicle scene features; according to the target routing decision, calling a local voice engine and / or a cloud voice engine to process the voice information input by the user, and generating a target voice instruction; and executing the target voice instruction to realize voice interaction between the user and the vehicle. According to the scheme, the adaptive capability of the voice interaction function in different interaction scenes is realized, and the user experience of voice interaction with the vehicle is improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Real-time system for spoken natural stylistic conversations with large language models

The techniques disclosed herein enable systems for spoken natural stylistic conversations with large language models. In contrast to many existing modalities for interacting with large language models that are limited to text, the techniques presented herein enable users to carry a fully spoken conversation with a large language model. This is accomplished by converting a user speech audio input to text and utilizing a prompt engine to analyze a sentiment expressed by the user. A large language model, having been trained on example conversations, by generating a text response as well as a style cue to express emotion in response to the sentiment expressed by speech audio input. A text-to-speech engine can subsequently interpret the text response and style cue to generate an audio output which emulates the sensation of human conversation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC