Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6 results about "Voice engine" patented technology

A voice engine is a software subsystem for bidirectional audio communication, typically used as part of a telecommunications system to simulate a telephone. It functions like a data pump for audio data, specifically voice data. The voice engine is typically used in an embedded system.

A method and system for virtual standardized patient avatar generation and dialogue

The application discloses a kind of virtual standardization patient image generation and dialogue method and system, comprising: using large language model to extract the multidimensional information of patient from original medical data and construct structured case data, infer the visual features of patient, and integrate visual features as portrait description prompt word;Using text encoder, portrait description prompt word is converted into high-dimensional semantic vector, guide latent diffusion model to generate patient portrait image in line with medical data;Convert user voice signal into natural language text as role prompt word;According to the structured case data of patient and preset emotional factor, system prompt word is constructed;Using large language model and based on role and system prompt word, the reply text of user voice signal is generated;Start voice engine to convert reply text into reply voice, and output reply voice through loudspeaker.The application can generate virtual patient entity with visual, audible and intelligent interaction capability according to static medical text.
Owner:THE SECOND XIANGYA HOSPITAL OF CENT SOUTH UNIV

Efficient human-to-machine and machine-to-human voice transmission

A human-to-machine transmission system and a machine-to-human transmission system includes a non-linear encoder that converts data representing an utterance into a frequency spectrum representing an aural range. An encoder in the human-to-machine transmission system that converts the output of the non-linear decoder into a bitstream and a decoder converts the bitstream into the frequency spectrum. An automatic speech recognition engine of the human-to-machine transmission system processes an output of the decoder to deliver an output. The machine-to-human transmission system includes a language model that replies to inquiries and a neural network that converts the text-to-speech received from a text-to-speech engine. It also includes a decoder that converts an input into the frequency spectrum, a vocoder that converts the frequency spectrum into audio frames, and a loudspeaker that converts the audio frame into audible sound.
Owner:SYNTHETIC MEDIA PROCESSING LABORATORY PTE LTD

Language-independent dictionary-trained grapheme-to-phoneme converter and text-to-speech engine for improved speech recognition

A language-independent dictionary-trained grapheme-to-phoneme converter and text-to-speech engine for improved speech recognition is disclosed. Techniques are described for recognizing a spoken wakeup word (WW) or command using a speech recognition system that does not need to be trained with any speech data that matches the WW / command for a human machine interface. Techniques train a word segmentation device and include decomposing a word into a plurality of combinations by splitting the word from a database at a plurality of different points for each of the plurality of combinations of unique sub-words. The word includes one or more writing units and one or more corresponding acoustic units. Techniques map acoustic units constituting a word to writing units constituting the word to generate an acoustic unit-to-writing unit mapping. The technique includes assigning a subset of the acoustic units to each of the unique sub-words based on the acoustic unit-to-write unit map to generate an acoustic unit-to-sub-word assignment of the word. Techniques accumulate acoustic units to sub-word assignments for a plurality of words from a database to create a sub-word likelihood dictionary.
Owner:INFINEON TECHNOLOGIES AMERICAS CORP

Universal voice command distribution method, electronic device, and computer-readable storage medium

The application belongs to the technical field of voice instructions, and specifically discloses a general voice instruction distribution method, an electronic device and a storage medium. The method of the application realizes the API by accessing a voice engine, converts returned semantic callback instructions into unified entity classes, defines two parameter variables, inputs the unified parameters into a semantic executor interface, instantiates the semantic executor interface through dynamic proxy, and the semantic executor interface obtains the callback interface modified by annotation, the semantic callback instruction method and the unified parameters of the semantic callback instruction method through reflection. The parameter variables in the unified parameters are matched with the callback interface modified by annotation and the parameters of the annotation, and it is judged whether to execute. If it is judged to be callback, the callback semantic instruction is executed. Otherwise, it is prompted that the command is not supported. When the voice engine is replaced, the new engine API can be flexibly accessed, the voice instructions of different voice engines are distributed and processed, and the research and development cost is reduced.
Owner:BEIJING MINGTUO HENGXIN TECH DEV CO LTD

Multiple modality human performance improvement systems

A system for improving human performance includes a software application with an onboarding engine, an initialization engine, an optimization engine, and a voice engine to populate a user profile, and use the profile to select and optimize a training protocol. The application is executed on a computer with processor, memory, user interface, and storage. The application exchanges information with the user and accesses a file containing the user profile. A method for improving human performance includes onboarding a new user by dynamically generating questions to create a profile; selecting an initial training protocol for the user based on the profile and verified training pathways; modifying the initial protocol based on user feedback to create an optimized training protocol, and guiding the user's training session with voice instructions that are generated by a voice engine.
Owner:MINDRIGHT HLDG INC

Voice interaction method and device, and storage medium

The invention discloses a voice interaction method and device and a storage medium, and relates to the technical field of vehicles, and the voice interaction method comprises the steps: collecting vehicle scene features under the condition that voice information input by a user is obtained, the vehicle scene features at least comprise one of network features, voice information features, vehicle state features and user personalized features; based on a preset multi-factor decision model, determining a target routing decision according to the vehicle scene features; according to the target routing decision, calling a local voice engine and / or a cloud voice engine to process the voice information input by the user, and generating a target voice instruction; and executing the target voice instruction to realize voice interaction between the user and the vehicle. According to the scheme, the adaptive capability of the voice interaction function in different interaction scenes is realized, and the user experience of voice interaction with the vehicle is improved.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1