Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

444 results about "Voice assistant" patented technology

AR glasses control system and method

The invention discloses an AR glasses control system and method, and belongs to the field of artificial intelligence and AR glasses application, and the method comprises the steps: capturing a user face image, combining a voice assistant, collecting user voiceprint data, and carrying out the matching through a recurrent neural network; constructing a three-dimensional environment model, positioning and marking obstacles and interactive objects, and identifying objects in the environment; the method comprises the following steps of: tracking a gesture, acquiring a key point coordinate by adopting an OpenPose algorithm, identifying a predefined gesture by utilizing a support vector machine, and mapping the gesture into a specific operation instruction; analyzing the user intention based on natural language processing, and calling a corresponding function module according to an analysis result; acquiring a gazing direction of a user by using an eye movement tracking sensor, calculating a screen focus position through an algorithm, and dynamically highlighting options or displaying information according to a gazing point position; through multi-modal data fusion, gestures, voices and fixation point input are analyzed in a unified manner. By accurately understanding the intention of the user, the interaction precision of the AR glasses is improved, and the user experience is optimized.
Owner:SHIJIAZHUANG TIEDAO UNIV

Dialogue state tracking for voice assistants

Dialogue state tracking for voice assistants involves correctly tracking intent and entities of a task that a user is performing. A dialogue state, having a tracked intent and one or more tracked entities, can then be used to perform the task. Building a dialogue state tracking system within a voice assistant is not trivial. In some embodiments, a dialogue state tracking system involving one or more large language models can be implemented downstream of a natural language understanding system to produce the tracked intent and the one or more tracked entities. In some embodiments, a dialogue state tracking system involving one or more large language models can be implemented upstream of a natural language understanding system to produce rephrased natural language text, which is in turn processed by the natural language understanding system to produce the tracked intent and the one or more tracked entities.
Owner:ROKU INC

Interface display method and device

An interface display method and a device are described. An original interface may be moved to reserve space for displaying an interface of a first application such as a voice assistant, so that the interface of the first application does not block the original interface. This helps a user interact with the original interface and the interface of the first application. For example, an electronic device displays a first interface. The electronic device detects an operation in which the user indicates to start a first application. The electronic device displays the first interface in a first area of a screen, and displays a second interface in a second area of the screen. The second interface is an interface of the first application, and the first area does not overlap the second area.
Owner:HUAWEI TECH CO LTD

Method and system for processing a voice input in voice assistant devices

The disclosure relates to a method and apparatus for processing a voice input. The method includes identifying a first candidate text output from a first list of candidate text outputs, the first list of candidate text outputs being generated based on a voice input being received from a user. The method includes validating a feasibility of executing the first candidate text output based on an analysis of a first contextual data related to the first candidate text output, user-specific data, historical data related to previous validations of the voice assistant device, and one of an execution result or a no-action result corresponding to the first candidate text output. The method includes, based on the feasibility of executing the first candidate text output being validated as not feasible for execution, identifying failure information associated with the first candidate text output.
Owner:SAMSUNG ELECTRONICS CO LTD

High-fidelity end-to-end acoustic modeling method combining speech intelligibility reconstruction and autoregressive feedback optimization

PendingCN121565148ASpeech recognitionFrequency spectrumWeighting filter
The invention belongs to the technical field of speech recognition and acoustic modeling, and particularly relates to a high-fidelity end-to-end acoustic modeling method combining speech intelligibility reconstruction and autoregressive feedback optimization, which comprises the following steps of: firstly, establishing an intelligibility loss function based on a human speech intelligibility model; subjective definition features of the voice are reconstructed through a frequency spectrum reconstruction network, intelligibility related features are extracted and embedded in combination with a perceptual weighted filter bank, and the intelligibility related features and Mel frequency spectrum features are input into an acoustic encoder in parallel; based on this, by introducing a voice intelligibility reconstruction and autoregressive feedback mechanism, while the conciseness of an end-to-end voice modeling framework is maintained, systematic improvement of voice sharpness, naturalness and stability is realized, and the method can be widely applied to scenes such as intelligent voice assistants, voice transcription, virtual anchors, voice restoration, cross-language voice generation and the like. And the method has extremely high practical application value and popularization prospect.
Owner:GUANGDONG UNIV OF TECH

Vehicle control method and device

The invention discloses a vehicle control method, device and equipment and a computer readable storage medium, and belongs to the technical field of vehicles. The method comprises the following steps: in response to a starting instruction for a voice assistant function configured in a vehicle, determining the running state of the vehicle and the remaining power percentage of the vehicle; under the condition that the running state of the vehicle meets the running requirement and the percentage of the remaining electric quantity of the vehicle is larger than a target threshold value, image acquisition equipment installed in the vehicle is controlled to be started; a first video collected by an image collection device in a first time period is obtained, the first video comprises a first gesture action of a driver and passengers of the vehicle, and the starting time of the first time period is the time for starting the image collection device; determining a first control instruction corresponding to the first gesture action; and controlling a control system corresponding to the first control instruction to execute a control operation corresponding to the first control instruction. According to the method, the use frequency of the voice assistant function is improved, and the vehicle control efficiency is higher.
Owner:CHERY AUTOMOBILE CO LTD +1

Collaboration Between a Recommendation Engine and a Voice Assistant

A method comprising causing a voice assistant and a recommendation engine that are executing in an infotainment system of a vehicle to cooperate in processing a vehicle occupant's acceptance of a recommendation proposed by the recommendation engine by having an interface to enable the recommendation engine to provide recommendation context to the voice assistant to enable the voice assistant to resolve an ambiguity in the occupant's acceptance of the recommendation.
Owner:CERENCE OPERATING CO

Multi-mode rehabilitation training system for children with cerebral palsy based on artificial intelligence

The invention belongs to the technical field of artificial intelligence, and discloses a cerebral palsy child multi-mode rehabilitation training system based on artificial intelligence. The system comprises a data acquisition and preprocessing module, a multi-modal information fusion module, a personalized rehabilitation training suggestion generation module, an interaction and excitation module, a progress tracking and abnormity early warning module, a doctor and parent end APP management module and a system control and management module. According to the invention, when the interaction and excitation module and the multi-modal information fusion module are combined and applied to rehabilitation training of children with cerebral palsy, the training effect can be significantly improved, and accurate evaluation and personalized training can be realized; the child patient can be trained in a more real and attractive environment, meanwhile, the enthusiasm of the child patient is further stimulated through the gamification excitation design, and the training motivation of the child patient is enhanced through real-time feedback encouragement.
Owner:HEBEI UNIVERSITY

Method and system for processing a voice input in voice assistant devices

PCT designated stageWO2025165147A1Speech recognitionEngineeringSpeech recognition
The disclosure relates to a method and apparatus for processing a voice input. The method includes identifying a first candidate text output from a first list of candidate text outputs, the first list of candidate text outputs being generated based on a voice input being received from a user. The method includes validating a feasibility of executing the first candidate text output based on an analysis of a first contextual data related to the first candidate text output, user-specific data, historical data related to previous validations of the voice assistant device, and one of an execution result or a no-action result corresponding to the first candidate text output. The method includes, based on the feasibility of executing the first candidate text output being validated as not feasible for execution, identifying failure information associated with the first candidate text output.
Owner:SAMSUNG ELECTRONICS CO LTD

Method and System to Integrate a Large Language Model with an In-Vehicle Voice Assistant

Methods, computing systems, and technology for personalizing a user experience in a vehicle. For example, a computing system may be configured to receive a user prompt from a user, wherein the user prompt is input into a voice assistant on-board the vehicle. The computing system may be configured to process the user prompt with a master language model agent. The master language model agent may determine, based on environmental data, one or more intermediate actions responsive to the user prompt. The master language model agent may evaluate, based on the environmental data, whether each intermediate action of the one or more intermediate actions satisfies the user prompt. The computing system may be configured to output one or more command instructions to the voice assistant, wherein the one or more command instructions cause the voice assistant to provide, using one or more human-machine interfaces, a response to the user prompt.
Owner:MERCEDES BENZ GROUP AG

Multi-modal emotion recognition fusion and communication method oriented to real-time human-computer interaction

The invention relates to a multi-modal emotion recognition fusion and communication method oriented to real-time human-computer interaction, and aims to solve the problems of asynchronous modal time sequence, non-uniform feature dimension, single fusion mode, high communication delay and the like of an existing emotion recognition system and improve the recognition precision and real-time performance of the system. According to the method, voice, images and physiological signals are synchronously collected through a multi-modal input module, feature alignment is achieved through time sequence interpolation, dynamic time warping and space coordinate mapping, double-domain feature fusion is conducted in combination with a time path network and a space path network, emotional state judgment is completed through a lightweight neural network, and the emotional state recognition accuracy is improved. And outputting six types of basic emotions and confidence coefficients. Meanwhile, a low-delay communication protocol based on UDP clipping extension realizes rapid feedback of emotion data, and in combination with modal priority scheduling, data compression and bandwidth sensing mechanisms, high-efficiency and low-delay transmission is ensured, and the real-time interaction requirement in a weak network environment is met. The method has the advantages of high accuracy, low power consumption, low time delay and flexible deployment, is suitable for various real-time interaction application scenes such as voice assistants, virtual customer service, emotion accompanying and telemedicine, and has wide application value and market prospect.
Owner:ZHONGSHAN INST OF CHANGCHUN UNIV OF SCI & TECH

Speech synthesis method and system for controllable latent variable modeling based on semantic distillation

The invention relates to the technical field of speech synthesis, and particularly discloses a speech synthesis method and system for controllable latent variable modeling based on semantic distillation, and the method comprises the steps: converting a Mel spectrum into continuous latent variable distribution through a speech coding module, generating continuous latent variables through re-parameterization sampling, introducing a self-supervised model for semantic distillation, and carrying out the semantic distillation. According to the method, alignment of latent variables and semantic features is constrained through marginal cosine similarity and distance matrix structure loss, a text encoder maps a phoneme sequence into latent variable distribution, time sequence alignment of a text and the latent variables is achieved in combination with monotonic alignment search, and a decoder reconstructs the latent variables into a Mel spectrum. According to the method, waveform synthesis through a vocoder and total loss function joint optimization reconstruction, KL divergence, distillation, text alignment and confrontation loss are carried out, discrete information loss is avoided through continuous latent variable modeling, semantic consistency and text alignment efficiency are enhanced, the naturalness, coherence and real-time performance of synthesized voice are improved, and the method is suitable for scenes such as voice assistants and virtual anchors.
Owner:BEIJING TIMES RUILANG TECH CO LTD

Conversation method and electronic equipment

The embodiment of the invention provides a dialogue method and electronic equipment, and the method comprises the steps: displaying an interface of a main session in a first application, the main session being a session between a user and a first intelligent assistant provided by the first application; in response to a first operation of the user in the interface of the main session, displaying an interface of at least one independent session in the first application, the at least one independent session including a first independent session, the first independent session being a session between the user and a second intelligent assistant, the second intelligent assistant calls the intelligent dialogue service for the first application aiming at the first scene. According to the dialogue method and the electronic equipment provided by the embodiment of the invention, the use efficiency of the intelligent voice assistant can be improved.
Owner:HUAWEI TECH CO LTD

Voice assistant awakening method, vehicle-mounted terminal and storage medium

The invention provides a voice assistant awakening method, a vehicle-mounted terminal and a storage medium, and relates to the technical field of voice assistants. According to the method, identity information of a user sending suspected wake-up voice data can be recognized, and an initial wake-up threshold value associated with the identity information of the user and historical wake-up habit data of a voice assistant are searched from a preset database; and adjusting the initial wake-up threshold according to the similarity between the real-time wake-up habit data and the historical wake-up habit data and the noise energy to obtain a real-time wake-up threshold. The real-time wake-up threshold is negatively correlated with the similarity and positively correlated with the noise energy, so that the reliability of the obtained real-time wake-up threshold is high, and the suspected wake-up voice data is evaluated to obtain the target wake-up value; and when the target wake-up value is greater than or equal to the real-time wake-up threshold value, waking up the voice assistant. Because the reliability of the real-time wake-up threshold is high, the wake-up of the voice assistant at the moment does not belong to false wake-up, and the sensitivity of wake-up of the voice assistant is also high.
Owner:VOYAH AUTOMOBILE TECH CO LTD

Temporary Configuration of A Media Playback System

Example techniques may involve temporary configuration of a media playback system in a place of accommodation, such as a hotel. In particular, the media playback system in a guest's room is configured with one or more settings of the guest's home media playback system. Example settings include user accounts of a various services, such as streaming audio services and / or voice assistant services. Other example settings include artists, albums, audio tracks, audio books, stations, and other audio content that the guest previously designated as a favorite using their home media playback system. When the guest leaves (e.g., checks-out of) of the place of accommodation, these settings are removed from the media playback system in the guest's room.
Owner:SONOS INC

Picture book interaction method and picture book content artificial intelligence generation method and system

The invention discloses a picture book interaction method and a picture book content artificial intelligence generation method and system.The picture book content artificial intelligence generation system comprises picture book application software and a server system, and the picture book application software comprises a picture book interaction AI engine front end SDK and a role playing mode program; the server system comprises a picture book interaction AI engine back-end service system, an interactive picture book media resource library system and a picture book content artificial intelligence generation system, and artificial intelligence processing of voice recognition is changed from a current television voice assistant to the picture book interaction AI engine front-end SDK and the picture book interaction AI engine back-end service system. Therefore, the picture book application software is independent of the artificial intelligence processing capability of the intelligent voice assistant of the existing television (set top box), not only is the adaptive work of the intelligent voice assistant of a television manufacturer reduced, but also the service data security protection of the picture book application software is realized.
Owner:华数传媒网络有限公司

AI voice assistant implementation method and equipment based on cloud computer terminal and medium

The invention discloses an AI voice assistant implementation method and equipment based on a cloud computer terminal and a medium, belongs to the technical field of cloud computers and voice assistants, and aims to solve the technical problems of how to apply an AI voice assistant to the cloud computer terminal and awaken the assistant by customizing an awakening word, fill the blank that the cloud computer terminal does not belong to a voice assistant of the cloud computer terminal and improve the voice assistant implementation efficiency. According to the technical scheme, voice wake-up and interaction based on the Android platform comprises the steps that a background thread decompresses a voice model, a microphone monitoring service is automatically started after decompression is completed, a voice recognition service continuously captures microphone input, and a recognition result is returned in real time; when a preset wake-up word is detected, popping up a floating window to enter an interaction mode, managing text display and recording states in an interaction process, and realizing timeout automatic processing through a delay task; speech recognition and interaction control based on streaming audio; performing asynchronous voice question-answer response; asynchronous speech synthesis and transmission; and playing the audio after the voice synthesis is successful.
Owner:INSPUR COMM TECH CO LTD

Instant intention unlocking system based on facial recognition and voice interaction

The invention discloses an instant intention unlocking system based on facial recognition and voice interaction, and aims at solving the problems that an existing voice assistant depends on wake-up words, interaction is low in efficiency, scene perception is weak, and energy consumption is high. The system adopts a three-layer architecture including a hardware layer (a camera, a microphone, a processor and a storage module), a software layer (a voice recognition engine, a face recognition algorithm, a scene perception module and an authority control module) and a logic layer (an intention fusion judgment module and a task routing and execution module). 'gaze + voice 'is taken as core triggering, and wakeup words are avoided; face recognition is used for confirming identity and watching state, voice recognition is used for analyzing instructions, and an application scene is fused to generate an instant intention triggering model; authority hierarchical management and control are realized, and energy consumption optimization, auxiliary triggering, cross-terminal synchronization and emergency help modules are arranged. Interaction efficiency can be improved, safety and convenience are balanced, energy consumption is reduced, and the method is adaptive to multiple terminals and multiple scenes.
Owner:冯东坡

Medical voice interaction dynamic privacy protection method and related equipment

The invention discloses a medical voice interaction dynamic privacy protection method and related equipment, and the method comprises the steps: carrying out the voiceprint recognition of a current user in response to a session request, determining the role of the current user based on a voiceprint recognition result, and dynamically authorizing the data access authority of the current user according to a predefined access control strategy; retrieving related historical dialogues and medical knowledge within the boundary of the data access permission, and generating personalized answers; voiceprint monitoring is continuously carried out during the session, and when switching of the current user is detected, a dynamic privacy protection mechanism is triggered; the dynamic privacy protection mechanism dynamically calculates privacy protection intensity based on multiple factors of voiceprint confidence, dialogue keyword sensitivity, dialogue intention sensitivity and a time decay factor, adds adaptive noise to the personalized answer according to the privacy protection intensity, and then outputs the personalized answer. The invention aims to improve the privacy protection capability of the intelligent voice assistant in the medical scene, so that the intelligent voice assistant can adapt to the complex multi-role cooperation requirement of the medical scene, and the dynamically changing session scene can be effectively processed.
Owner:XI AN JIAOTONG UNIV

Controllable voice generation method and device based on multi-agent dynamic scheduling

The invention discloses a controllable voice generation method and device based on multi-agent dynamic scheduling, and belongs to the technical field of voice synthesis, and the method comprises the steps: constructing a multi-agent voice generation frame comprising a central scheduling module, an identity agent, an emotion agent and an environment agent; a central scheduling module analyzes a user instruction and outputs a structured task plan to drive each agent to generate primary voice output of identity, emotion and environment dimensions, and a cooperation cost matrix between the agents is constructed according to the primary voice output; the optimal execution path of the multiple agents is solved through an optimization algorithm based on the cost matrix, the agents are executed in a cascading mode based on the optimal execution path, and finally sound mixing is completed and high-quality voice is output. According to the method, the naturalness, the semantic consistency and the overall quality of the synthesized voice in a complex scene can be remarkably improved, and the method is suitable for various man-machine interaction application scenes such as intelligent voice assistants, virtual digital humans, immersive entertainment, barrier-free voice services, personalized content creation and the like.
Owner:ZHEJIANG UNIV OF TECH

Device, system and method for configuring a voice assistant feature for rental radio

A system, method, and device are provided for configuring a voice assistant feature in portable radios for the rental radio market. The portal receives a request to configure the rental radio with a voice assistant feature as part of a rental registration for the event. The portal obtains access to a customer database. Based on that access, the portal obtains customer deployment context data. Customer environment functions are identified from the customer deployment context data for the event. Voice commands are generated by the portal for the identified customer environment functions. A mapping of voice commands to respective customer environment functions for the event is created and provided for customer verification and, if desired, customization. The portal then programs the portable rental radio(s) with the mapped voice commands to respective customer environment functions to complete the voice assistant configuration. The voice assistant configuration does not rely on natural language processing.
Owner:MOTOROLA SOLUTIONS INC

Generative offer system with multilingual ai-driven voice assistant for real-time commerce

A system and method for generating and delivering personalized promotional offers in real time in response to user voice input. The system processes natural language queries and analyzes prior purchase history, inferred preferences, historical interactions, environmental context including device type, location, and time, and ongoing behavioral signals. Machine learning models adapt future offers based on accumulated user data to provide evolving personalization. The system supports multilingual speech recognition and can operate across smartphones, voice assistants, wearables, smart televisions, extended reality environments, and large language models. Offers may be delivered through audio, visual, or haptic interfaces to enable integration into commerce and advertising platforms. The combination of speech-triggered interaction, adaptive personalization, and contextual awareness enables dynamic, AI-driven promotional flows that operate consistently across multiple platforms and languages.
Owner:BYRD STEPHEN M

AI voice assistant semantic enhancement method oriented to education scene

The invention discloses an AI voice assistant semantic enhancement method oriented to an education scene, and the method comprises the following steps: S1, collecting classroom voice signals in real time through an annular multi-microphone array which is installed at the edge of a blackboard in an embedded manner, and obtaining the real-time three-dimensional coordinates of a teacher through a UWB positioning module; s2, constructing a spatial sound source mapping model based on the three-dimensional coordinates of the teacher; s3, acquiring blackboard writing character content through a visual sensor mounted on the side edge of the blackboard; s4, injecting the extracted feature keywords into a dynamic term library generation module to generate a standardized teaching instruction, and synchronizing the standardized teaching instruction to student terminal equipment; and S5, when detecting that the confidence of the voice instruction is lower than a preset threshold value, jointly analyzing the time sequence characteristics of the lip movement track and the blackboard writing handwriting to perform semantic completion. According to the method, an intelligent semantic analysis system with acoustic enhancement, term adaptation and multimode verification is constructed, and the semantic distortion problem caused by environmental interference in traditional speech recognition is eliminated through directional pickup and noise suppression of an intelligent array.
Owner:GUANGZHOU DAZZLE VIEW INTELLIGENT TECH CO LTD

Display device for onboarding plurality of voice assistants using plurality of QR codes, electronic device, and method for controlling display device and electronic device

A display device may include a display, and one or more processors which receive a user command for onboarding a plurality of voice assistants to the display device, and controls the display to display a screen including a plurality of QR codes corresponding to the plurality of voice assistants based on order information of the plurality of voice assistants, wherein each of the plurality of QR codes includes a URL of a server at which login of user accounts corresponding to the voice assistants are performed.
Owner:SAMSUNG ELECTRONICS CO LTD

Chinese dialect English speech processing system and method

According to the method, a Chinese dialect English speech corpus covering multiple regions and multiple dimensional variables is constructed, the influence of Chinese dialect phoneme migration on speech parameter differences is deeply analyzed, and a Fujissaki model, a pinyin phoneme theory and a hidden Markov model (HMM) are creatively fused; high-precision automatic speech recognition (ASR), high-naturalness speech synthesis (TTS) and intelligent pronunciation correction of Chinese dialect English are realized. The method is suitable for the application fields of domestic English education, cross-regional multi-language interaction, intelligent voice assistants and the like, and effectively solves the problems of low recognition rate, separation of synthetic voice and actual accent and the like when a traditional system processes Chinese dialect English.
Owner:SHANDONG FOREIGN TRADE VOCATIONAL COLLEGE

Voice assistant interaction method and system

The invention discloses a voice assistant interaction method and system, the system comprises an end side device layer, a cloud service layer and a communication interaction layer, the end side device layer comprises at least two terminal devices, and each terminal device is provided with a multi-mode sensor. According to the invention, single-mode noise interference is compensated through multi-mode data fusion, which is significantly superior to the existing voice assistant; a personalized strategy is generated based on a triple of intention, emotion and equipment, so that the interaction experience is improved; local priority synchronization is realized by adopting a double-layer MQTTBroker architecture, the state synchronization delay of cross-network segment equipment is reduced, multi-equipment collaboration is supported, and seamless switching among equipment is realized.
Owner:NANJING HUJU LONGPAN INTELLIGENT TECH CO LTD

System and method for activating a voice assistant for a vehicle

A system for activating a voice assistant of a vehicle includes a microphone, a first wireless module, a second wireless module, and a controller in electrical communication with the microphone, the first wireless module, and the second wireless module. The controller is programmed to transmit a plurality of original training signals on a plurality of subcarriers using the first wireless module. The controller is further programmed to receive a plurality of propagated training signals using the second wireless module. The controller is further programmed to determine a deviation between the plurality of original training signals and the plurality of propagated training signals. The controller is further programmed to identify a motion marker based at least in part on the deviation. The controller is further programmed to activate the voice assistant of the vehicle to receive a voice command using the microphone based at least in part on the motion marker.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Vehicle-mounted voice assistant testing method, device and equipment

The invention provides a vehicle-mounted voice assistant testing method, device and equipment, and is applied to the technical field of voice testing. The method comprises the following steps: generating voice test cases in various test scenes according to vehicle information of a vehicle; and playing the corresponding test voice based on the voice test case. Through multiple acquisition ways, the execution information of the vehicle machine system of the vehicle for the test voice is obtained. The execution information of each collection path is compared with the expected information corresponding to the voice test case, the test result of the vehicle-mounted voice assistant is obtained, and the problems that a traditional vehicle-mounted voice assistant test method is high in test cost, low in test efficiency and limited in test function can be solved.
Owner:CHERY AUTOMOBILE CO LTD

Robot with charging stand

This product is a robot equipped with a charging stand that allows for two-way communication with the user using artificial intelligence. It has a voice assistant function and displays animated facial expressions on its display unit.
Owner:ANKER INNOVATIONS TECH CO LTD