Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6040results about "Sound input/output" patented technology

Auditory user interfaces and associated systems, methods, devices, and non-transitory computer-readable media

An auditory operating system designed to facilitate context-aware, audio-based user interactions, particularly with artificial intelligence agents or applications. An auditory operating system shell serves as the primary interface, managing and coordinating multiple specialized agents that handle specific domains like music streaming, scheduling, or home automation. Using natural language processing, the auditory operating system shell identifies the appropriate agent or application for a user's command or query and ensures task execution and context preservation across interactions. The auditory operating system shell enforces privacy and stability by controlling agents' and applications' access to data and system privileges. The auditory operating system shell also supports dynamic context management, enabling seamless handoffs between agents when user requests span multiple domains. This auditory operating system shell may reduce the need for users to memorize specific wake words or commands, as the auditory operating system shell may determine the user's intent from speech or contextual cues.
Owner:IYO INC

Smart volume controller and method thereof

A controller configured to control a sound producing module includes a volume controlling unit configured to determine a demodulation amplitude and a modulation amplitude corresponding to a target volume. The sound producing module comprises a driving circuit and an air-pulse generating device. The driving circuit generates a demodulation driving signal according to the demodulation amplitude and generates a modulation driving signal according to the modulation amplitude, so as to drive the air-pulse generating device. The air-pulse generating device produces sound via generating a plurality of air pulses at an ultrasonic pulse rate.
Owner:XMEMS LABS INC

Digital human interaction system and method based on multi-modal emotion recognition

ActiveCN121116129ASemantic analysisSpeech analysisInteractive modelingData stream
The embodiment of the invention provides a digital human interaction system and method based on multi-modal emotion recognition, and belongs to the technical field of digital human interaction. The system comprises a multi-modal sensing module used for collecting multi-modal data and preprocessing the multi-modal data to generate a standardized data stream; the cross-modal fusion and emotion recognition module is used for carrying out interactive modeling on the multi-modal features and outputting a current emotion label and emotion intensity; the reaction planning module is used for generating a composite reaction strategy; and the digital human rendering module is used for mapping the composite reaction strategy into control signals corresponding to the voice, the facial expression and the action respectively, and driving a digital human to execute corresponding voice output, facial expression change and limb action through the control signals so as to realize interaction. According to the method, multi-modal data are deeply fused through the cross-modal graph neural network and comparative learning, the weight is dynamically adjusted in combination with the modal confidence, and the emotion recognition accuracy and robustness are improved.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Low-delay audio input switching method and system, storage medium and equipment

The invention relates to the technical field of audio control, and discloses a low-delay audio input switching method and system, a storage medium and equipment, and the method comprises the steps: carrying out the parallel pre-initialization of a plurality of pieces of audio input equipment, and building an independent parallel audio data cache for each piece of equipment; monitoring the state of each audio input device and the audio stream quality in real time, and judging whether switching is triggered or not based on a multi-factor decision model; after the switching decision is triggered, seamless audio data stream switching is executed, and format unification processing and cross fade-in and fade-out transition are included; a unified equipment operation interface is provided through the hardware abstraction layer, and system resources are optimized and managed; a delay sensing closed-loop control mechanism is constructed, processing delay of each link is monitored in real time, a caching strategy, a processing algorithm and resource allocation parameters are dynamically adjusted, self-adaptive balance of low delay and high tone quality is achieved, and through the method, quick and smooth switching of audio input equipment is achieved, delay is remarkably reduced, and real-time audio experience is improved.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Immersive audio-video follow-up adjustment method and system

The invention is applicable to the field of intelligent audio adjustment, and provides an immersive audio-video follow-up adjustment method and system, and the method comprises the steps: constructing a multi-dimensional perception system, and collecting multi-source information in real time; carrying out fusion processing on the collected multi-source information based on a deep neural network model, mining a dynamic mapping relation between a user state and the video content through space-time correlation analysis, and identifying a user interaction intention and an emotional tone and a space scene attribute of the video content; according to a fusion processing result, calling a dynamic parameter adjustment engine, and generating an audio parameter adjustment scheme in real time; a user experience feedback closed loop is constructed, a visual attention area of a user is collected through eye movement tracking equipment, and personalized adjustment preference parameters are generated; according to the method and the device, the audio effect is accurately matched with the user state and the audio and video content, the naturalness and the adaptability of immersive experience are remarkably improved, universality and individual differences are considered, and the audio experience which is more suitable for scenes and needs of the user is brought to the user.
Owner:SHENZHEN ZIDOO TECH CO LTD

Storing generated digital objects on a distributed ledger

Generative media content (e.g., generative audio) can be dynamically generated based on various inputs, which can include blockchain data. A playback device accesses blockchain data stored via a distributed ledger and generates media content based at least in part on the blockchain data. The playback device can access a library of pre-existing media segments and arrange a selection of pre-existing media segments from the library for playback according to a generative media content model and based at least in part on the blockchain data. The generated media content can then be played back via the playback device.
Owner:SONOS INC

Modular digital audio multi-channel IO control system and method for aeronautical communication

The invention belongs to the technical field of data processing, relates to a modular digital audio multi-channel IO control system and method for aviation communication, and aims to solve the problems of poor environmental adaptability, static synchronization parameters, calibration lag and insufficient intelligent fault-tolerant capability of a synchronization technology in the aviation field. The method comprises the steps that a scene sensing module fuses multi-domain data to generate a scene vector, synchronous parameters are generated through a look-ahead control module, clock skew of multiple IO channels is calibrated in combination with a synchronous execution module, and synchronous control is achieved; the health monitoring module quantifies the full-link health degree and dynamically adapts to the self-healing strategy through the self-healing learning module. According to the method, high-precision synchronization and high-robustness fault tolerance are balanced through perspective prediction and response type self-healing closed-loop cooperation, the system is endowed with intelligent self-adaptive evolution ability, and the synchronization performance and task reliability of the aviation audio system in a complex dynamic environment are improved.
Owner:CHINA SOUTHERN TECHNOLOGY (GUANGDONG HENGQIN) CO LTD

Digital large screen interaction method and device based on multi-agent cooperation and medium

The invention discloses a digital large-screen interaction method and device based on multi-agent cooperation and a medium, and relates to the technical field of digital large-screen interaction.The method comprises the steps that postures, expressions and voice data of a user are collected in real time, the attention weight is calculated in combination with spatial position information, the emotional state change rate is monitored, and the user experience is improved in combination with personal historical browsing preferences; acquiring a main interaction user, an emotional state feature and an emotional change trend; according to the main interaction user, the emotional state features and the emotional change trend, obtaining an initial confidence coefficient of the suspected interested field through a semantic understanding model; and based on the candidate guide images, identifying a user selection intention through a multi-modal fusion processing mechanism, calculating a selection probability, triggering content activation of the main interaction area and content dynamic generation of the auxiliary information area, synchronizing state information, and starting an immersive collaborative interaction process. Through three-level collaborative service configuration, the beneficial effects of improving large screen content organization efficiency and enhancing immersive interaction experience are achieved.
Owner:YLZ INFORMATION TECHNOLOGY CO LTD

Ending active noise cancellation based on a detected audio source

In aspects of ending active noise cancellation based on a detected audio source, a mobile device implements an audio playback manager that detects a headset employing active noise cancellation (ANC) of audio in communication with a mobile device in an environment of the mobile device. The audio playback manager receives a selection of at least one audio source that triggers adjustment of the ANC and detects sound originating from the at least one audio source that triggers adjustment of the ANC. Based on detecting the sound originating from the at least one audio source, the audio playback manager adjusts the ANC.
Owner:MOTOROLA MOBILITY LLC

Intelligent man-machine interaction system based on element universe

The invention relates to the technical field of man-machine interaction, and discloses an intelligent man-machine interaction system based on meta-universe, which is used for solving the problem of lagging caused by the fact that each virtual module possibly initiates response at the same time during man-machine interaction, and comprises the following steps: constructing a meta-universe scene, collecting multi-modal behavior information of a user in the meta-universe scene, and displaying the multi-modal behavior information of the user in the meta-universe scene. Performing unified formatting processing on the multi-modal behavior information to obtain formatted multi-modal behavior information, and performing joint reasoning on the formatted multi-modal behavior information by adopting a natural language processing engine and a multi-modal fusion model to obtain structured command data; the method comprises the following steps: obtaining a virtual module in a meta universe scene needing to make a response action according to structured command data, recording the virtual module as a driven module, obtaining priority information of all the driven modules, evaluating to obtain a priority index, and driving the modules to make the response action according to the priority index, thereby effectively improving the interaction response efficiency and the user experience continuity.
Owner:HANGZHOU FACTORIAL CLOUD TECHNOLOGY CO LTD

Interaction control method of intelligent glasses

The invention relates to the technical field of computers, and discloses an interaction control method of intelligent glasses. The method comprises the following steps: synchronously acquiring multi-modal data such as eye movement, voice, gestures and head postures and environment and application context information; carrying out independent time sequence feature coding on each modal data; generating a modulation vector in combination with the context, and outputting probability distribution of user intentions through a cross-modal attention fusion network; and a unique execution instruction is determined through an instruction arbitration module based on rules and a state machine. The system comprises corresponding function modules. According to the method, the accuracy, robustness and naturalness of interaction are improved through multi-modal synchronous fusion and a context self-adaption mechanism.
Owner:NINGBO JINSHENGXIN IMAGE TECH CO LTD

Real-time emotion perception and voice interaction system for intelligent cockpit

The invention discloses a real-time emotion perception and voice interaction system for an intelligent cabin. The system comprises a multi-modal data acquisition module used for synchronously acquiring a facial image, a voice signal and text input of a driver; the visual feature enhancement unit is used for carrying out restoration and emotion distribution extraction on the low-quality image; the audio noise reduction and feature extraction unit is used for extracting voice emotion features; the text emotion coding unit is used for fusing relative position coding and context semantic information; the cross-modal fusion module is used for outputting an emotion classification result by integrating visual, audio and text features; the personalized emotion database is used for storing historical emotion data of the user and performing emotion trend prediction and early warning judgment; the large language model feedback module is used for generating structured cue words according to the emotion recognition result and the driving situation and generating natural language feedback; the voice synthesis and output module is used for adjusting voice parameters and performing feedback output through a vehicle-mounted multi-channel; according to the invention, the emotion expression of human-vehicle interaction is enhanced.
Owner:SUZHOU UNIV

Transfusion monitoring system

The invention discloses an infusion monitoring system, which relates to the technical field of infusion monitoring, and comprises a controller with a display screen, an infusion pump arranged at the top of a main body case and used for regulating and controlling the infusion speed, an infusion tube fixing frame arranged on one side of the main body case, and an alarm device arranged on the surface of the main body case, a physiological parameter acquisition module, a medicine characteristic analysis module, a dynamic dripping speed regulation and control module, a liquid abnormity detection module, an AR interaction control module and an alarm cooperative processing module are arranged in the controller; a physiological parameter acquisition module, a medicine characteristic identification module, a dynamic dripping speed regulation module, a liquid abnormity detection module, an AR interaction control module and an alarm cooperative processing module are integrated. The intelligent and closed-loop infusion monitoring system is constructed, self-adaptive dripping speed adjustment based on the physiological state of the patient and the medicine characteristics is achieved, and the safety of infusion treatment is improved.
Owner:DONGGUAN BINHAI BAY CENT HOSPITAL

Machine learning based voice control for audio device

Various implementations include approaches for voice control in audio devices. In some cases, a method includes: listening, using at least one audio capture device, for user input to control at least one attribute of an audio device; routing the user input through a machine learning (ML) model to determine a control action for the at least one attribute based on the user input; and causing the determined control action to be performed, wherein the ML model need not have been pre-trained with the user input to determine the control action for the at least one attribute of the audio device.
Owner:BOSE CORP

Ending audio playback based on a trigger word

In aspects of ending audio playback based on a trigger word, a mobile device implements an audio playback manager that detects a headset in communication with a mobile device, the headset causing active noise cancellation (ANC) and outputting audio playback communicated from the mobile device. The audio playback manager detects a trigger word that is audible in an environment of the mobile device. In response to detecting the trigger word, the audio playback manager ends the audio playback and adjusts the ANC of the headset.
Owner:MOTOROLA MOBILITY LLC

Apparatuses and methods for an interactive device

A method includes receiving acoustic signals from a user. The acoustic signals are transformed into data. The data are transmitted to a voice-to-text conversion and artificial intelligence system over a network. Information is received from the voice-to-text conversion and artificial intelligence system. The information was obtained from the data and the information is used to create a request to a hotel. A computer readable medium contains executable computer program instructions, which when executed by a data processing system, cause the data processing system to perform a process that includes; receiving acoustic signals from a user; transforming the acoustic signals into data; transmitting the data to a voice-to-text conversion and artificial intelligence system over a network; receiving information from the voice-to-text conversion and artificial intelligence system, wherein the information was obtained from the data; and using the information to create a request to a hotel through a hotel information system.
Owner:ELECTRIC MIRROR LLC

Image display system and method of intelligent voice interactive electronic ink screen

The invention relates to the technical field of display, in particular to an image display system and method for an intelligent voice interactive electronic ink screen, and the system comprises a language interaction module which is used for carrying out the semantic processing of an accessed voice instruction, and generating corresponding description information; the image generation module is connected with the language interaction module and is used for generating target image data according to the description information and the display characteristic parameters of the electronic ink screen; the image display module is connected with the image generation module and used for driving the electronic ink screen to display images according to the target image data, and the problem that the existing image display method of the intelligent voice control electronic ink screen does not fully consider the specific display characteristic parameters of the electronic ink screen in the image generation process; therefore, the technical problems of display ghosting and slow refreshing are solved.
Owner:SHENZHEN WAVESHARE ELECTRONICS

End side assistant construction method based on multi-mode multi-agent driving

The invention discloses an end-side assistant construction method based on multi-mode multi-agent driving. The method comprises the following steps: defining an instruction mode and constructing a template library; an end-to-end reasoning framework based on a multi-modal large model is introduced to serve as a core base of a GUI agent; and based on a long-range task execution scheme of reinforcement learning and hierarchical planning, constructing a double-layer solution of reinforcement learning training and hierarchical planning execution. According to the method, the generalization ability, the single-step operation accuracy and the long-range task planning execution ability can be improved.
Owner:SHANGHAI JIAOTONG UNIV

Exhibition hall display method and equipment based on digital multimedia

The invention discloses an exhibition hall display method and device based on digital multimedia, relates to the technical field of digital multimedia, and discloses an exhibition hall display method and device based on digital multimedia, which are characterized in that the positions of audiences are acquired in real time, associated device subsets are matched, and cooperative control parameters are generated in combination with environment data and attention data; the multi-device cooperative display strategy is dynamically adjusted, the multimedia cooperative control precision and the display effect are improved, and the exhibition viewing experience of audiences is improved.
Owner:SICHUAN MOCAI CULTURE MEDIA CO LTD

Devices, Methods, And Graphical User Interfaces For Interacting With System User Interfaces Within Three-Dimensional Environments

While a view of an environment is visible via one or more display generation components of a computer system, and while the view of the environment includes a respective object that moves as a hand of a user moves, the computer system detects, via one or more input devices, a respective input. In response to detecting the respective input: in accordance with a determination that first criteria are met, the first criteria including a requirement that the hand of the user is holding a controller, the computer system displays, via the one or more display generation components, a first user interface object at a first location relative to the respective object; and, in accordance with a determination that the first criteria are not met, the computer system forgoes displaying the first user interface object at the first location.
Owner:APPLE INC

Methods and systems for generating and rendering object based audio with conditional rendering metadata

Methods and audio processing units for generating an object based audio program including conditional rendering metadata corresponding to at least one object channel of the program, where the conditional rendering metadata is indicative of at least one rendering constraint, based on playback speaker array configuration, which applies to each corresponding object channel, and methods for rendering audio content determined by such a program, including by rendering content of at least one audio channel of the program in a manner compliant with each applicable rendering constraint in response to at least some of the conditional rendering metadata. Rendering of a selected mix of content of the program may provide an immersive experience.
Owner:DOLBY LABORATORIES LICENSING CORP +1

Multi-mode sensing cooperation and real-time synchronization system

The invention discloses a multi-mode perception collaboration and real-time synchronization system, which breaks through the limitation of vision / hearing in the prior art, covers vision, hearing, smell and movement, and realizes more accurate cooperation through a perception layer, a control layer and an execution layer. According to the invention, an immersive experience space full of shock and reality can be created, synchronous output of visual, auditory, tactile and olfactory signals is realized by means of high-precision equipment, and a series of cross-sensory signal superposition schemes are designed on the basis of response characteristics of human physiology and psychology to various sensory stimuli. Therefore, the overall perception of the experiencer to the space atmosphere is enhanced. Meanwhile, through weight distribution and conflict resolution, the system can flexibly adjust the sensing mode according to the environment and the user state. The multi-device cooperation error can be controlled at millisecond level and millimeter level, and the natural fusion of virtuality and reality is realized.
Owner:BEIJING FANTASY PAI SHUSHI TECH CO LTD

User interfaces for audio routing

The present disclosure generally relates to user interface for audio routing. In some examples, a computer system detects an audio device and displays a selectable option in accordance with a determination that a set of criteria is met. The selectable option, when selected, enables the computer system to connect to the audio device and transmit audio data to the audio device.
Owner:APPLE INC

Unidentified aerial phenomena field disturbance detector

An unidentified aerial phenomena field disturbance detector includes an accelerometer configured to measure a gravitational field surrounding the unidentified aerial phenomena field disturbance detector; a magnetometer configured to measure a magnetic field surrounding the unidentified aerial phenomena field disturbance detector; a microwave frequency detector configured to measure a specific band of microwave frequencies surrounding the unidentified aerial phenomena field disturbance detector; a controller; and an optical indicator configured to optically communicate information to a user of the unidentified aerial phenomena field disturbance detector. The controller determines a change in the gravitational field surrounding the unidentified aerial phenomena field disturbance detector when generates gravitational field values are outside a predetermined range of non-event field values. The controller controls the optical indicator to optically communicate the determined change in the gravitational field surrounding the unidentified aerial phenomena field disturbance detector.
Owner:DIPOALA WILLIAM SAMUEL

Driving processing method and device, storage medium and electronic equipment

The embodiment of the invention discloses a driving processing method and device, a storage medium and electronic device.The method comprises the steps that in response to a voice recording instruction of a target user in the vehicle driving process, the voice stream of the target user is collected, driving context data is generated, the voice stream of the target user is converted into a target transliteration text, and the target transliteration text is sent to the storage medium; based on the target transcription text and the driving context data, a driving processing large model is adopted to carry out record event analysis to obtain event element fragments, and based on the driving context data, event semantic completion is carried out on the event element fragments to obtain event record information, based on the event record information, performing data mapping by adopting a pre-equipment-forgetting event data model to generate an event task execution script and an event task record, displaying the event task record to the target user, and responding to a record confirmation instruction of the target user. And starting an event task corresponding to the event task record based on the event task execution script.
Owner:BEIJING VISION WORLD TECH CO LTD

Smart Audio System For Use With An Information Handling System Audio System

A system, method, and computer-readable medium for performing a smart audio operation. The smart audio operation includes detecting a plurality of audio devices available for use by an information handling system; monitoring information handling system audio related information; analyzing the information handling system audio related information using a trained smart audio artificial intelligence model; and, automatically connecting the information handling system to a particular audio device of the plurality of audio devices based upon the information handling system audio related information.
Owner:DELL PROD LP

Immersive medical operation AR practical training system based on mixed reality and reinforcement learning and training method thereof

The invention discloses a medical operation AR practical training system based on mixed reality and reinforcement learning and a training method thereof, and relates to the technical field of medical education and medical operation training. The system comprises an MR terminal presentation and interaction module, a case and scene management module, a multi-mode operation data acquisition and evaluation module, a reinforcement learning training strategy engine module, a practical training process control module, a training examination management module and the like, a vivid medical scene is constructed in an MR environment, and data such as operation tracks, force and step time sequences of students are acquired in real time; the evaluation result is coded into a state to be input into the reinforcement learning model, the case difficulty, the prompting mode and the training rhythm are adjusted in a self-adaptive mode, individualized, procedural and quantifiable skill training and assessment are achieved, and the method is suitable for various operations such as cardio-pulmonary resuscitation, endoscopic surgery and puncture catheterization.
Owner:ARTIFACT R&D (SUZHOU) CO LTD

System and methods for audio data analysis and tagging

A system for automated processing and analysis of audio files for large data sets in a cloud environment. A unified analytic environment can integrate audio machine learning models for processing and analysis with a knowledge management system, including graph presentations of tracked entities, linked to audio files and / or associated translations and transcripts. Entities within such data can be searched or filtered and proposed for tracking, or identified as tracked objects. These features can allow triage and prioritization of audio files for analysis. User interfaces can facilitate feedback on transcription and translation outputs, thereby improving present outputs and future inputs and outputs. Entities speaking or referred to can be found, tagged, and distinguished in audio files (e.g., using speaker identification in audio files, text searching in transcripts, etc.) Users can provide feedback and input on various aspects of a system, to enhance or adjust initial automated or other machine learning outputs.
Owner:PALANTIR TECHNOLOGIES INC

Adjusting audio output in a vehicle

In aspects of adjusting audio output in a vehicle, a vehicle audio system implements an audio playback manager that detects a location of a first person in the vehicle, the first person located in proximity of a first speaker device configured for audio output. The audio playback manager also detects an additional location of a second person in the vehicle, the second person located in proximity of a second speaker device configured for the audio output. The audio playback manager determines a preference for the first person related to the audio output and an additional preference for the second person related to the audio output. Based on the preference of the first person or the additional preference of the second person, the audio playback manager adjusts a volume of at least one of the first speaker device or the second speaker device.
Owner:MOTOROLA MOBILITY LLC

Ai-based models for assistance with near-repetitve computer-based tasks

In one aspect, a device includes a processor system and storage accessible to the processor system. The storage includes instructions executable by the processor system to track user inputs as a user interacts with a first graphical user interface (GUI) presented on a display. The instructions are also executable to execute a prediction model to identify a near-repetitive action from the user inputs. Based on the identification of the near-repetitive action, the instructions are executable to present a foreshadow action on the display. The foreshadow action is selectable to command the processor system to autonomously perform a real action corresponding to the foreshadowed action.
Owner:LENOVO (SINGAPORE) PTE LTD