Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

40 results about "Interactive audio" patented technology

Parkinson's disease auxiliary diagnosis method and system based on multi-stage interactive audio-visual fusion

The invention provides a Parkinson's disease auxiliary diagnosis method and system based on multi-stage interactive audio-visual fusion, and belongs to the technical field of auxiliary diagnos.The Parkinson's disease auxiliary diagnosis method comprises the steps that image data and audio data of a patient to be diagnosed are obtained, and the image data comprise a dynamic optical flow image and a static apparent image; a multi-scale dynamic and static feature extraction module is adopted to extract dynamic features and static features of the dynamic optical flow image and the static apparent image; inputting the dynamic features and the static features of each scale into a dynamic and static feature fusion module for fusion to obtain video modal features of each scale; a time-frequency attention speech feature extraction module is adopted to extract speech modal features from the audio data; fusing the video modal features of each scale and the voice modal features of the corresponding scale by using a global and local attention fusion module to obtain audio-visual fusion features; and inputting the audio-visual fusion features into a full-connection layer classification network, and outputting a classification result to assist diagnosis.
Owner:SHANDONG UNIV

Wearable interactive audio device

PendingUS20250392849A1Input/output for user-computer interactionMicrophonesTouch SensesPHYSICAL MANIPULATIONS
Embodiments are directed to a wearable audio device, such as an earbud. The earbud may be configured to detect input using various sensors and structures. For example, the earbud may be configured to detect gestures, physical manipulations, and so forth performed along or on the earbud. In response to the detected inputs, the earbud may be configured to change various outputs, such as an audio output or a haptic output of the device. The earbud may also include a microphone to register voice commands. In some cases, the microphone may be used to control the earbud using the registered voice command in response to one or more detected gestures or physical manipulations.
Owner:APPLE INC

Interactive audio-visual for english teaching

ActiveCN309850791SAudiovisual deviceIndustrial engineering
1. Name of the product in this design: Interactive Audiovisual Device for English Teaching. 2. Purpose of this design: For use in English teaching. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design's key points: 3D view 1.
Owner:黄靖 +1

Systems and methods for interactive video presentation of transactional information

ActiveUS12694381B2Visual presentationLanguage processor
Methods, apparatuses, and computer program products are described for presenting an interactive audio-visual presentation of transaction documents. A method can include receiving a bill associated with a payor and payee, using a textual language processor or the like to identify content fields from the bill and assign markups and / or metadata to content fields, and using the content fields, markups, and / or metadata to generate an audio-visual presentation associated with the bill. This audio-visual presentation can be presented to the payor. The payee may then interact with the audio-visual presentation, for instance by verbal, visual, manual, or textual response. A verbal language processing engine, natural language processing engine, audio-visual language processing engine, or visual-manual language processing engine can be initiated to facilitate interpretation of the payee response and generate a further audio-visual presentation.
Owner:PAYMENTUS CORP

Control method of interactive audio middleware, electronic equipment, medium and product

PendingCN121957931ALower the technical threshold for automated processingImprove versatilityInterprogram communicationInference methodsUser inputReusability
The invention discloses a control method of interactive audio middleware, electronic equipment, a medium and a product. The control method comprises the following steps: acquiring demand information input by a user; based on the demand information, obtaining a tool capable of being called in a model context protocol server and tool information corresponding to the tool, and determining a currently called target function and corresponding description information; determining a called target tool and related parameters in the model context protocol server, and initiating a calling request to the target tool through the model context protocol host; and judging the received execution result in combination with the demand information, and sending the execution result to the user through the model context protocol host, or determining a next tool called in the model context protocol server and related parameters corresponding to the next tool. According to the method, intelligent operation is carried out on the Wwise editor through the natural language, the technical threshold of Wwise automatic processing is lowered, and a Wwise operation tool with high universality and reusability is achieved.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

AI dialogue idle state detection method, AI dialogue system, equipment and storage medium

The invention discloses an AI dialogue idle state detection method, an AI dialogue system, an AI dialogue device and a storage medium. The method comprises the following steps: collecting audio data in real time and reporting the audio data to a server side; receiving an interactive audio packet replied by the server in response to the audio data, and reading the byte length of the interactive audio packet; in response to the situation that the byte length of the interactive audio packet does not exceed the byte length of the silent audio packet, counting the silent duration of continuously receiving the interactive audio packet which does not exceed the byte length of the silent audio packet; and in response to the silence duration reaching a preset duration threshold, disconnecting the connection with the server side. By means of the mode, whether the equipment end is in the idle state or not can be simply and effectively detected, the equipment end is actively disconnected with the server end in the idle state, unnecessary access to the server end is reduced, service charges are reduced, and service cost and equipment end operation cost are simply and effectively reduced.
Owner:SHENZHEN NEOWAY TECH

System and method for generating interactive media

An attraction system includes a display that operates to present augmented reality and / or virtual reality (AR / VR) imagery to a guest in an interactive space. The system includes an audio controller that operates an array of speakers that are distributed throughout the interactive space and a controller having one or more processors. The one or more processors are operable to receive data indicative of a state of the guest (e.g., an action, movement, or gesture of the guest, an input received from an input device). The one or more processors are also operable to adjust the AR / VR imagery in response to the state of the guest, and instruct the audio controller to operate the array of the speakers to provide interactive audio based on the state of the guest and such that the interactive audio presents to the guest as though originating from a dynamic portion of the AR / VR imagery.
Owner:UNIVERSAL CITY STUDIOS LLC

A classroom teaching mode analysis system and method based on a large language model

The application belongs to the teaching application field of information technology, and provides a classroom teaching mode analysis system and method based on a large language model, which comprises the following steps: (1) interactive audio extraction in a classroom teaching video; (2) conversion of dialogue audio into interactive text, matching of the text and a voice actor, and output of the interactive text in a dialogue form; (3) construction of an education mode prompt word; and (4) classroom teaching mode analysis, wherein the interactive text in the dialogue form obtained in step (2) is subjected to block coding, and then each block coding vector and the prompt word generated in step (3) are jointly input into a large language model, all output results are fused based on probability and weight, and a final classroom teaching mode analysis result is obtained. The application provides a new approach for classroom teaching video teaching mode analysis, and promotes intelligent understanding of a classroom teaching process.
Owner:HUAZHONG NORMAL UNIV

AI interactive audio learning machine

ActiveCN309858368SLearning machineInteractive audio
1. The name of the design product: AI interactive listening machine. 2. The use of the design product: for story listening, AI interaction. 3. The design points of the design product: in the combination of shape and pattern. 4. The picture or photo that best indicates the design points: perspective drawing.
Owner:BEIJING LINGJI TIANCI TECHNOLOGY CO LTD

Interactive audio entertainment system for vehicles

A system for interacting with an audio stream to obtain lyric information, control playback of the audio stream, and control aspects of the audio stream. In some instances, end users can request that the audio stream play with a lead vocal track or without a lead vocal track. Obtaining lyric information includes receiving via a text to speech module an audio playback of the lyric information.
Owner:CERENCE OPERATING CO

Virtual prop-based interaction method, apparatus and device, and storage medium

The invention discloses an interaction method and device based on virtual props, equipment and a storage medium, and belongs to the technical field of computers. The method comprises the steps that in the process of playing the interactive audio, a prompt animation of the interactive audio is displayed based on the aiming position of a virtual prop, and the prompt animation is used for prompting the playing progress of the interactive audio; when the aiming position is the position of the to-be-attacked object, in response to a release operation of the virtual prop, displaying a hit effect matched with the generation moment of the release operation; wherein the hit effects matched with different generation moments are different, and different generation moments correspond to different playing progresses of the interactive audio. Therefore, the interaction interest of the interaction object is improved, and the man-machine interaction rate is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Audio interaction method and program product

PendingCN122369434AUser inputSemantic feature
This application discloses an audio interaction method and program product. It includes: obtaining the acoustic feature vector of user-input audio; searching a composite vector database based on the acoustic feature vector to obtain target question-and-answer text that matches semantic feature vectors and sentiment feature vectors; and generating interactive audio based on the target question-and-answer text and the acoustic feature vector of the user-input audio. By retrieving and obtaining question-and-answer text that matches the acoustic feature vector of the user-input audio, and generating interactive audio based on the matched question-and-answer text and acoustic feature vector, the interactive audio can match the user's emotional tone while responding to the user's actual needs (i.e., the needs of semantic feature representation), achieving adaptive response to different user acoustic feature vectors. Furthermore, the analysis and matching of acoustic feature vectors reduces computational resource consumption to a certain extent, accelerates the generation speed of interactive audio, and further improves the user experience.
Owner:CHENGDU BOSS INNOVATION TECH CO LTD

Autonomous mobile platform with 3-d imaging system

An autonomous system having mobility, navigation, power, and general purpose computing. In some embodiments, the system comprises a base unit capable of sensing its environment and computing navigation instructions to direct the system to move to particular locations and execute functions as directed by a set of programmed instructions. In some embodiments, two or more sensors, such as 3-D cameras, with a field of view larger than 180° are attached to measure distance to objects in the environment. Cameras may also be used to recognize objects in the environment, and may also be used by the navigation system. In some embodiments, a coupling exists on the base unit to attach additional structures and mechanisms. These structures may be elements such as a means for carrying packages or other items, robotic manipulators to grab and move objects, interactive audio and video displays, or devices for serving food and drink.
Owner:UBIQUITY ROBOTICS INC

Cross-device collaborative film watching interaction method and device, smart television and storage medium

The invention discloses a cross-device collaborative film watching interaction method and device, a smart television and a storage medium, and relates to the technical field of smart device control, the method is applied to the smart television, the smart television is connected with an audio output device, and the method comprises the following steps: obtaining a preset plot element corresponding to to-be-played media data selected by a user; receiving user voice data collected by the audio output device, performing data matching on the user voice data and preset plot elements, and generating response text data according to a matching result; and converting the response text data into interactive audio data, and distributing the interactive audio data to an audio output device for playing when the media data to be played is output.
Owner:SHENZHEN ZHIXIAN VISION SOFTWARE TECHNOLOGY CO LTD

Quantitative analysis method and system for social behaviors of children and application

The invention belongs to the technical field of behavior recognition and medical auxiliary diagnosis, and relates to a child social behavior quantitative analysis method and system and application. The analysis method comprises the steps that interactive videos and interactive audios in a natural state are collected, multi-modal social behavior data of children in the natural state are obtained from the interactive videos and the interactive audios, and the multi-modal social behavior data comprise face data, audio data, text data and posture data; constructing corresponding quantitative features based on the data and fusing the quantitative features to obtain fused quantitative features; and configuring a pre-trained large language model, and inputting the fused quantitative features into the large language model to obtain an ASD risk probability, a multi-dimensional quantitative score vector, a frame-level anomaly probability sequence and a microscopic anomaly positioning timestamp set.
Owner:TIANJIN UNIV +1

English interactive audio-visual device

1. The name of the design product: English interactive audio-visual device. 2. The use of the design product: The design product is used for English teaching. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view 1.
Owner:沈军红

Method for identifying real-time interactive audio application traffic of mobile platform in encrypted tunnel

The application discloses a kind of identification methods for real-time interactive audio application flow in encrypted tunnel, the method includes: for typical real-time interactive audio application, analyze the correlation atlas of decompilation code, dynamic debugging application, build application behavior and traffic characteristics;Based on the protocol level rule, the packet level hierarchical space-time representation is carried out to encrypted tunnel flow, and the application is identified using the Transformer classifier;Finally, based on the application behavior-traffic characteristics correlation atlas, the packet level and flow level characteristics of the application flow in the encrypted tunnel are extracted in window, a feature matrix is constructed, and an integrated learning model is used to make a behavior decision for each window, and finally a complete behavior description of the traffic sample is formed.The application can well identify the attribution application and application behavior of real-time interactive audio application flow in encrypted tunnel by combining application reverse analysis and behavior traffic characteristics analysis with mainstream machine learning classifiers, which is of great significance for strengthening network supervision and maintaining public opinion environment.
Owner:NANJING UNIV OF SCI & TECH

English interactive audio-visual aid (teaching)

ActiveCN309581285SEngineeringInteractive audio
1. The name of the design product: English interactive audio-visual device (for teaching). 2. The use of the design product: The design product is used for English teaching. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view 1.
Owner:刘丹

Language interaction text information calibration method based on humiture large language model Agent

PendingCN121687015ASemantic analysisSpeech recognitionAudio segmentationEngineering
The invention discloses a language interaction text information calibration method based on a humiture large language model Agent, which relates to the field of speech recognition, solves the problem of poor calibration effect of the existing language interaction text information calibration method, and comprises the following steps: S1, acquiring real-time interaction audio and sampling to obtain a sampled audio clip, and storing the sampled audio clip in a database; the method comprises the following steps: S1, sampling an audio sample, carrying out audio analysis to obtain a speech loudness intrusive ratio and a fragment statement recognition coincidence ratio, and obtaining interactive audio collection data, S2, carrying out type division on real-time interactive audio by analyzing the interactive audio collection data, respectively carrying out audio statement segmentation to obtain real-time interactive audio segmentation data, and carrying out audio analysis on the real-time interactive audio segmentation data; and S3, performing vocabulary semantic verification on the text vocabularies according to the real-time interaction audio segmentation data, and performing vocabulary translation and outputting according to the verification result. The method can improve the pertinence and accuracy of the language interaction text information calibration method.
Owner:XINJIANG UYGUR AUTONOMOUS REGION INST OF MEASUREMENT & TESTING

System and method for interactive audio and visual accompaniment

PCT designated stageWO2026142475A1PersonalizationEngineering
The invention relates to the field of information technology and multimedia systems, and more particularly to methods and systems for creating personalized audio and visual accompaniment to reading written content, listening, inter alia to speech, or other ways of consuming content. The technical result consists in significantly shortening learning time, increasing reading and listening speed, improving comprehension and recall of a read text, and increasing the user's concentration. A system for creating interactive audio and video accompaniment to reading, listening, narrating or other ways of consuming content comprises a data input module, a data preprocessing module, a data analysis and interpretation module, a synchronization module, a physiological data module, an environmental data module, a data combining module, an audio and video accompaniment generating module, an enhancement module, an external device integration module, a personalization and settings module, a feedback and self-learning module, a collaboration module, and an audio and visual accompaniment display module.
Owner:KRIKUNOV ALEKSEY NIKOLAEVICH

Interaction processing method and device, electronic equipment, medium and program product

The invention relates to an interaction processing method and device, electronic equipment, a medium and program product quality. The method comprises the following steps: displaying a first preset recommendation media corresponding to a preset business party on a preset page; in the process of displaying the first preset recommendation media, displaying the first interaction guide information on a preset page; and in response to the preset interaction operation, playing a preset interaction audio corresponding to the first interaction guide information, and displaying a preset entry of a corresponding associated page of the preset service party, the preset interaction audio being an interaction audio customized by the preset service party for a media browsing account corresponding to the preset service party. According to the technical scheme provided by the embodiment of the invention, interactivity and interestingness of the service party in the recommendation process can be improved, meanwhile, the enthusiasm of the user for further knowing the recommendation information corresponding to the service party is better improved, and the recommendation effect of the related content of the service party is improved.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Voice control method and device for clothes processing equipment, equipment and storage medium

The invention provides a voice control method, device and equipment for clothes processing equipment and a storage medium, and the voice control method comprises the steps: a voice receiving step: receiving user voice information; a voice analysis step: analyzing the voice information to obtain clear control information, and generating a control command; in the voice analysis step, whether the voice information contains clothes processing related information or not is judged, if yes, the voice information serves as clear control information, if not, the voice information serves as fuzzy control information, and set interaction audio for guiding complementation is generated and played. And skipping to the voice receiving step. According to the method and the device, judgment is carried out through the clothes processing related information in the voice information, the user is guided to complement the information with fuzzy demands so as to clarify the demands, so that the control command can be accurately generated for control, mistaken setting caused by fuzzy voice commands is reduced, and the use experience of the user is improved.
Owner:PANASONIC APPLIANCES (CHINA) CO LTD +1

method

A method executed by an information processing apparatus that includes a controller, an imager, and an input interface includes executing, by the controller, operations including acquiring consent from a customer regarding recording of customer engagement audio via the input interface, starting audio recording after the consent is acquired, detecting a staff member from an image of the imager, and interrupting the audio recording in a case in which a predetermined condition is met, and the predetermined condition includes a first condition that the staff member has disappeared from the image of the imager.
Owner:TOYOTA JIDOSHA KK

method

A method executed by an information processing apparatus that includes a controller, an imager, and an input interface includes executing, by the controller, operations including acquiring consent from a customer regarding recording of customer engagement audio via the input interface, starting capturing an image using the imager after the consent is acquired, and starting audio recording when a predetermined condition is met, and the predetermined condition includes a first condition that a staff member is reflected in an image of the imager.
Owner:TOYOTA JIDOSHA KK

Apparatus, system, and method for motion sensing

Methods and devices provide physiological movement detection, such as gesture, breathing, cardiac and / or gross body motion, with active sound generation such as for an interactive audio device. The processor may evaluate, via a microphone coupled to the interactive audio device, a sensed audible verbal communication. The processor may control producing, via a speaker coupled to the processor, a sound signal in a user's vicinity. The processor may control sensing, via a microphone coupled to the processor, a reflected sound signal. This reflected sound signal is a reflection of the generated sound signal from the vicinity or user. The processor may process the reflected sound, such as by a demodulation technique, to derive a physiological movement signal. The processor may generate, in response to the sensed audible verbal communication, an output based on an evaluation of the derived physiological movement signal.
Owner:RESMED SENSOR TECH LTD

Bank customer insight and portrait updating method and system based on voice interaction analysis

PendingCN121306177AFinanceSpeech recognitionQuality of serviceCustomer insight
The invention discloses a bank customer insight and portrait updating method and system based on voice interaction analysis. The method comprises the following steps: collecting interactive audio of a bank staff and a customer through a portable audio collection device; transmitting to a remote processing end; processing the audio by using a natural language processing model, extracting customer financial demand characteristics based on a predefined financial product label library and a demand intention recognition model, extracting staff financial service performance characteristics based on a preset compliance rule library and a service quality evaluation model, and generating a bidirectional analysis result; generating output information based on the result; receiving feedback input of the staff and adjusting model parameters or strategies according to the feedback input; and automatically sending the customer insight data to the bank CRM system to update the portrait. According to the invention, by fusing hardware acquisition, AI analysis and system integration technologies, the technical problem that offline interaction data are difficult to carry out real-time structured analysis and form an automatic closed loop with an online business system is solved.
Owner:李雪梅

Method, system and device for dynamically adjusting audio of virtual scene and medium

The embodiment of the invention provides a method, a system, equipment and a medium for dynamically adjusting audio of a virtual scene, and relates to the technical field of audio processing, the method comprises the following steps: presenting an interactive interface of the virtual scene, an interactive object list exists in the interactive interface, and when the number of times of interaction between a user and Ai meets an interactive requirement corresponding to an interactive object, executing the interactive object list; obtaining an interactive audio playing at the current time point as an intermediate audio, obtaining an audio of a preset beat number of the intermediate audio after the current time point as a future audio clip, obtaining a first feature value list based on a frequency spectrum corresponding to the future audio clip, obtaining a second feature value list corresponding to the Bij based on a frequency spectrum corresponding to an initial audio clip of the Bij, the initial matching degree of the first feature value list and the second feature value list corresponding to the Bij is obtained, and the interactive audio and the intermediate audio with the highest initial matching degree are played together, so that the user experience is improved.
Owner:TIANJIN SENYUEXING INTELLIGENT TECH CO LTD

Interactive audio book with copper foil conductive principle

ActiveCN224457499UCopper foilAcoustics
This utility model discloses an interactive audiobook based on the conductivity principle of copper foil, belonging to the field of learning tools technology. It includes a cover and a spine. The spine has a side edge with a pre-set sliding track. Several distributed conductive lines are integrated inside the sliding track, and contact points are provided on the conductive lines. A matching slider assembly is slidably installed inside the sliding track, and the slider assembly is used to press against the contact points to conduct electricity through the conductive lines. A power supply is embedded in the spine and electrically connected to the conductive lines for power supply. This utility model utilizes the conductivity of copper foil and activates sound effects through a physical sliding mechanism along the track, providing children with a new interactive reading experience, enhancing the interactivity and fun of reading, and promoting the development of children's auditory perception, hand-eye coordination, and cognitive abilities.
Owner:HUAHUI PRINTNG PROD SHENZHEN CO LTD

Transmission of interactive audio content

Techniques for transmitting audio content are provided. In some embodiments, the techniques may involve capturing audio content from a plurality of microphones. The techniques may further involve encoding a subset of audio content channels to form a first audio stream and encoding the remaining subset of audio content channels to form a second audio stream. The techniques may further involve determining sensor data and device information. The techniques may further involve storing data associated with the second audio stream, the sensor data, and the device information as metadata, wherein the metadata includes data usable to decode and reconstruct the second audio stream. The techniques may further involve generating an audio package comprising the first audio stream and the metadata, wherein the audio package is usable by an interactive-compatible audio playback device to render interactive audio content and is usable by a non-compatible audio playback device to render non-interactive audio content.
Owner:DOLBY LABORATORIES LICENSING CORP