Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

31 results about "Auditory information" patented technology

Auditory memory is the ability to process information presented orally, analyze it mentally, and store it to be recalled later. Those with a strong capacity for this type of memory are called auditory learners. The ability to learn from oral instructions and explanations is a fundamental skill required throughout life.

system

We provide the system. [Solution] A device for analyzing visual and auditory information collected in the work environment, A device that identifies tasks that can be automated based on the aforementioned analysis, A device that proposes an optimal automation method for the aforementioned task, A system that includes this.
Owner:SOFTBANK GROUP CORP

Communication method and system of AR glasses with earphone and glasses separated, and medium

The invention provides a communication method and system for AR glasses with separated earphones and glasses, and a medium, and the method comprises the steps: initializing the AR glasses, and obtaining the initialization state information; scanning compatible state information between the AR glasses and the separated audio equipment based on the initialized state information; establishing communication state information between the AR glasses and the separated audio equipment based on the compatible state information; performing communication matching authentication based on the communication state information between the AR glasses and the separated audio equipment to obtain an authentication result; analyzing whether the communication is successful or not based on the authentication result, if the communication is successful, outputting audio information and instruction response, and if the communication is failed, re-pairing the separated audio equipment; communication matching is completed by analyzing the communication state of the AR glasses and the separated audio equipment, so that successful communication is realized, cooperative transmission of visual information at a glasses end and auditory information at an earphone end is realized, and synchronism of audio and pictures is ensured.
Owner:SHENZHEN ORANGE ELECTRONICS CO LTD

System

An object of a system according to an embodiment is to perform information presentation utilizing multiple senses using generated AI.SOLUTION: A system includes a visual information presentation part, an auditory information presentation part, a tactile information presentation part, a taste information presentation part, and an olfactory information presentation part. The visual information presenting unit presents visual information by using the generated AI. The audio information presenting unit presents audio information by using the generated AI. The tactile information presenting unit presents tactile information by using the generated AI. The taste information presenter presents taste information by using the generated AI. The olfactory information presenting unit presents olfactory information by using the generated AI.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Control method of teaching device, control device, teaching system, and storage medium

The application discloses a teaching equipment control method, a control device, a teaching system and a storage medium. The teaching equipment control method comprises the following steps: image and audio of a target in a teaching space are collected to obtain image data and audio data of the target, wherein the teaching space comprises the teaching equipment; visual information of the target is extracted by using the image data of the target, and auditory information of the target is extracted by using the audio data of the target; and the teaching equipment is controlled based on the visual information and the auditory information of the target. In the foregoing manner, the convenience of the teaching equipment control can be improved, and the accuracy is relatively high.
Owner:IFLYTEK LINGZHI (JIANGSU) TECH CO LTD +1

system

We provide the system. [Solution] A device that performs information processing to generate an individually optimized learning plan based on the learner's abilities, progress, and goals, A device that provides educational materials based on a generated learning plan, A device that analyzes learners' activities with educational materials in real time and generates evaluation results, A device that analyzes a learner's past learning history to identify learning challenges and provides additional exercises to improve those challenges, A sensor that acquires sound and video, and a device that uses the acquired data to perform interactive speech practice in real time, A device that provides immediate instruction using visual and auditory information based on speech evaluation data, A system that includes this.
Owner:SOFTBANK GROUP CORP

Robot control system, robot control program

PendingCN122319059ALittle fingerEngineering
The hand tool 50 is equipped with sensor groups with different information acquisition functions on each of the fingers 22A, 22B, 22C, 22D, and 22E. In other words, it is equipped with sensor groups for acquiring various information required for multimodal robot control. A visual sensor 51A for detecting visual information is installed on the thumb, an auditory sensor 51B for detecting auditory information is installed on the index finger, an olfactory sensor 51C is installed on the middle finger, a tactile sensor 51D for detecting tactile information is installed on the ring finger, and a taste sensor 51E for detecting taste information is installed on the little finger.
Owner:SOFTBANK GROUP CORP

Artificial intelligence-based english spoken language pronunciation correction system

PendingCN122392380AVisually impairedMemory retention
The application belongs to the field of intelligent teaching, and particularly relates to an English oral pronunciation correction system based on artificial intelligence, which comprises a multi-modal perception module, a deep pronunciation feature extraction module, a cognitive compensation fusion module and an intelligent correction and training planning module; the abstract pronunciation is converted into a tactile experience through a multi-modal learning process, the tactile information and the auditory information are synchronously input, the memory retention rate of the visually impaired students is improved, and the students are guided to deeply remember the acoustic characteristics of the pronunciation; the application simulates the auditory-tactile joint representation ability of the visually impaired brain through the learnable attention weight and the cognitive compensation prior mask, realizes the real sense of "touch instead of vision", simultaneously identifies the error reasons through multi-task learning, provides scientific basis for targeted training, designs a special tactile generation network, maps the acoustic characteristics and the pronunciation error types into tactile patterns, provides instant tactile correction feedback for the students, and forms a closed-loop multi-modal learning experience.
Owner:SHANDONG VOCATIONAL COLLEGE OF SPECIAL EDUCATION

system

We provide the system. [Solution] Features that acquire the learner's visual and auditory information in real time, The function analyzes the data obtained by the aforementioned function and identifies the mental state related to the learner's level of understanding, A function that generates individualized instruction and practice problems based on the mental state identified by the aforementioned function, The function presents the instruction and problems generated by the aforementioned function to the learner, The machine acts as a learning supporter within the home, providing optimal learning materials according to the learner's progress, A system that includes this.
Owner:SOFTBANK GROUP CORP

AI vision camera (HUSKYLENS)

1. Name of the product in this design: AI Vision Camera (HUSKYLENS). 2. Purpose of this design: This product is used to collect, identify and process visual and auditory information, and output the results. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design's key points: a 3D model.
Owner:SHANGHAI ZHIWEI ROBOT CO LTD

system

The system according to this embodiment aims to enable visually impaired or bedridden people to enjoy video content. [Solution] The system according to the embodiment comprises an acquisition unit, an analysis unit, a language conversion unit, a speech conversion unit, and a provision unit. The acquisition unit acquires auditory information from video content. The analysis unit analyzes the visual information contained in the video based on the auditory information acquired by the acquisition unit. The language conversion unit translates the visual information analyzed by the analysis unit into language. The speech conversion unit converts the visual information translated by the language conversion unit into speech. The provision unit provides audio content by combining the speech converted by the speech conversion unit with the auditory information.
Owner:SOFTBANK GROUP CORP

A multi-modal fusion video classification method and system based on brain-like feedback interaction

The application discloses a kind of multi-modal fusion video classification method and system based on brain-like feedback interaction, method includes: video pre-processing, obtain visual information and auditory information in video;After feature extraction, input into the multi-modal fusion framework based on neural network architecture search to obtain fusion information representation after fusion;Fusion information representation, visual information and auditory information are input into feedback module based on feedback modulation effect generated by single sensory cortex after auditory and visual information integration of human brain in superior temporal sulcus, and the output multi-modal fusion visual information and multi-modal fusion auditory information are respectively through fully connected layer to obtain the confidence of each classification;The confidence of each classification is input into DS decision fusion module, and the final classification result is obtained.The application learns from the processing mode of human brain to perceive external environment, integrates multiple senses from various modal perception information, and effectively improves the accuracy of recognizing and classifying the expressions of people in the video.
Owner:HOHAI UNIV

Speedometers that utilize touch and hearing

PendingJP2026112346ARoad vehicles traffic controlAudible meter reading indicationAcousticsTesting Methods
The present invention provides a device that divides the vehicle's speed into multiple stepped speed ranges and outputs different tactile or auditory information according to the number of steps in each speed range. [Solution] The number of speed range steps in the tactile speedometer is changed from the usual decimal number to a lower base n number, and the number of different values ​​used by each digit of the new base n number is reduced, thereby reducing the total number of numerical patterns used by the tactile speedometer to display numerical values, making it easier to understand the number of speed range steps via touch. Furthermore, it is possible to change the type of base n number used for each speed range step, so that the speed range indicated by the first digit of each base n number is shown in the same speed range unit across the entire speed range, making it possible to unify the speed range of the first digit. This enables the speedometer to display speed in a unified speed range unit.
Owner:大庭 有二

Multi-modal collaborative perception based multi-level privacy protection method, device and equipment for care scene

This invention provides a multi-level privacy protection method, device, and equipment for caregiving scenarios based on multimodal collaborative perception. By constructing a multimodal neural network model at the caregiving data acquisition terminal and fusing visual and auditory information for deep semantic analysis, it effectively overcomes the limitations of traditional single-modal perception in complex environments such as low light and occlusion. This invention employs a collaborative design of delay buffering and pre-judgment, forcibly placing privacy level determination before media stream distribution, fundamentally eliminating the risk of instantaneous leakage of privacy images due to inference delays. By introducing a two-level dynamic judgment strategy of prioritizing high-risk feature blocking and scene adaptive fusion, the robustness of the system in real-world home environments is improved. Based on the matching relationship between user permissions and privacy levels, this invention performs differentiated media stream processing on the terminal side, achieving fine-grained control of privacy protection granularity and perfectly balancing the dual needs of caregiving functions and privacy protection.
Owner:KAIWANG (HANGZHOU) TECH CO LTD

Interspecies communication systems and programs

ActiveJP7910817B1Interspecies communicationBiological body
We provide an interspecies communication system and program that enables high-fidelity interspecies communication. [Solution] The interspecies communication system 100 includes an interface device 10 that detects information of multiple different types of modalities output by the organism and converts the detected modality information into input information, and a translation device 30 that inputs the input information into the trained model and generates translation information that the organism can receive based on the output information output from the trained model. The interface device 10 detects information of multiple different types of modalities, which include at least one of either chemical substances or olfactory information, the organism's biological information, auditory information, or visual information, either the chemical substance or the olfactory information, or the organism's biological information.
Owner:CONTRACT CO WAKU WAKU DIGITAL CONSULTING

System, device, and method for improving visual and / or auditory tracking of a presentation given by a presenter

ActiveUS12614299B2Image enhancementImage analysisAuditory trackingEngineering
A system, device, and method to improve visual and / or auditory tracking of a presentation given by a presenter, the system having a first electronic device integrating a first piece of software for obtaining information run in the device; a second electronic device integrating a second piece of software; a microphone for obtaining auditory information of the presentation; a compact module comprising a single-board computer, a router, a power supply, a fixed camera for acquiring information of the presentation shown in a support, and a moving camera for acquiring information of the presenter's position; tracking device for obtaining presenter tracking information based on the information of the position. The second piece of software is adapted for showing the information run in the first electronic device, the auditory information of the presentation, the information shown through the support, and the presenter tracking information.
Owner:BEMYVEGA SL

Vehicle operation apparatus

To provide a vehicle operation apparatus which can reduce complexity in confirmation work of a user.SOLUTION: A vehicle operation apparatus 100 comprises: an operation panel 21 for displaying a plurality of operation buttons 22; a touch detection part 30 for detecting, as a touch operation, a contact to any one of the plurality of operation buttons 22; an output unit 50 capable of outputting visual information, tactile information and auditory information; and a control unit 10 for outputting, to the output unit, the visual information, the tactile information and the auditory information as feedback information to the touch operation, in response to detection of the touch operation by the touch detection unit 30. The control unit 10 outputs, to the output unit 50, the feedback information different in at least one of the visual information, the tactile information and the auditory information for each of the operation buttons 22.SELECTED DRAWING: Figure 1
Owner:U SHIN LTD

Information processing system and information processing method

[Problem] To provide assistance to a user in a situation where visual recognition of an object is difficult. [Solution] This information processing system comprises a processing circuit that performs: processing for detecting an object on the basis of sensing data obtained by sensing the surroundings using a sensor unit provided in a device worn by a user; processing for generating auditory information on the basis of information on the detected object; and processing for performing control so as to notify the user of the generated auditory information through auditory stimulation.
Owner:SONY SEMICON SOLUTIONS CORP

Campus bullying behavior detection method and system based on video phone voice recognition

The application relates to the technical field of speech recognition, in particular to a campus bullying behavior detection method and system based on video telephone speech recognition. By acquiring sound signals and continuous image frames in a video telephone monitoring area, effective fusion of visual and auditory information is realized, so that abnormal conditions in the campus can be more comprehensively captured and analyzed. The feature points of a person in the image frame are recognized and tracked, and the position distribution between persons is combined to obtain a gathering characteristic value of the person in the monitoring area, which is of great significance to identifying potential bullying scenes. When the gathering characteristic value is abnormal, the sound signals are further analyzed, sound data segments are divided, acoustic characteristics are extracted, and an abnormal factor of each sound data segment is calculated. Finally, the abnormal factors of all sound data segments and the person gathering characteristic value are combined to obtain an abnormal characteristic value for monitoring the campus bullying behavior, so that the risk of missed reports and false reports is reduced.
Owner:GUANGDONG AOXUNDA INFORMATION TECH CO LTD

Bar interactive ordering management system based on augmented reality

The invention relates to the technical field of catering management systems, in particular to a bar interactive ordering management system based on augmented reality, comprising: A1, user front-end equipment integrated with a multi-modal data acquisition module for displaying augmented reality content and acquiring visual information, auditory information and three-dimensional space information in an environment in real time; and A2, a data preprocessing module which is connected with the multi-mode data acquisition module. Visual, auditory and three-dimensional space information is acquired through the multi-modal data acquisition module, and is subjected to deep fusion processing through the multi-modal fusion recognition module after being optimized through the data preprocessing module, so that the defect that single visual recognition is insufficient in precision in a bar weak light and shielding scene is avoided, and high-precision recognition of a target object, user gestures and environmental characteristics is realized; by means of an AI recommendation engine in combination with user historical consumption records, real-time preferences, inventory information and environmental factors, personalized drink or package recommendation is provided, and the problem of blind recommendation of a traditional system is solved.
Owner:BEIJING HOLOGRAPHIC JULANG TECH CO LTD

Electronic device for outputting answer to question of user by using artificial intelligence model and operating method thereof

According to an embodiment, this electronic device may comprise: a camera; at least one processor (220) including a processing circuit; and a memory (230) including one or more storage media and storing instructions. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to receive a question from a user while capturing images at a designated period through the camera. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to obtain a first keyword of the question on the basis of providing the question to an artificial intelligence model. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to, on the basis of identifying that the first keyword matches with at least one keyword among keywords stored in the memory, obtain location information indicating a location at which at least one image corresponding to the at least one keyword is stored in an external database. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to identify at least one first image most recently captured from a time point at which the question is received. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to obtain an answer to the question on the basis of providing the at least one first image, the location information, and the question-related information to the artificial intelligence model. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to output the answer as at least one of visual information or auditory information.
Owner:SAMSUNG ELECTRONICS CO LTD

Interactive language learning system based on semantic scene generation

The invention relates to the technical field of language learning, in particular to an interactive language learning system based on semantic scene generation, and the system comprises a concept building module which is used for building the concept of a target language in the cognition of a learner through the association of visual information and / or auditory information and a unified voice form; the note recording number learning module is used for associating a unified voice form with a corresponding note recording number on the basis of the concept construction module, and training a learner to master the conversion between the voice and the note recording number; and the character learning module is used for training the learner to master spelling rules between the note numbers and the corresponding writing forms according to the note characters. According to the interactive language learning system generated based on the semantic scene, the memory efficiency and durability of vocabularies and syntax are greatly improved, more importantly, the language intuition and direct application ability of a learner are cultivated, and the key transformation from learning of knowledge about languages to learning of the languages is achieved.
Owner:赖诚诚

Wearable electronic device for providing auditory information and method therefor

A wearable electronic device according to various embodiments disclosed in the present document may comprise an image sensor, a speaker, a memory for storing instructions, and at least one processor including processing circuitry, wherein the instructions, when executed individually or collectively by the at least one processor, cause the wearable electronic device to: identify an action event of a user; analyze an action intent on the basis of the action event; identify a target object of the action event and an attribute of the target object on the basis of the action intent; identify, in correspondence to the attribute of the target object, a body part of the user and at least a partial area of the target object corresponding to the action intent; identify spatial information on the basis of the attribute of the target object, the body part, and the at least partial area of the target object; set, on the basis of the spatial information, a path through which the body part reaches the at least partial area of the target object; and provide, via the speaker, sound information including a first sound and a second sound for guiding the execution of the action event on the basis of the path.
Owner:SAMSUNG ELECTRONICS CO LTD

Driver fatigue monitoring method based on multi-modal fusion

The invention discloses a driver fatigue monitoring method based on multi-modal fusion, and belongs to the field of computer vision and artificial intelligence. The method comprises the following steps: constructing a visual perception assembly line, and extracting visual physiological features of a driver from a video frame; constructing an auditory perception assembly line, and extracting a fatigue-related acoustic event from the audio stream; a multi-modal fusion decision engine is constructed, and a final fatigue driving early warning decision is generated by adopting a double-track parallel mechanism according to the visual physiological features and the acoustic events; and the deployment performance is optimized by adopting a model lightweight technology. According to the invention, visual and auditory information is fused, and dual-channel complementation is utilized, so that the detection robustness in a complex environment with weak light, high noise and face shielding is improved; and meanwhile, a lightweight model is selected and combined with INT8 integer quantization optimization, so that the problem that a high-precision model is difficult to deploy on a low-computing-power vehicle-mounted platform is effectively solved, and the balance of high precision and high efficiency is realized.
Owner:BEIHANG UNIV

Evaluation device, evaluation method, and computer program

ActiveJP2026137513AEngineeringDevice Evaluation Method
Improve the accuracy of video ad evaluation. [Solution] The evaluation device for evaluating video advertisements comprises an acquisition unit and an evaluation unit. The acquisition unit is configured to acquire feature quantities relating to one or more expressive elements contained in the video advertisement. The evaluation unit is configured to input the feature quantities into a trained machine learning model and acquire evaluation values ​​for the video advertisement output by the machine learning model using one or more evaluation indicators based on the feature quantities from the machine learning model. One or more expressive elements include at least one of visual information and auditory information. One or more evaluation indicators include one or more indicators relating to advertising effectiveness.
Owner:HAKUHODO INC

System

An object of the system according to the embodiment is to convert visual information into sound and provide the sound to a visually impaired person or a bedridden person.SOLUTION: A system according to an embodiment includes a visual information analysis unit, a decoder, a voice conversion unit, and an integration unit. The visual information analysis unit analyzes visual information of the video. The decoder decodes the visual information analyzed by the visual information analyzer. The voice conversion unit converts the visual information verbalized by the verbalization unit into voice. The integration unit integrates the visual information and the audio information converted into the voice by the voice conversion unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

system

The system according to the embodiment aims to transmit visual and auditory information directly to the brains of people with visual or hearing impairments. [Solution] A system according to an embodiment includes a collection unit, an analysis unit, a generation unit, and a transmission unit. The collection unit collects video and audio data. The analysis unit analyzes the data collected by the collection unit and learns how to generate electrical signals. The generation unit converts the video and audio data into electrical signals based on the data learned by the analysis unit. The transmission unit sends the electrical signals generated by the generation unit to the brain.
Owner:SOFTBANK GROUP CORP

system

Provide a system. 【Solution means】 Means for analyzing the user's learning history to estimate the area of interest, Means for selecting and displaying appropriate learning materials based on the estimated area of interest, Means for analyzing voice commands to operate learning content, Means for promoting learning in a game-like manner by means of a point system and a reward system, Means for managing the learning schedule and giving reminders at set times, Means for providing the learning progress data to the guardian, Means for providing feedback to the user using visual information and auditory information to achieve two-way learning, Means for providing customized quizzes and information based on the user's interests, Means for providing virtual rewards when a specific point is reached, Means for generating prompts based on the learning theme A system including.
Owner:SOFTBANK GROUP CORP

Audio-visual saliency prediction method based on implicit neural representation

The invention relates to an implicit neural representation-based audio-visual saliency prediction method, which comprises the following steps of: acquiring data containing audios and videos, performing pairwise processing on the data, taking video frames of adjacent time as a group of data, expressing the time as a time interval of two frames, constructing a PAVS model and performing training, and obtaining an audio-visual saliency prediction result. Two-dimensional discrete image coordinates, time dynamic features and corresponding weights are generated through a dynamic perception generator module; in the parameterized audio-visual feature fusion module, discrete image coordinates and time dynamic features are mapped, the average absolute error loss between a predicted saliency map and a real saliency map is calculated, the optimal training result weight is updated, and the optimal weight is used to predict a test set. The problems of weak audio-visual interactivity, low efficiency and poor performance of an audio-visual saliency prediction task are solved, the method can effectively realize internal interaction of visual and auditory information streams, and audio-visual features are adaptively fused to obtain excellent performance.
Owner:DONGHUA UNIV +2

An Active Sound-Based Anti-Drowsiness System and Method for Electric Vehicles Based on Multimodal Information Fusion

PendingCN122078413AOvercome hysteresisovercome inaccuraciesElectric/fluid circuitMotion sicknessEngineering
This invention discloses an active acoustic motion sickness suppression system and method for electric vehicles based on multimodal information fusion. The system includes a signal acquisition module, a data processing and motion sickness level recognition module, an acoustic signal generation module, and an onboard sound device. The signal acquisition module collects vehicle motion parameters and occupant physiological state data. The data processing and motion sickness level recognition module fuses and analyzes the raw EEG signals and vehicle motion parameters, outputting a motion sickness level judgment result. The acoustic signal generation module uses an order synthesis algorithm to enhance the corresponding parameters of specific order harmonics, amplifying the synthesized audio signal as an acoustic intervention signal. The onboard sound device receives the acoustic intervention signal and plays the intervention sound. This invention achieves accurate judgment of motion sickness and can dynamically generate highly correlated acoustic intervention signals based on the unique vehicle motion parameters of electric vehicles, compensating for auditory information gaps and achieving a motion sickness suppression effect.
Owner:WUHAN UNIV OF TECH