Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

43 results about "Auditory information" patented technology

Auditory memory is the ability to process information presented orally, analyze it mentally, and store it to be recalled later. Those with a strong capacity for this type of memory are called auditory learners. The ability to learn from oral instructions and explanations is a fundamental skill required throughout life.

Multi-mode sensing intelligent microphone array signal processing method and system

The invention provides a multi-mode sensing intelligent microphone array signal processing method and system, and belongs to the technical field of signal processing, and the method comprises the steps: obtaining a visual signal through employing a multi-directional visual sensor, and obtaining a sound signal through employing a microphone array; extracting visual features and acoustic features; constructing an audio-visual topological feature space, mapping visual features and acoustic features to the space, and establishing a sound source probability distribution model; processing the sound signal by adopting a multi-dimensional discriminant adversarial generative network, and separating out a target voice signal; the acoustic environment state is evaluated in real time, and processing parameters are dynamically adjusted; quality evaluation is carried out on the separated multiple paths of target voice signals, the voice signal with the highest quality is selected as output, audio-visual multi-mode information is deeply fused and cooperatively processed, and a topology enhanced adversarial generative network architecture and an environment self-adaptive mechanism are combined, so that the voice separation effect in a complex environment is remarkably improved; and more than 85% of speech intelligibility can still be kept in a scene that six persons speak simultaneously.
Owner:GUANGZHOU OPSMEN TECH CO LTD

Behavior data analysis method and system based on intelligent perception

PendingCN121071570ABiological modelsKnowledge based modelsChild behaviourEngineering
The invention discloses a behavior data analysis method and system based on intelligent perception. The method and system are used for improving the reliability of child behavior analysis. The method comprises the following steps: acquiring multi-dimensional monitoring data of a plurality of children in a monitoring environment, wherein the multi-dimensional monitoring data comprises visual information, auditory information and physiological information of each child; performing multi-source fusion on the visual information, the auditory information and the physiological information of each child to generate a behavior feature set; inputting the behavior feature set into a preset behavior and emotion classification model for classification to obtain a behavior mode and an emotion state of each child; on the basis of individual information of different children, the behavior modes and the emotional states in combination with scene information of the monitoring environment, knowledge graph construction on a time sequence is carried out, a children behavior graph is obtained, and the children behavior graph is used for displaying behavior events of the children in a preset time period.
Owner:SHEN ZHEN GUANG QI JI SUAN KE JI YOU XIAN GONG SI

Communication robot, communication robot control method, and program

A communication robot includes an auditory information processing portion configured to recognize a volume of voice collected by a sound collection portion and generate an auditory attention map by projecting a sound position in a three-dimensional space onto a two-dimensional attention map in which the robot is located at a center, a visual information processing portion configured to generate a visual attention map using a face detection result obtained by detecting a face of a person from an image captured by an imaging portion and a motion detection result obtained by detecting a motion of the person, an attention map generation portion configured to generate an attention map by integrating the auditory attention map and the visual attention map, and a motion processing portion configured to control eyeball movements and motions of the communication robot using the attention map.
Owner:HONDA MOTOR CO LTD

Multimodal perception based smart microphone array signal processing method and system

The application provides a multi-modal perception intelligent microphone array signal processing method and system, and belongs to the technical field of signal processing, which comprises the following steps: acquiring visual signals by using multi-directional visual sensors and acquiring sound signals by using microphone arrays; extracting visual features and acoustic features; constructing a visual-auditory topological feature space, mapping the visual features and acoustic features to the space, and establishing a sound source probability distribution model; processing the sound signals by using a multi-dimensional discriminative generative adversarial network to separate target speech signals; real-time evaluating acoustic environment states, dynamically adjusting processing parameters; evaluating the quality of the separated multi-path target speech signals, selecting the highest quality speech signal as the output, and deeply fusing and cooperatively processing multi-modal visual-auditory information, combining a topological enhanced generative adversarial network architecture and an environment adaptive mechanism, and significantly improving the speech separation effect in a complex environment, which can still maintain a speech intelligibility of more than 85% in a 6-person simultaneous speaking scene.
Owner:GUANGZHOU OPSMEN TECH CO LTD

Intelligent audio-visual auxiliary glasses, audio-visual auxiliary method, electronic equipment and storage medium

The invention relates to the field of audio-visual conversion, and provides intelligent audio-visual auxiliary glasses, an audio-visual auxiliary method, electronic equipment and a storage medium, and the intelligent audio-visual auxiliary glasses comprise a glasses frame which comprises a microphone and lenses; the character conversion unit is used for receiving external sound information through the microphone, converting the sound information into characters and displaying the characters on the lens; and the voice conversion unit is used for receiving external visual information through the visual sensor, converting the visual information into voice and playing the voice. The invention provides an integrated audio-visual conversion device, which is used for solving the defect of lack of audio-visual conversion devices for visually impaired people and hearing impaired people in the prior art, can help the visually impaired people to obtain front visual information and can also help the hearing impaired people to obtain sound information, so that the audio-visual conversion effect is improved. And bidirectional enhancement of visual information and auditory information is realized.
Owner:CHONGQING UNIV

Autonomous collaborative decision-making method, system and medium for industrial robots based on multimodal perception

The present application provides an autonomous collaborative decision-making method, system and medium for industrial robots based on multimodal perception, which belongs to the field of intelligent control technology for industrial robots. The method includes: collecting visual, tactile and auditory information, performing cross-modal spatiotemporal alignment to eliminate spatiotemporal differences in data, combining a dynamic weight adjustment mechanism and an attention calculation model to achieve multimodal feature fusion, and generating a joint decision-making strategy through a deep learning optimization model. The system dynamically allocates sensor weights according to the task type, and introduces a multi-objective optimization mechanism for energy consumption, accuracy and safety, while setting fault-tolerant rules to automatically restore weights or recalibrate sensors. The present invention solves the problems of rigid data fusion and single optimization dimension in traditional multimodal decision-making, and significantly improves the accuracy, response speed and environmental adaptability of collaborative operations of industrial robots.
Owner:SHENZHEN HUAZHONG NUMERICAL CONTROL

system

We provide the system. [Solution] A device for analyzing visual and auditory information collected in the work environment, A device that identifies tasks that can be automated based on the aforementioned analysis, A device that proposes an optimal automation method for the aforementioned task, A system that includes this.
Owner:SOFTBANK GROUP CORP

Communication method and system of AR glasses with earphone and glasses separated, and medium

The invention provides a communication method and system for AR glasses with separated earphones and glasses, and a medium, and the method comprises the steps: initializing the AR glasses, and obtaining the initialization state information; scanning compatible state information between the AR glasses and the separated audio equipment based on the initialized state information; establishing communication state information between the AR glasses and the separated audio equipment based on the compatible state information; performing communication matching authentication based on the communication state information between the AR glasses and the separated audio equipment to obtain an authentication result; analyzing whether the communication is successful or not based on the authentication result, if the communication is successful, outputting audio information and instruction response, and if the communication is failed, re-pairing the separated audio equipment; communication matching is completed by analyzing the communication state of the AR glasses and the separated audio equipment, so that successful communication is realized, cooperative transmission of visual information at a glasses end and auditory information at an earphone end is realized, and synchronism of audio and pictures is ensured.
Owner:SHENZHEN ORANGE ELECTRONICS CO LTD

System

An object of a system according to an embodiment is to perform information presentation utilizing multiple senses using generated AI.SOLUTION: A system includes a visual information presentation part, an auditory information presentation part, a tactile information presentation part, a taste information presentation part, and an olfactory information presentation part. The visual information presenting unit presents visual information by using the generated AI. The audio information presenting unit presents audio information by using the generated AI. The tactile information presenting unit presents tactile information by using the generated AI. The taste information presenter presents taste information by using the generated AI. The olfactory information presenting unit presents olfactory information by using the generated AI.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Control method of teaching device, control device, teaching system, and storage medium

The application discloses a teaching equipment control method, a control device, a teaching system and a storage medium. The teaching equipment control method comprises the following steps: image and audio of a target in a teaching space are collected to obtain image data and audio data of the target, wherein the teaching space comprises the teaching equipment; visual information of the target is extracted by using the image data of the target, and auditory information of the target is extracted by using the audio data of the target; and the teaching equipment is controlled based on the visual information and the auditory information of the target. In the foregoing manner, the convenience of the teaching equipment control can be improved, and the accuracy is relatively high.
Owner:IFLYTEK LINGZHI (JIANGSU) TECH CO LTD +1

system

We provide the system. [Solution] A device that performs information processing to generate an individually optimized learning plan based on the learner's abilities, progress, and goals, A device that provides educational materials based on a generated learning plan, A device that analyzes learners' activities with educational materials in real time and generates evaluation results, A device that analyzes a learner's past learning history to identify learning challenges and provides additional exercises to improve those challenges, A sensor that acquires sound and video, and a device that uses the acquired data to perform interactive speech practice in real time, A device that provides immediate instruction using visual and auditory information based on speech evaluation data, A system that includes this.
Owner:SOFTBANK GROUP CORP

Robot control system, robot control program

PendingCN122319059ALittle fingerEngineering
The hand tool 50 is equipped with sensor groups with different information acquisition functions on each of the fingers 22A, 22B, 22C, 22D, and 22E. In other words, it is equipped with sensor groups for acquiring various information required for multimodal robot control. A visual sensor 51A for detecting visual information is installed on the thumb, an auditory sensor 51B for detecting auditory information is installed on the index finger, an olfactory sensor 51C is installed on the middle finger, a tactile sensor 51D for detecting tactile information is installed on the ring finger, and a taste sensor 51E for detecting taste information is installed on the little finger.
Owner:SOFTBANK GROUP CORP

Vehicle operation device

This vehicle operation device (100) comprises: an operation panel (21) on which a plurality of operation buttons (22) are displayed; a touch detection unit (30) which detects, as a touch operation, contact with any of the plurality of operation buttons (22); an output unit (50) which can output visual information, tactile information, and auditory information; and a control unit (10) which, in response to the detection of the touch operation by the touch detection unit (30), causes the output unit to output the visual information, tactile information, and auditory information as feedback information for the touch operation. The control unit (10) causes the output unit (50) to output, for each of the operation buttons (22), feedback information in which at least one of the visual information, tactile information, and auditory information is different.
Owner:U SHIN LTD

Artificial intelligence-based english spoken language pronunciation correction system

PendingCN122392380AVisually impairedMemory retention
The application belongs to the field of intelligent teaching, and particularly relates to an English oral pronunciation correction system based on artificial intelligence, which comprises a multi-modal perception module, a deep pronunciation feature extraction module, a cognitive compensation fusion module and an intelligent correction and training planning module; the abstract pronunciation is converted into a tactile experience through a multi-modal learning process, the tactile information and the auditory information are synchronously input, the memory retention rate of the visually impaired students is improved, and the students are guided to deeply remember the acoustic characteristics of the pronunciation; the application simulates the auditory-tactile joint representation ability of the visually impaired brain through the learnable attention weight and the cognitive compensation prior mask, realizes the real sense of "touch instead of vision", simultaneously identifies the error reasons through multi-task learning, provides scientific basis for targeted training, designs a special tactile generation network, maps the acoustic characteristics and the pronunciation error types into tactile patterns, provides instant tactile correction feedback for the students, and forms a closed-loop multi-modal learning experience.
Owner:SHANDONG VOCATIONAL COLLEGE OF SPECIAL EDUCATION

Cow artificial insemination simulation method and device

The invention relates to the technical field of animal husbandry breeding, and discloses a cow artificial insemination analog simulation method and device, and the method comprises the following steps: a computer terminal loads a preset standardized operation parameter set which is used for defining objective training standards; collecting multi-modal data generated by multiple types of sensors arranged on the simulation cattle model and the operation instrument in real time, and analyzing the multi-modal data into a real-time state vector representing the current operation; dynamically evaluating the real-time state vector based on a loaded standardized operation parameter set and preset collaborative logic to generate a structured evaluation result; and driving the computer terminal to provide real-time multimedia interaction feedback containing visual and auditory information to the operator according to the generated evaluation result. Through multi-modal sensing, collaborative logic and a quantitative evaluation model, objective evaluation and dynamic physiological feedback of operation are realized, and the trueness and standardization level of training are comprehensively improved.
Owner:NANJING LAIYITE ELECTRONIC TECH CO LTD

system

We provide the system. [Solution] Features that acquire the learner's visual and auditory information in real time, The function analyzes the data obtained by the aforementioned function and identifies the mental state related to the learner's level of understanding, A function that generates individualized instruction and practice problems based on the mental state identified by the aforementioned function, The function presents the instruction and problems generated by the aforementioned function to the learner, The machine acts as a learning supporter within the home, providing optimal learning materials according to the learner's progress, A system that includes this.
Owner:SOFTBANK GROUP CORP

AI vision camera (HUSKYLENS)

1. Name of the product in this design: AI Vision Camera (HUSKYLENS). 2. Purpose of this design: This product is used to collect, identify and process visual and auditory information, and output the results. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design's key points: a 3D model.
Owner:SHANGHAI ZHIWEI ROBOT CO LTD

system

The system according to this embodiment aims to enable visually impaired or bedridden people to enjoy video content. [Solution] The system according to the embodiment comprises an acquisition unit, an analysis unit, a language conversion unit, a speech conversion unit, and a provision unit. The acquisition unit acquires auditory information from video content. The analysis unit analyzes the visual information contained in the video based on the auditory information acquired by the acquisition unit. The language conversion unit translates the visual information analyzed by the analysis unit into language. The speech conversion unit converts the visual information translated by the language conversion unit into speech. The provision unit provides audio content by combining the speech converted by the speech conversion unit with the auditory information.
Owner:SOFTBANK GROUP CORP

A multi-modal fusion video classification method and system based on brain-like feedback interaction

The application discloses a kind of multi-modal fusion video classification method and system based on brain-like feedback interaction, method includes: video pre-processing, obtain visual information and auditory information in video;After feature extraction, input into the multi-modal fusion framework based on neural network architecture search to obtain fusion information representation after fusion;Fusion information representation, visual information and auditory information are input into feedback module based on feedback modulation effect generated by single sensory cortex after auditory and visual information integration of human brain in superior temporal sulcus, and the output multi-modal fusion visual information and multi-modal fusion auditory information are respectively through fully connected layer to obtain the confidence of each classification;The confidence of each classification is input into DS decision fusion module, and the final classification result is obtained.The application learns from the processing mode of human brain to perceive external environment, integrates multiple senses from various modal perception information, and effectively improves the accuracy of recognizing and classifying the expressions of people in the video.
Owner:HOHAI UNIV

Speedometers that utilize touch and hearing

PendingJP2026112346ARoad vehicles traffic controlAudible meter reading indicationAcousticsTesting Methods
The present invention provides a device that divides the vehicle's speed into multiple stepped speed ranges and outputs different tactile or auditory information according to the number of steps in each speed range. [Solution] The number of speed range steps in the tactile speedometer is changed from the usual decimal number to a lower base n number, and the number of different values ​​used by each digit of the new base n number is reduced, thereby reducing the total number of numerical patterns used by the tactile speedometer to display numerical values, making it easier to understand the number of speed range steps via touch. Furthermore, it is possible to change the type of base n number used for each speed range step, so that the speed range indicated by the first digit of each base n number is shown in the same speed range unit across the entire speed range, making it possible to unify the speed range of the first digit. This enables the speedometer to display speed in a unified speed range unit.
Owner:大庭 有二

Multi-modal collaborative perception based multi-level privacy protection method, device and equipment for care scene

This invention provides a multi-level privacy protection method, device, and equipment for caregiving scenarios based on multimodal collaborative perception. By constructing a multimodal neural network model at the caregiving data acquisition terminal and fusing visual and auditory information for deep semantic analysis, it effectively overcomes the limitations of traditional single-modal perception in complex environments such as low light and occlusion. This invention employs a collaborative design of delay buffering and pre-judgment, forcibly placing privacy level determination before media stream distribution, fundamentally eliminating the risk of instantaneous leakage of privacy images due to inference delays. By introducing a two-level dynamic judgment strategy of prioritizing high-risk feature blocking and scene adaptive fusion, the robustness of the system in real-world home environments is improved. Based on the matching relationship between user permissions and privacy levels, this invention performs differentiated media stream processing on the terminal side, achieving fine-grained control of privacy protection granularity and perfectly balancing the dual needs of caregiving functions and privacy protection.
Owner:KAIWANG (HANGZHOU) TECH CO LTD

Interspecies communication systems and programs

ActiveJP7910817B1Interspecies communicationBiological body
We provide an interspecies communication system and program that enables high-fidelity interspecies communication. [Solution] The interspecies communication system 100 includes an interface device 10 that detects information of multiple different types of modalities output by the organism and converts the detected modality information into input information, and a translation device 30 that inputs the input information into the trained model and generates translation information that the organism can receive based on the output information output from the trained model. The interface device 10 detects information of multiple different types of modalities, which include at least one of either chemical substances or olfactory information, the organism's biological information, auditory information, or visual information, either the chemical substance or the olfactory information, or the organism's biological information.
Owner:CONTRACT CO WAKU WAKU DIGITAL CONSULTING

System, device, and method for improving visual and / or auditory tracking of a presentation given by a presenter

ActiveUS12614299B2Image enhancementImage analysisAuditory trackingEngineering
A system, device, and method to improve visual and / or auditory tracking of a presentation given by a presenter, the system having a first electronic device integrating a first piece of software for obtaining information run in the device; a second electronic device integrating a second piece of software; a microphone for obtaining auditory information of the presentation; a compact module comprising a single-board computer, a router, a power supply, a fixed camera for acquiring information of the presentation shown in a support, and a moving camera for acquiring information of the presenter's position; tracking device for obtaining presenter tracking information based on the information of the position. The second piece of software is adapted for showing the information run in the first electronic device, the auditory information of the presentation, the information shown through the support, and the presenter tracking information.
Owner:BEMYVEGA SL

Vehicle operation apparatus

To provide a vehicle operation apparatus which can reduce complexity in confirmation work of a user.SOLUTION: A vehicle operation apparatus 100 comprises: an operation panel 21 for displaying a plurality of operation buttons 22; a touch detection part 30 for detecting, as a touch operation, a contact to any one of the plurality of operation buttons 22; an output unit 50 capable of outputting visual information, tactile information and auditory information; and a control unit 10 for outputting, to the output unit, the visual information, the tactile information and the auditory information as feedback information to the touch operation, in response to detection of the touch operation by the touch detection unit 30. The control unit 10 outputs, to the output unit 50, the feedback information different in at least one of the visual information, the tactile information and the auditory information for each of the operation buttons 22.SELECTED DRAWING: Figure 1
Owner:U SHIN LTD

Information processing system and information processing method

[Problem] To provide assistance to a user in a situation where visual recognition of an object is difficult. [Solution] This information processing system comprises a processing circuit that performs: processing for detecting an object on the basis of sensing data obtained by sensing the surroundings using a sensor unit provided in a device worn by a user; processing for generating auditory information on the basis of information on the detected object; and processing for performing control so as to notify the user of the generated auditory information through auditory stimulation.
Owner:SONY SEMICON SOLUTIONS CORP

Auditory operation training method for hearing-impaired children

The invention relates to an auditory operation training method for hearing-impaired children, and aims to solve the problem of hearing matching and difficulty in auditory understanding caused by insufficient high-order cognitive skills. Existing training mainly focuses on sound identification and word and sentence understanding, and lacks complex discourse reasoning multi-step information integration and processing process training. According to the method, an auditory operation stage is added on four stages of auditory development of Erber. The training method comprises the following steps: S1, visual understanding: designing a dual story, and enabling students to watch a teacher operation desktop doll, answer facts and infer questions; s2, multi-modal memory: students listen to teacher instructions, operate a desktop doll, and convert auditory information into multi-modal memory; s3, semi-open auditory operation: the students answer rational questions under the condition of no visual clues; and S4, opening auditory operation: presenting a new discourse material, and completing an rational problem through auditory single-channel information input by a student without the steps S1 to S3. According to the method, support type auditory operation training is provided for the hearing-impaired children.
Owner:陈清怡

Campus bullying behavior detection method and system based on video phone voice recognition

The application relates to the technical field of speech recognition, in particular to a campus bullying behavior detection method and system based on video telephone speech recognition. By acquiring sound signals and continuous image frames in a video telephone monitoring area, effective fusion of visual and auditory information is realized, so that abnormal conditions in the campus can be more comprehensively captured and analyzed. The feature points of a person in the image frame are recognized and tracked, and the position distribution between persons is combined to obtain a gathering characteristic value of the person in the monitoring area, which is of great significance to identifying potential bullying scenes. When the gathering characteristic value is abnormal, the sound signals are further analyzed, sound data segments are divided, acoustic characteristics are extracted, and an abnormal factor of each sound data segment is calculated. Finally, the abnormal factors of all sound data segments and the person gathering characteristic value are combined to obtain an abnormal characteristic value for monitoring the campus bullying behavior, so that the risk of missed reports and false reports is reduced.
Owner:GUANGDONG AOXUNDA INFORMATION TECH CO LTD

Bar interactive ordering management system based on augmented reality

The invention relates to the technical field of catering management systems, in particular to a bar interactive ordering management system based on augmented reality, comprising: A1, user front-end equipment integrated with a multi-modal data acquisition module for displaying augmented reality content and acquiring visual information, auditory information and three-dimensional space information in an environment in real time; and A2, a data preprocessing module which is connected with the multi-mode data acquisition module. Visual, auditory and three-dimensional space information is acquired through the multi-modal data acquisition module, and is subjected to deep fusion processing through the multi-modal fusion recognition module after being optimized through the data preprocessing module, so that the defect that single visual recognition is insufficient in precision in a bar weak light and shielding scene is avoided, and high-precision recognition of a target object, user gestures and environmental characteristics is realized; by means of an AI recommendation engine in combination with user historical consumption records, real-time preferences, inventory information and environmental factors, personalized drink or package recommendation is provided, and the problem of blind recommendation of a traditional system is solved.
Owner:BEIJING HOLOGRAPHIC JULANG TECH CO LTD

Electronic device for outputting answer to question of user by using artificial intelligence model and operating method thereof

According to an embodiment, this electronic device may comprise: a camera; at least one processor (220) including a processing circuit; and a memory (230) including one or more storage media and storing instructions. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to receive a question from a user while capturing images at a designated period through the camera. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to obtain a first keyword of the question on the basis of providing the question to an artificial intelligence model. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to, on the basis of identifying that the first keyword matches with at least one keyword among keywords stored in the memory, obtain location information indicating a location at which at least one image corresponding to the at least one keyword is stored in an external database. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to identify at least one first image most recently captured from a time point at which the question is received. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to obtain an answer to the question on the basis of providing the at least one first image, the location information, and the question-related information to the artificial intelligence model. According to an embodiment, the instructions, when collectively or individually executed by the at least one processor, may cause the electronic device to output the answer as at least one of visual information or auditory information.
Owner:SAMSUNG ELECTRONICS CO LTD