Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

64 results about "Auditory information" patented technology

Auditory memory is the ability to process information presented orally, analyze it mentally, and store it to be recalled later. Those with a strong capacity for this type of memory are called auditory learners. The ability to learn from oral instructions and explanations is a fundamental skill required throughout life.

Industrial robot autonomous collaborative decision-making method and system based on multi-modal perception and medium

The invention provides an industrial robot autonomous collaborative decision-making method and system based on multi-modal sensing and a medium, and belongs to the technical field of industrial robot intelligent control. The method comprises the steps of performing cross-modal space-time alignment to eliminate data space-time differences by collecting visual, tactile and auditory information, realizing multi-modal feature fusion in combination with a dynamic weight adjustment mechanism and an attention calculation model, and generating a joint decision strategy through a deep learning optimization model. The system dynamically allocates sensor weights according to task types, introduces a multi-objective optimization mechanism of energy consumption, precision and safety, and sets a fault-tolerant rule to automatically recover the weights or recalibrate the sensors. According to the method, the problems of rigid data fusion and single optimization dimension in traditional multi-modal decision making are solved, and the precision, the response speed and the environmental adaptability of collaborative operation of the industrial robot are remarkably improved.
Owner:SHENZHEN HUAZHONG NUMERICAL CONTROL

Multi-mode sensing intelligent microphone array signal processing method and system

The invention provides a multi-mode sensing intelligent microphone array signal processing method and system, and belongs to the technical field of signal processing, and the method comprises the steps: obtaining a visual signal through employing a multi-directional visual sensor, and obtaining a sound signal through employing a microphone array; extracting visual features and acoustic features; constructing an audio-visual topological feature space, mapping visual features and acoustic features to the space, and establishing a sound source probability distribution model; processing the sound signal by adopting a multi-dimensional discriminant adversarial generative network, and separating out a target voice signal; the acoustic environment state is evaluated in real time, and processing parameters are dynamically adjusted; quality evaluation is carried out on the separated multiple paths of target voice signals, the voice signal with the highest quality is selected as output, audio-visual multi-mode information is deeply fused and cooperatively processed, and a topology enhanced adversarial generative network architecture and an environment self-adaptive mechanism are combined, so that the voice separation effect in a complex environment is remarkably improved; and more than 85% of speech intelligibility can still be kept in a scene that six persons speak simultaneously.
Owner:GUANGZHOU OPSMEN TECH CO LTD

Adaptive noise reduction method for positive pressure type air breathing machine based on airflow frequency spectrum characteristics

The invention discloses a self-adaptive noise reduction method for a positive pressure type air respirator based on airflow spectrum characteristics, and relates to the technical field of positive pressure type air respirators.The self-adaptive noise reduction method comprises the steps that a respirator air path pressure signal and an environment noise signal are synchronously collected through an airflow sensor and a microphone array; performing time-frequency analysis on the collected signals, and extracting airflow spectrum features; and establishing a mapping relation between the working state of the breathing machine and the noise spectrum based on the historical record according to the mapping relation between the working state of the breathing machine and the noise spectrum. Through real-time acquisition and analysis of airflow frequency spectrum characteristics, accurate identification of useful airflow signals and environmental noise signals, dynamic adjustment of noise reduction parameters, effective suppression of high-frequency noise and low-frequency vibration, and improvement of noise reduction effect, compared with a traditional fixed frequency band filtering or passive sound insulation material, the method can better cope with a non-stationary noise environment, and has good application prospects. The noise reduction strategy is adjusted in real time, clear auditory information is provided in a complex noise environment, and therefore the auditory perception ability of firemen is improved.
Owner:NINGBO JIUYUN HUASHENG TECHNOLOGY CO LTD

Vehicle automatic driving auditory information semantic understanding learning method and device

The invention discloses a vehicle automatic driving auditory information semantic understanding learning method and device, and the method comprises the steps: collecting sound events in a virtual environment through a virtual audio collection module, obtaining a virtual audio data set, classifying the virtual audio data in the virtual audio data set, and generating a pseudo tag; acquiring scene information of the target virtual audio data, and taking the scene information as a scenarized semantic tag of the target virtual audio data; the virtual audio data set with the label is used for deep learning model training to obtain an audio analysis model; acquiring real-time audio data, and extracting audio features of the real-time audio data through a feature extraction layer of the audio analysis model; acquiring a context feature corresponding to the audio feature in the time sequence, and generating an audio time feature according to the target audio feature and the context feature; audio time features are analyzed through an output layer of the audio analysis model to obtain classification and scenes corresponding to real-time audio data, and the perception and understanding ability of the automatic driving vehicle to auditory information is improved.
Owner:FUJIAN UNIV OF TECH

Control system for robot and control program for robot

To enable installing, at appropriate positions, respective sensors for acquiring various kinds of information necessary for controlling a robot in a multimodal manner; and to acquire information from the respective sensors in an appropriate state and at appropriate timing.SOLUTION: In a hand tool 50, sensors having respective different information acquisition functions are attached to respective finger parts 22A, 22B, 22C, 22D, and 22E. In other words, sensors for acquiring various kinds of information necessary for controlling a robot are attached. A visual sensor 51A for detecting visual information is attached to a thumb, an auditory sensor 51B for detecting auditory information is attached to an index finger, an olfactory sensor 51C is attached to a middle finger, a tactile sensor 51D for detecting tactile information is attached to a ring finger, and a taste sensor 51E for detecting taste information is attached to a little finger.SELECTED DRAWING: Figure 4
Owner:SOFTBANK GROUP CORP

Behavior data analysis method and system based on intelligent perception

PendingCN121071570ABiological modelsKnowledge based modelsChild behaviourEngineering
The invention discloses a behavior data analysis method and system based on intelligent perception. The method and system are used for improving the reliability of child behavior analysis. The method comprises the following steps: acquiring multi-dimensional monitoring data of a plurality of children in a monitoring environment, wherein the multi-dimensional monitoring data comprises visual information, auditory information and physiological information of each child; performing multi-source fusion on the visual information, the auditory information and the physiological information of each child to generate a behavior feature set; inputting the behavior feature set into a preset behavior and emotion classification model for classification to obtain a behavior mode and an emotion state of each child; on the basis of individual information of different children, the behavior modes and the emotional states in combination with scene information of the monitoring environment, knowledge graph construction on a time sequence is carried out, a children behavior graph is obtained, and the children behavior graph is used for displaying behavior events of the children in a preset time period.
Owner:SHEN ZHEN GUANG QI JI SUAN KE JI YOU XIAN GONG SI

Communication robot, communication robot control method, and program

A communication robot includes an auditory information processing portion configured to recognize a volume of voice collected by a sound collection portion and generate an auditory attention map by projecting a sound position in a three-dimensional space onto a two-dimensional attention map in which the robot is located at a center, a visual information processing portion configured to generate a visual attention map using a face detection result obtained by detecting a face of a person from an image captured by an imaging portion and a motion detection result obtained by detecting a motion of the person, an attention map generation portion configured to generate an attention map by integrating the auditory attention map and the visual attention map, and a motion processing portion configured to control eyeball movements and motions of the communication robot using the attention map.
Owner:HONDA MOTOR CO LTD

A method and device for fusing visual and auditory information in vehicle autonomous driving

The present invention discloses a method and device for fusing visual and auditory information in vehicle autonomous driving, including: acquiring visual information and auditory information collected by the vehicle; the auditory information includes the sound source direction; projecting the auditory information into the coordinate system of the visual information; fusing the visual information and the auditory information in the same coordinate system through an attention mechanism to obtain a fused feature; projecting the auditory information into the visual coordinate system of the visual information includes: acquiring the internal parameter matrix of the vehicle camera; obtaining the distance value from the vehicle camera to the sound source according to the sound source direction; converting the sound source direction into the coordinates in the visual coordinate system according to the internal parameter matrix and the distance value. After acquiring the visual information and the auditory information of the vehicle, by projecting the auditory information into the coordinate system of the visual information, the auditory information and the visual information can be effectively fused in the same coordinate system, so as to improve the perception accuracy and robustness of the vehicle for complex scenarios through the fused feature for target detection and positioning.
Owner:FUJIAN UNIV OF TECH

Multimodal perception based smart microphone array signal processing method and system

The application provides a multi-modal perception intelligent microphone array signal processing method and system, and belongs to the technical field of signal processing, which comprises the following steps: acquiring visual signals by using multi-directional visual sensors and acquiring sound signals by using microphone arrays; extracting visual features and acoustic features; constructing a visual-auditory topological feature space, mapping the visual features and acoustic features to the space, and establishing a sound source probability distribution model; processing the sound signals by using a multi-dimensional discriminative generative adversarial network to separate target speech signals; real-time evaluating acoustic environment states, dynamically adjusting processing parameters; evaluating the quality of the separated multi-path target speech signals, selecting the highest quality speech signal as the output, and deeply fusing and cooperatively processing multi-modal visual-auditory information, combining a topological enhanced generative adversarial network architecture and an environment adaptive mechanism, and significantly improving the speech separation effect in a complex environment, which can still maintain a speech intelligibility of more than 85% in a 6-person simultaneous speaking scene.
Owner:GUANGZHOU OPSMEN TECH CO LTD

Intelligent audio-visual auxiliary glasses, audio-visual auxiliary method, electronic equipment and storage medium

The invention relates to the field of audio-visual conversion, and provides intelligent audio-visual auxiliary glasses, an audio-visual auxiliary method, electronic equipment and a storage medium, and the intelligent audio-visual auxiliary glasses comprise a glasses frame which comprises a microphone and lenses; the character conversion unit is used for receiving external sound information through the microphone, converting the sound information into characters and displaying the characters on the lens; and the voice conversion unit is used for receiving external visual information through the visual sensor, converting the visual information into voice and playing the voice. The invention provides an integrated audio-visual conversion device, which is used for solving the defect of lack of audio-visual conversion devices for visually impaired people and hearing impaired people in the prior art, can help the visually impaired people to obtain front visual information and can also help the hearing impaired people to obtain sound information, so that the audio-visual conversion effect is improved. And bidirectional enhancement of visual information and auditory information is realized.
Owner:CHONGQING UNIV

Autonomous collaborative decision-making method, system and medium for industrial robots based on multimodal perception

The present application provides an autonomous collaborative decision-making method, system and medium for industrial robots based on multimodal perception, which belongs to the field of intelligent control technology for industrial robots. The method includes: collecting visual, tactile and auditory information, performing cross-modal spatiotemporal alignment to eliminate spatiotemporal differences in data, combining a dynamic weight adjustment mechanism and an attention calculation model to achieve multimodal feature fusion, and generating a joint decision-making strategy through a deep learning optimization model. The system dynamically allocates sensor weights according to the task type, and introduces a multi-objective optimization mechanism for energy consumption, accuracy and safety, while setting fault-tolerant rules to automatically restore weights or recalibrate sensors. The present invention solves the problems of rigid data fusion and single optimization dimension in traditional multimodal decision-making, and significantly improves the accuracy, response speed and environmental adaptability of collaborative operations of industrial robots.
Owner:SHENZHEN HUAZHONG NUMERICAL CONTROL

system

We provide the system. [Solution] A device for analyzing visual and auditory information collected in the work environment, A device that identifies tasks that can be automated based on the aforementioned analysis, A device that proposes an optimal automation method for the aforementioned task, A system that includes this.
Owner:SOFTBANK GROUP CORP

Communication method and system of AR glasses with earphone and glasses separated, and medium

The invention provides a communication method and system for AR glasses with separated earphones and glasses, and a medium, and the method comprises the steps: initializing the AR glasses, and obtaining the initialization state information; scanning compatible state information between the AR glasses and the separated audio equipment based on the initialized state information; establishing communication state information between the AR glasses and the separated audio equipment based on the compatible state information; performing communication matching authentication based on the communication state information between the AR glasses and the separated audio equipment to obtain an authentication result; analyzing whether the communication is successful or not based on the authentication result, if the communication is successful, outputting audio information and instruction response, and if the communication is failed, re-pairing the separated audio equipment; communication matching is completed by analyzing the communication state of the AR glasses and the separated audio equipment, so that successful communication is realized, cooperative transmission of visual information at a glasses end and auditory information at an earphone end is realized, and synchronism of audio and pictures is ensured.
Owner:SHENZHEN ORANGE ELECTRONICS CO LTD

System

An object of a system according to an embodiment is to perform information presentation utilizing multiple senses using generated AI.SOLUTION: A system includes a visual information presentation part, an auditory information presentation part, a tactile information presentation part, a taste information presentation part, and an olfactory information presentation part. The visual information presenting unit presents visual information by using the generated AI. The audio information presenting unit presents audio information by using the generated AI. The tactile information presenting unit presents tactile information by using the generated AI. The taste information presenter presents taste information by using the generated AI. The olfactory information presenting unit presents olfactory information by using the generated AI.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Control method of teaching device, control device, teaching system, and storage medium

The application discloses a teaching equipment control method, a control device, a teaching system and a storage medium. The teaching equipment control method comprises the following steps: image and audio of a target in a teaching space are collected to obtain image data and audio data of the target, wherein the teaching space comprises the teaching equipment; visual information of the target is extracted by using the image data of the target, and auditory information of the target is extracted by using the audio data of the target; and the teaching equipment is controlled based on the visual information and the auditory information of the target. In the foregoing manner, the convenience of the teaching equipment control can be improved, and the accuracy is relatively high.
Owner:IFLYTEK LINGZHI (JIANGSU) TECH CO LTD +1

A Cloud-Edge-Terminal Collaborative Hearing Inference Method with Controllable Multimodal Perception Flow

The present invention relates to a cloud-edge-end collaborative hearing aid inference method with controllable multi-modal perception streams. First, GRU is used for audiovisual event recognition, visual and auditory features are extracted respectively, and under the condition of controlling the visual features, a recurrent neural network is used for audiovisual event recognition. At the decision-making layer, the results of audiovisual information are fused by means of dynamic weighting. According to the influence of visual perception amount on the recognition accuracy and the influence of the change of audiovisual information weight on the final result, two parameters meeting the requirements are determined according to the actual situation. Secondly, cross-modal data generation is carried out based on DCGAN, visual and auditory features are extracted, and a part of visual features are fused with auditory features. Under the condition of controlling the visual information perception amount for fusion, a generative adversarial network is used for cross-modal data generation. According to the evaluation method of the performance of the generative adversarial network and the influence of visual information perception amount on the quality of the generated samples, parameters meeting the actual requirements are determined.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

system

We provide the system. [Solution] A device that performs information processing to generate an individually optimized learning plan based on the learner's abilities, progress, and goals, A device that provides educational materials based on a generated learning plan, A device that analyzes learners' activities with educational materials in real time and generates evaluation results, A device that analyzes a learner's past learning history to identify learning challenges and provides additional exercises to improve those challenges, A sensor that acquires sound and video, and a device that uses the acquired data to perform interactive speech practice in real time, A device that provides immediate instruction using visual and auditory information based on speech evaluation data, A system that includes this.
Owner:SOFTBANK GROUP CORP

Robot control system, robot control program

PendingCN122319059ALittle fingerEngineering
The hand tool 50 is equipped with sensor groups with different information acquisition functions on each of the fingers 22A, 22B, 22C, 22D, and 22E. In other words, it is equipped with sensor groups for acquiring various information required for multimodal robot control. A visual sensor 51A for detecting visual information is installed on the thumb, an auditory sensor 51B for detecting auditory information is installed on the index finger, an olfactory sensor 51C is installed on the middle finger, a tactile sensor 51D for detecting tactile information is installed on the ring finger, and a taste sensor 51E for detecting taste information is installed on the little finger.
Owner:SOFTBANK GROUP CORP

Vehicle operation device

This vehicle operation device (100) comprises: an operation panel (21) on which a plurality of operation buttons (22) are displayed; a touch detection unit (30) which detects, as a touch operation, contact with any of the plurality of operation buttons (22); an output unit (50) which can output visual information, tactile information, and auditory information; and a control unit (10) which, in response to the detection of the touch operation by the touch detection unit (30), causes the output unit to output the visual information, tactile information, and auditory information as feedback information for the touch operation. The control unit (10) causes the output unit (50) to output, for each of the operation buttons (22), feedback information in which at least one of the visual information, tactile information, and auditory information is different.
Owner:U SHIN LTD

Vehicle automatic driving visual and auditory information fusion method and device

The invention discloses a vehicle automatic driving visual and auditory information fusion method and device. The method comprises the steps of obtaining visual information and auditory information collected by a vehicle; the auditory information comprises a sound source direction; projecting the auditory information into a coordinate system of the visual information; fusing the visual information and the auditory information in the same coordinate system through an attention mechanism to obtain fused features; projection of the auditory information to a visual coordinate system of the visual information includes: acquiring an internal reference matrix of a vehicle camera; obtaining a distance value from the vehicle camera to the sound source according to the sound source direction; and converting the sound source direction into coordinates in a visual coordinate system according to the internal reference matrix and the distance value. After the visual information and the auditory information of the vehicle are obtained, the auditory information is projected into the coordinate system of the visual information, so that the auditory information and the visual information can be effectively fused in the same coordinate system, target detection and positioning are performed through the fused features, and the perception precision and robustness of the vehicle to a complex scene are improved.
Owner:FUJIAN UNIV OF TECH

Artificial intelligence-based english spoken language pronunciation correction system

PendingCN122392380AVisually impairedMemory retention
The application belongs to the field of intelligent teaching, and particularly relates to an English oral pronunciation correction system based on artificial intelligence, which comprises a multi-modal perception module, a deep pronunciation feature extraction module, a cognitive compensation fusion module and an intelligent correction and training planning module; the abstract pronunciation is converted into a tactile experience through a multi-modal learning process, the tactile information and the auditory information are synchronously input, the memory retention rate of the visually impaired students is improved, and the students are guided to deeply remember the acoustic characteristics of the pronunciation; the application simulates the auditory-tactile joint representation ability of the visually impaired brain through the learnable attention weight and the cognitive compensation prior mask, realizes the real sense of "touch instead of vision", simultaneously identifies the error reasons through multi-task learning, provides scientific basis for targeted training, designs a special tactile generation network, maps the acoustic characteristics and the pronunciation error types into tactile patterns, provides instant tactile correction feedback for the students, and forms a closed-loop multi-modal learning experience.
Owner:SHANDONG VOCATIONAL COLLEGE OF SPECIAL EDUCATION

A method and device for learning semantic understanding of auditory information for autonomous driving of a vehicle

The present invention discloses a method and device for learning semantic understanding of auditory information for autonomous driving of a vehicle. The method comprises the following steps: collecting sound events in a virtual environment through a virtual audio collection module to obtain a virtual audio data set, classifying the virtual audio data in the virtual audio data set and generating pseudo labels; obtaining scene information of target virtual audio data and using it as scene-based semantic labels for the target virtual audio data; using the labeled virtual audio data set for deep learning model training to obtain an audio analysis model; obtaining real-time audio data, and extracting audio features of the real-time audio data through a feature extraction layer of the audio analysis model; obtaining context features corresponding to the audio features in a time series, and generating audio time features based on the target audio features and the context features; and analyzing the audio time features through an output layer of the audio analysis model to obtain classifications and scenes corresponding to the real-time audio data, thereby improving the perception and understanding capabilities of autonomous driving vehicles for auditory information.
Owner:FUJIAN UNIV OF TECH

Cow artificial insemination simulation method and device

The invention relates to the technical field of animal husbandry breeding, and discloses a cow artificial insemination analog simulation method and device, and the method comprises the following steps: a computer terminal loads a preset standardized operation parameter set which is used for defining objective training standards; collecting multi-modal data generated by multiple types of sensors arranged on the simulation cattle model and the operation instrument in real time, and analyzing the multi-modal data into a real-time state vector representing the current operation; dynamically evaluating the real-time state vector based on a loaded standardized operation parameter set and preset collaborative logic to generate a structured evaluation result; and driving the computer terminal to provide real-time multimedia interaction feedback containing visual and auditory information to the operator according to the generated evaluation result. Through multi-modal sensing, collaborative logic and a quantitative evaluation model, objective evaluation and dynamic physiological feedback of operation are realized, and the trueness and standardization level of training are comprehensively improved.
Owner:NANJING LAIYITE ELECTRONIC TECH CO LTD

Information processing system, information processing method, and program

PCT designated stage expiredWO2025104969A1Data processing applicationsInformation processingEngineering
[Problem] To provide, for example, an information processing system capable of offering a novel user experience with a highly soothing effect. [Solution] One aspect of the present invention provides an information processing system. This information processing system includes at least one processor, and the processor is configured to execute a program so that each of the following steps is performed. In a setting step, theme information indicating at least one of a plurality of pre-managed themes is set. In a first presentation step, multimodal information corresponding to the theme information is presented to a user located in a closed space. The multimodal information is a combination of information comprising at least two of visual information, auditory information, tactile information, and olfactory information.
Owner:KK TOYOTA CHUO KENKYUSHO +1

system

We provide the system. [Solution] Features that acquire the learner's visual and auditory information in real time, The function analyzes the data obtained by the aforementioned function and identifies the mental state related to the learner's level of understanding, A function that generates individualized instruction and practice problems based on the mental state identified by the aforementioned function, The function presents the instruction and problems generated by the aforementioned function to the learner, The machine acts as a learning supporter within the home, providing optimal learning materials according to the learner's progress, A system that includes this.
Owner:SOFTBANK GROUP CORP

AI vision camera (HUSKYLENS)

1. Name of the product in this design: AI Vision Camera (HUSKYLENS). 2. Purpose of this design: This product is used to collect, identify and process visual and auditory information, and output the results. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design's key points: a 3D model.
Owner:SHANGHAI ZHIWEI ROBOT CO LTD

system

The system according to this embodiment aims to enable visually impaired or bedridden people to enjoy video content. [Solution] The system according to the embodiment comprises an acquisition unit, an analysis unit, a language conversion unit, a speech conversion unit, and a provision unit. The acquisition unit acquires auditory information from video content. The analysis unit analyzes the visual information contained in the video based on the auditory information acquired by the acquisition unit. The language conversion unit translates the visual information analyzed by the analysis unit into language. The speech conversion unit converts the visual information translated by the language conversion unit into speech. The provision unit provides audio content by combining the speech converted by the speech conversion unit with the auditory information.
Owner:SOFTBANK GROUP CORP

A multi-modal fusion video classification method and system based on brain-like feedback interaction

The application discloses a kind of multi-modal fusion video classification method and system based on brain-like feedback interaction, method includes: video pre-processing, obtain visual information and auditory information in video;After feature extraction, input into the multi-modal fusion framework based on neural network architecture search to obtain fusion information representation after fusion;Fusion information representation, visual information and auditory information are input into feedback module based on feedback modulation effect generated by single sensory cortex after auditory and visual information integration of human brain in superior temporal sulcus, and the output multi-modal fusion visual information and multi-modal fusion auditory information are respectively through fully connected layer to obtain the confidence of each classification;The confidence of each classification is input into DS decision fusion module, and the final classification result is obtained.The application learns from the processing mode of human brain to perceive external environment, integrates multiple senses from various modal perception information, and effectively improves the accuracy of recognizing and classifying the expressions of people in the video.
Owner:HOHAI UNIV

Speedometers that utilize touch and hearing

PendingJP2026112346ARoad vehicles traffic controlAudible meter reading indicationAcousticsTesting Methods
The present invention provides a device that divides the vehicle's speed into multiple stepped speed ranges and outputs different tactile or auditory information according to the number of steps in each speed range. [Solution] The number of speed range steps in the tactile speedometer is changed from the usual decimal number to a lower base n number, and the number of different values ​​used by each digit of the new base n number is reduced, thereby reducing the total number of numerical patterns used by the tactile speedometer to display numerical values, making it easier to understand the number of speed range steps via touch. Furthermore, it is possible to change the type of base n number used for each speed range step, so that the speed range indicated by the first digit of each base n number is shown in the same speed range unit across the entire speed range, making it possible to unify the speed range of the first digit. This enables the speedometer to display speed in a unified speed range unit.
Owner:大庭 有二