Bionic echo positioning method based on machine vision
Through machine vision and artificial intelligence technology, spatial depth information is converted into audio signals, imitating bats' auditory perception methods, solving the problem that blind people and robots find it difficult to perceive spatial depth and obstacles, and realizing the function of sound sensing spatial depth.
Patent Information
- Application Number
- CN202510360317.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-11-20
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-13
AI Technical Summary
It is difficult for blind people and robots to perceive and identify spatial depth and obstacle locations in traditional ways.
Using machine vision and artificial intelligence technology, it imitates the auditory perception method of bats, converts spatial depth information into audio signals through echolocation, and uses audio to express the location and depth of obstacles in the space.
The ability of blind people and robots to perceive spatial depth and identify obstacles through sound, enhancing their spatial perception capabilities.
Smart Images

Figure CN120148463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for converting spatial depth information into audio using machine vision technology. Using this method, an audio signal that does not contain semantics but can express the positions of obstacles in space and spatial depth can be generated. The technical fields involved include machine vision, artificial intelligence, bionics, etc. Background Art
[0002] The present invention innovates a method that enables robots or blind people to perceive spatial depth using sound and identify obstacles, and is expected to make up for the spatial perception ability of blind people through machine vision and artificial intelligence technologies. This method may be implemented as an AI glasses for guiding the blind. Summary of the Invention
[0003] The core principle of the present invention: Imitate the ability of bats to perceive space through hearing, and the process of creating "echoes" is realized by machine vision and artificial intelligence technologies. After machine vision collects spatial information, artificial intelligence analyzes the positions of obstacles and spatial depth. The result is a two-dimensional tensor expressing three-dimensional information, where the rows and columns correspond to the X and Y axes (width and height of the frame), and the element values represent the Z axis (spatial depth). Further, the Z axis is encoded as time series (the earlier or later the echo represents the distance), and the X and Y axes can be encoded as pitch and frequency respectively, or the frequencies can be superimposed to further reduce the dimension. The result is an audio containing spatial description information. This audio does not directly contain semantics, but can be restored to a spatial depth matrix to a certain extent.
[0004] Typical implementation scheme of the present invention: Use an avatar sensor or radar to collect spatial images or 3D point clouds, and segment, identify, and estimate the depth of the above data through an artificial intelligence model. After obtaining three-dimensional depth data, map the depth (distance) values of the data to the time domain of the audio, and map the left-right, up-down position values to the frequency domain (normalized to the frequency space distinguishable by hearing) and pitch, thereby obtaining an audio segment. During the process of implementing it as a product, periodic markers can be added to the audio segment as needed to completely simulate the cycle from one sound emission to the echo. Brief Description of the Drawings
[0005] The attached drawing [Figure 1] is a flowchart of the method described in the present invention, revealing the key process of bionic echolocation based on machine vision. Among them, the process node P1 represents the start of the process, that is, using an image sensor to collect spatial images (P2), and then handing over the above image data to P3 - an artificial intelligence depth estimation model to obtain three-dimensional data expressing spatial depth information (P4). Thus, it comes to a key step of the present invention - expressing the Z axis (depth) of the spatial three-dimensional data as the time series of the audio (P5.1), and corresponding the X and Y axes to frequency and pitch respectively, or superimposing the frequencies to a single audio track to achieve dimension reduction (P5.2). Finally, as shown in P6 in the figure, an audio signal is output.
Claims
1. A method for generating bionic echolocation audio based on machine vision, the key feature of which is the conversion of spatial depth (far and near) into the time domain of audio.
2. Based on the above 1, the width and height of the spatial image correspond to the frequency and pitch of the audio respectively.
3. Based on the above 1, the image width or height is matched to different frequencies and then superimposed and output.