Bionic echo positioning method based on machine vision

Through machine vision and artificial intelligence technology, spatial depth information is converted into audio signals, imitating bats' auditory perception methods, solving the problem that blind people and robots find it difficult to perceive spatial depth and obstacles, and realizing the function of sound sensing spatial depth.

CN120148463APending Publication Date: 2025-06-13OR (BEIJING) SCIENCE & TECHNOLOGY CONSULTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510360317.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-20
Filing Date
2025-03-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

It is difficult for blind people and robots to perceive and identify spatial depth and obstacle locations in traditional ways.

Method used

Using machine vision and artificial intelligence technology, it imitates the auditory perception method of bats, converts spatial depth information into audio signals through echolocation, and uses audio to express the location and depth of obstacles in the space.

Benefits of technology

The ability of blind people and robots to perceive spatial depth and identify obstacles through sound, enhancing their spatial perception capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148463A_ABST
    Figure CN120148463A_ABST
Patent Text Reader

Abstract

The invention relates to a method for converting space depth information into audio by utilizing a machine vision technology, and by utilizing the method, bionic echo positioning can be realized, namely, an audio signal which does not contain semantics but can express the position of an obstacle in a space and the space depth is generated. According to the principle, after space information is collected through machine vision, the position and the space depth of an obstacle are analyzed through artificial intelligence, the result is a two-dimensional tensor expressing three-dimensional information, then the Z axis expressing the space depth corresponds to an audio time domain, the X axis and the Y axis expressing the width and the height of a picture correspond to frequency pitches, and therefore a section of audio capable of simulating echoes can be generated. According to the scheme, auditory sense is utilized to make up the spatial perception of the blind, and the final product may be AI blind guide glasses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for converting spatial depth information into audio using machine vision technology. Using this method, an audio signal that does not contain semantics but can express the positions of obstacles in space and spatial depth can be generated. The technical fields involved include machine vision, artificial intelligence, bionics, etc. Background Art

[0002] The present invention innovates a method that enables robots or blind people to perceive spatial depth using sound and identify obstacles, and is expected to make up for the spatial perception ability of blind people through machine vision and artificial intelligence technologies. This method may be implemented as an AI glasses for guiding the blind. Summary of the Invention

[0003] The core principle of the present invention: Imitate the ability of bats to perceive space through hearing, and the process of creating "echoes" is realized by machine vision and artificial intelligence technologies. After machine vision collects spatial information, artificial intelligence analyzes the positions of obstacles and spatial depth. The result is a two-dimensional tensor expressing three-dimensional information, where the rows and columns correspond to the X and Y axes (width and height of the frame), and the element values represent the Z axis (spatial depth). Further, the Z axis is encoded as time series (the earlier or later the echo represents the distance), and the X and Y axes can be encoded as pitch and frequency respectively, or the frequencies can be superimposed to further reduce the dimension. The result is an audio containing spatial description information. This audio does not directly contain semantics, but can be restored to a spatial depth matrix to a certain extent.

[0004] Typical implementation scheme of the present invention: Use an avatar sensor or radar to collect spatial images or 3D point clouds, and segment, identify, and estimate the depth of the above data through an artificial intelligence model. After obtaining three-dimensional depth data, map the depth (distance) values of the data to the time domain of the audio, and map the left-right, up-down position values to the frequency domain (normalized to the frequency space distinguishable by hearing) and pitch, thereby obtaining an audio segment. During the process of implementing it as a product, periodic markers can be added to the audio segment as needed to completely simulate the cycle from one sound emission to the echo. Brief Description of the Drawings

[0005] The attached drawing [Figure 1] is a flowchart of the method described in the present invention, revealing the key process of bionic echolocation based on machine vision. Among them, the process node P1 represents the start of the process, that is, using an image sensor to collect spatial images (P2), and then handing over the above image data to P3 - an artificial intelligence depth estimation model to obtain three-dimensional data expressing spatial depth information (P4). Thus, it comes to a key step of the present invention - expressing the Z axis (depth) of the spatial three-dimensional data as the time series of the audio (P5.1), and corresponding the X and Y axes to frequency and pitch respectively, or superimposing the frequencies to a single audio track to achieve dimension reduction (P5.2). Finally, as shown in P6 in the figure, an audio signal is output.

Claims

1. A method for generating bionic echolocation audio based on machine vision, the key feature of which is the conversion of spatial depth (far and near) into the time domain of audio.

2. Based on the above 1, the width and height of the spatial image correspond to the frequency and pitch of the audio respectively.

3. Based on the above 1, the image width or height is matched to different frequencies and then superimposed and output.