Three-dimensional space obstacle perception data processing method, device and electronic equipment

The method of generating three-dimensional data and converting it into sound data through binocular cameras and graphics recognition algorithms solves the problem of low perception accuracy of blind walking glasses equipment, and realizes accurate perception and autonomous actions of blind people about the environment.

CN120279495BActive Publication Date: 2025-08-19NANJING TAIRUI HUIZHI ELECTRONIC INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510758295.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-19
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

Existing blind walking glasses equipment cannot accurately perceive environmental information, resulting in a low accuracy in perceiving the surrounding environment for blind people.

Method used

The environment information image is extracted through a binocular camera, and a graphic recognition algorithm is used to generate three-dimensional data of the object, and convert it into dimensional data for sound frequency, volume and playback time. The sound playback is performed through blind action assisted glasses so that blind people can perceive the location and outline of obstacles in the environment.

Benefits of technology

The blind people's three-dimensional spatial obstacle perception ability of the environment is improved, allowing the blind people to perceive environmental information more accurately, realize the continuous picture perception of the information, and enhance the blind people's movement autonomy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279495B_ABST
    Figure CN120279495B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, and electronic device for processing three-dimensional spatial obstacle perception data, relating to the field of electronic technology, and addressing the technical issue of low accuracy in the perception of environmental information by blind people using existing mobility-aiding glasses. The method comprises: extracting an environmental information image corresponding to the direction in which the user's head is facing using a binocular camera; performing three-dimensional spatial recognition of a plurality of physical objects using a graphic recognition algorithm based on the environmental information image to generate three-dimensional object data corresponding to the plurality of physical objects; converting the three spatial dimensions of the object data into sound representation dimension data; and playing sound using a device such as mobility-aiding glasses based on the target sound frequency representation dimension data, target sound volume representation dimension data, and target sound playback time representation dimension data corresponding to the three spatial dimensions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electronic technology, and in particular to a method, device and electronic device for processing three-dimensional space obstacle perception data. Background Art

[0002] Currently, available eyewear products for the blind, based on camera-based video recognition or acoustic ranging radar, can provide features such as environmental alerts, obstacle proximity warnings, object recognition, and text reading, making life easier for the blind. However, existing devices fail to provide specific awareness of their surroundings, relying solely on voice-based alerts and non-directional proximity alarms. These existing eyewear products are limited in use and incomplete, resulting in low accuracy in the blind's perception of environmental information. Summary of the Invention

[0003] The purpose of the present invention is to provide a three-dimensional space obstacle perception data processing method, device and electronic equipment to solve the technical problem that the existing blind walking glasses make the blind people's perception of environmental information have low accuracy.

[0004] In a first aspect, the present application provides a method for processing three-dimensional spatial obstacle perception data, which is applied to a blind person's mobility assistance glasses-type device, wherein the blind person's mobility assistance glasses-type device is provided with a binocular camera and the blind person's mobility assistance glasses-type device is worn on the user's head; the method comprises:

[0005] Extracting an environmental information image corresponding to the direction in which the user's head is facing by the binocular camera; the environment corresponding to the environmental information image includes several physical objects closest to the user;

[0006] Based on the environmental information image, a graphic recognition algorithm is used to perform three-dimensional spatial recognition on the plurality of physical objects to generate three-dimensional object data corresponding to the plurality of physical objects;

[0007] Converting the three spatial dimension data in the three-dimensional object data into sound representation dimension data respectively; the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, and each of the three spatial dimension data is converted into a corresponding sound representation dimension data;

[0008] Based on the target sound frequency representation dimension data, target sound volume representation dimension data and target sound playback time representation dimension data converted into the three spatial dimension data, sound is played through the blind mobility assistive glasses type device, so that the user can perceive the outlines of the several physical objects in the environment and their relative positions relative to the user through the target sound frequency representation dimension data, the target sound volume representation dimension data and the target sound playback time representation dimension data.

[0009] In one possible implementation, performing three-dimensional spatial recognition on the plurality of physical objects using a graphic recognition algorithm based on the environmental information image to generate three-dimensional object data corresponding to the plurality of physical objects includes:

[0010] orthographically projecting the environmental information image onto a plurality of pixel windows formed by a plurality of rows and columns, and determining up, down, left, and right position information of the plurality of physical objects in the environmental information image;

[0011] Performing three-dimensional spatial recognition on the multiple physical objects based on the up, down, left, and right position information and the sizes of the multiple physical objects in the multiple pixel windows, and determining front-back distance information of the multiple physical objects in the multiple pixel windows;

[0012] Three-dimensional object data corresponding to the plurality of physical objects is generated based on the up, down, left, and right position information and the front and back distance information.

[0013] In one possible implementation, converting the three spatial dimension data in the three-dimensional object data into sound representation dimension data respectively includes:

[0014] Converting the up-down position information in the up-down, down-down, left-right position information in the three-dimensional data of the object into the sound frequency representation dimensional data, so as to represent a vertically continuous column of the pixel windows from top to bottom in the up-down, down-down, left-right position information by different single-frequency sounds; wherein the higher the pixel position in the up-down, down-down, left-right position information, the lower the frequency corresponding to the single-frequency sound, and the lower the pixel position in the up-down, down-down, left-right position information, the higher the frequency corresponding to the single-frequency sound;

[0015] Converting the front-to-back distance information in the three-dimensional object data into sound volume representation dimensional data, so that the front-to-back distance information of the corresponding pixel position of the physical object is represented by the sound volume of each single-frequency sound; wherein the closer the front-to-back distance corresponding to the front-to-back distance information is, the louder the sound volume is, and the farther the front-to-back distance is, the quieter the sound volume is; the closer the front-to-back distance is, the faster the sound volume decreases, and the farther the front-to-back distance is, the slower the sound volume decreases;

[0016] generating a mixed sound signal using a sound wave mixer algorithm based on the different single-frequency sounds and the sound volume of each of the single-frequency sounds, wherein the mixed sound signal represents the up-down position information and the front-to-back distance information corresponding to a column of vertically continuous pixel windows;

[0017] The plurality of mixed sound signals are played in sequence in the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are continuous in the left and right horizontal directions, so that the plurality of mixed sound signals played in sequence represent the comprehensive three-dimensional information corresponding to all the pixel windows, and the comprehensive three-dimensional information includes the up, down, left, and right position information and the front and back distance information.

[0018] In one possible implementation, converting the front-to-back distance information in the three-dimensional object data into sound volume representation dimensional data, so as to represent the front-to-back distance information of the corresponding pixel position of the physical object by the sound volume of each single-frequency sound, includes:

[0019] Performing a logarithmic calculation on the front-to-back distance information in the three-dimensional data of the object to obtain a logarithmic calculation result, and based on the logarithmic calculation result, converting the front-to-back distance information into a volume value using the following sound volume-to-front-to-back distance correspondence conversion formula:

[0020] A= 30 / X 1 / 3 ;

[0021] Wherein, X is the front-to-back distance information, and the calculation result of the reciprocal function of the X power operation represents the volume value A corresponding to the front-to-back distance information. The unit of the volume value A is dB, and A represents the average front-to-back distance corresponding to the physical object contained in a pixel window. The front-to-back distance information of the pixel position corresponding to the physical object contained in the pixel window is represented by the volume value of each single-frequency sound.

[0022] In a possible implementation, the blind person mobility assist glasses type device corresponds to a plurality of rows and a plurality of columns;

[0023] The projecting of the environmental information image into a plurality of pixel windows formed by a plurality of rows and a plurality of columns comprises:

[0024] Testing the user's auditory recognition ability to obtain a test result of the user's auditory recognition ability; wherein the test result of the user's auditory recognition ability includes a test result of the user's auditory recognition ability of sound frequency;

[0025] Adjusting the number of the plurality of rows and the number of the plurality of columns according to the test result of the user's auditory recognition ability of sound frequencies to obtain a pixel window fineness adjustment result;

[0026] A plurality of pixel windows formed by a plurality of rows and a plurality of columns are determined according to the pixel window fineness adjustment result, and the environmental information image is orthographically projected into the plurality of pixel windows formed by the plurality of rows and a plurality of columns.

[0027] In one possible implementation, the user's auditory recognition ability test result further includes the user's auditory recognition ability test result for sound reception; and the sequentially playing the multiple mixed sound signals in the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are horizontally continuous includes:

[0028] adjusting the sound output length of each of the mixed sound signals according to the test result of the user's auditory recognition ability of sound reception, to obtain a sound output length adjustment result;

[0029] According to the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are continuous in the horizontal direction, and according to the sound output length corresponding to the sound output length adjustment result, the multiple mixed sound signals are played in sequence.

[0030] In a possible implementation, playing the plurality of mixed sound signals in sequence according to the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are horizontally continuous includes:

[0031] During actual playback, a 1000 Hz head positioning signal of a first specified number of seconds is played before the multiple mixed sound signals corresponding to each frame of the environmental information image are played. After the head positioning signal is played, the multiple mixed sound signals are played sequentially according to the order in which the vertically continuous pixel windows in each column of the up, down, left, and right position information are horizontally continuous, with each column of the pixel windows having a sound output length of a second specified number of seconds, so as to synthesize and play the multiple mixed sound signals corresponding to multiple frames of the environmental information image through the multiple head positioning signals.

[0032] The first specified number of seconds is less than the second specified number of seconds.

[0033] In one possible implementation, the binocular camera is a structure of two sets of dual camera modules, and the two sets of dual camera modules correspondingly extract two sets of environmental information images, and the two sets of environmental information images represent the left field of view environmental information image and the right field of view environmental information image of the user;

[0034] Playing the plurality of mixed sound signals in sequence according to the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are horizontally continuous includes:

[0035] Playing, through the left audio channel of the blind person mobility assist glasses-type device, the plurality of mixed sound signals corresponding to the left visual field environment information image in sequence according to the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are horizontally continuous;

[0036] Through the right channel of the blind person's mobility assistance glasses type device, the multiple mixed sound signals corresponding to the right field of view environmental information image are played in sequence according to the order of the left and right horizontal continuity of the vertically continuous pixel windows in each column in the up, down, left and right position information.

[0037] In a second aspect, the present application provides a three-dimensional spatial obstacle perception data processing device, which is applied to a blind person's mobility assistance glasses-type device, wherein the blind person's mobility assistance glasses-type device is provided with a binocular camera, and the blind person's mobility assistance glasses-type device is worn on the user's head; the device comprises:

[0038] an extraction module, configured to extract, through the binocular camera, an environmental information image corresponding to the direction in which the user's head is facing; wherein the environment corresponding to the environmental information image contains a plurality of physical objects;

[0039] A recognition module, configured to perform three-dimensional spatial recognition of the plurality of physical objects using a graphic recognition algorithm based on the environmental information image, and generate three-dimensional object data corresponding to the plurality of physical objects;

[0040] a conversion module, configured to convert the three spatial dimension data in the three-dimensional object data into sound representation dimension data; wherein the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, and each of the three spatial dimension data is converted into a corresponding sound representation dimension data;

[0041] A playback module is used to play sound through the blind person's mobility assist glasses-type device based on the target sound frequency representation dimension data, target sound volume representation dimension data, and target sound playback time representation dimension data converted into the three spatial dimension data, so that the user can perceive the position and outline of the multiple physical objects in the environment through the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data.

[0042] In a third aspect, the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the method described in the first aspect is implemented.

[0043] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute the method described in the first aspect above.

[0044] This application brings the following beneficial effects:

[0045] The present application provides a three-dimensional spatial obstacle perception data processing method, device and electronic device, which can extract an environmental information image corresponding to the direction facing the user's head through a binocular camera; the environment corresponding to the environmental information image contains several physical objects closest to the user; based on the environmental information image, three-dimensional spatial recognition is performed on the several physical objects through a graphic recognition algorithm to generate three-dimensional object data corresponding to the several physical objects; the three spatial dimension data in the object three-dimensional data are respectively converted into sound representation dimension data; the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data and sound playback time representation dimension data, and each of the three spatial dimension data is converted into a corresponding sound representation dimension data; based on the target sound frequency representation dimension data, target sound volume representation dimension data and target sound playback time representation dimension data converted from the three spatial dimension data, sound is played through a blind person's mobility assist glasses type device, so that the user can perceive the outline of several physical objects in the environment and their relative position relative to the user through the target sound frequency representation dimension data, target sound volume representation dimension data and target sound playback time representation dimension data. In this solution, by representing dimensional data through sound frequency, sound volume and sound playback time, the blind can perceive the real three-dimensional data of each entity object in the environment, thereby enabling the blind to perceive more comprehensive environmental information in the front, back, left, right, up and down directions, and enabling the blind to gain the ability to perceive three-dimensional spatial obstacles, thereby presenting a continuous picture perception of information, that is, providing the blind with the ability to perceive spatial obstacles, thereby improving the accuracy of the blind's perception of environmental information, and solving the technical problem of low accuracy of the blind's perception of environmental information through existing blind walking glasses.

[0046] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0048] Figure 1 A flowchart of a method for processing three-dimensional space obstacle perception data provided in an embodiment of the present application;

[0049] Figure 2 This is an example of a glasses-type device for assisting blind people in their mobility, in the three-dimensional space obstacle perception data processing method provided in the embodiments of the present application;

[0050] Figure 3 An example of an environmental information image in the three-dimensional space obstacle perception data processing method provided in the embodiment of the present application;

[0051] Figure 4 An example of a three-dimensional object outline in the three-dimensional space obstacle perception data processing method provided in an embodiment of the present application;

[0052] Figure 5 This is an example of up, down, left, right, and front-to-back distance information corresponding to multiple pixel windows in the three-dimensional space obstacle perception data processing method provided in an embodiment of the present application;

[0053] Figure 6 This is an example of comparing the sound frequencies and sound volumes corresponding to multiple pixel windows in the three-dimensional space obstacle perception data processing method provided in an embodiment of the present application;

[0054] Figure 7 A schematic diagram of the structure of a three-dimensional space obstacle perception data processing device provided in an embodiment of the present application;

[0055] Figure 8 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0056] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0057] The terms "including," "having," and any variations thereof, as used in the embodiments of this application, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.

[0058] Currently, the unresolved issue is how to enable blind people to quickly visualize their surroundings, allowing them to move freely. Currently, devices available for the blind include the following: using cameras to identify the surrounding environment; using ultrasound probes to scan for obstacles in the forward direction; relying on voice prompts, such as "There is an obstacle one meter to the left, go around to the right." Another method involves projecting a camera image and then placing an electrode array on the tongue. This converts image pixels into electrode stimulation voltages, which are sensed by the tongue and then reflected by the brain as the external image. Currently, these methods have been tested and proven to be effective, demonstrating that the human body can form images in the brain through various sensory perceptions.

[0059] However, existing devices cannot allow blind people to perceive the specific conditions of their surroundings; they only provide voice alerts and non-directional proximity alarms. Furthermore, while supralingual sensing can provide graphical perception, it is only a two-dimensional image, and its limited, inconvenient use makes it virtually unusable. Therefore, the accuracy of environmental information perceived by blind people using existing mobility aids is low.

[0060] Based on this, the embodiments of the present application provide a three-dimensional space obstacle perception data processing method, device and electronic device, which can solve the technical problem that the existing blind walking glasses make the blind people's perception of environmental information have low accuracy.

[0061] The embodiments of the present invention are further described below with reference to the accompanying drawings.

[0062] Figure 1 This is a flow chart of a method for processing three-dimensional space obstacle perception data provided by an embodiment of the present application. The method is applied to a blind person's mobility assistance glasses type device, which is provided with a binocular camera and is worn on the user's head. Figure 1 As shown, the method includes:

[0063] Step S110 , extracting an environmental information image corresponding to the direction the user's head is facing through a binocular camera.

[0064] The environment corresponding to the environmental information image contains several physical objects closest to the user. As an example, Figure 2The camera module technology on the blind person's mobility assistance glasses type device shown uses two sets of binocular cameras to complete two sets of three-dimensional space scene information extraction, and can further obtain the location information of obstacles in the environment through the dual camera module.

[0065] Step S120 , performing three-dimensional spatial recognition on a plurality of physical objects through a graphic recognition algorithm based on the environmental information image, and generating three-dimensional object data corresponding to the plurality of physical objects.

[0066] In an optional embodiment, the environment stereo image (environmental information image) can be obtained by a 3D contour extraction algorithm to calculate the distance to the nearest obstacle in each pixel window. Figure 5 As shown in the figure, the obstacles in the field of view are decomposed into 8×8 pixel windows, and the distance information of each pixel window is calculated.

[0067] Exemplarily, step S120 may include the following steps: orthographically projecting the environmental information image onto a plurality of pixel windows formed by a plurality of rows and columns to determine the up, down, left, and right positional information of the plurality of physical objects within the environmental information image; performing three-dimensional spatial recognition of the plurality of physical objects based on the up, down, left, and right positional information and the sizes of the plurality of physical objects within the plurality of pixel windows to determine the front-to-back distance information of the plurality of physical objects within the plurality of pixel windows; and generating three-dimensional object data corresponding to the plurality of physical objects based on the up, down, left, and right positional information and the front-to-back distance information. This data processing method makes the generated three-dimensional object data more accurate.

[0068] For example, Figure 3 and Figure 4 As shown, the environmental information image captured by the binocular camera is processed by the pattern recognition algorithm to form a three-dimensional stereo contour, so that the three-dimensional image information can be expressed by sound later. Figure 5 As shown, a three-dimensional image is projected onto a 64-pixel window of 8 rows and 8 columns, presenting up, down, left, and right position information. The distance information of this 8×8 pixel window is then marked to describe the three-dimensional contour information of the environment.

[0069] Step S130 : converting the three spatial dimension data in the three-dimensional object data into sound representation dimension data respectively.

[0070] The sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, and each of the three spatial dimension data is converted into a corresponding sound representation dimension data. As an optional implementation, this step S130 may specifically include the following steps:

[0071] Converting the up-down position information in the up-down, down-down, left-right position information in the three-dimensional data of the object into sound frequency representation dimensional data, so as to represent a vertically continuous column of pixel windows from top to bottom in the up-down, down-down, left-right position information by different single-frequency sounds; wherein, the higher the pixel position in the up-down, down-down, left-right position information, the smaller the frequency corresponding to the single-frequency sound, and the lower the pixel position in the up-down, down-down, left-right position information, the larger the frequency corresponding to the single-frequency sound;

[0072] Converting the front-to-back distance information in the three-dimensional data of the object into sound volume representation dimensional data, so that the front-to-back distance information of the corresponding pixel position of the physical object is represented by the sound volume of each single-frequency sound; wherein, the closer the front-to-back distance information corresponding to the front-to-back distance information is, the louder the sound volume is, and the farther the front-to-back distance information is, the smaller the sound volume is; the closer the front-to-back distance information is, the faster the frequency of the single-frequency sound decreases from top to bottom, and the farther the front-to-back distance information is, the slower the frequency of the single-frequency sound decreases from top to bottom;

[0073] A mixed sound signal is generated using a sound wave mixer algorithm based on different single-frequency sounds and the sound volume of each single-frequency sound, where the mixed sound signal represents up-and-down position information and front-and-back distance information corresponding to a column of vertically continuous pixel windows;

[0074] Multiple mixed sound signals are played in sequence according to the order in which each column of vertically continuous pixel windows in the up, down, left, and right position information is continuous in the left and right horizontal directions, so that the multiple mixed sound signals played in sequence represent the comprehensive three-dimensional information corresponding to all pixel windows, and the comprehensive three-dimensional information includes the up, down, left, and right position information and the front and back distance information.

[0075] Through this data processing method, the converted sound representation dimensional data is easier for the blind to perceive, improving the efficiency and convenience of perception. In actual applications, most blind people have more developed hearing. In terms of frequency resolution, ordinary people can distinguish 2 pitches, while blind people can distinguish more than 8 pitches. By using this ability, the three-dimensional contour information of the environment can be transmitted to the blind. Figure 5 and Figure 6As shown, 8 different single-frequency sounds are grouped to represent a column of pixels from top to bottom, and the volume of each single-frequency sound represents the distance of the corresponding pixel position. The volume is louder when it is near and smaller when it is far, and it drops faster at the near end and slower at the far end. These 8 single-frequency sounds use the sound wave mixer algorithm to generate a mixed sound signal. This sound depicts the position and distance information in a longitudinal two-dimensional contour section of three-dimensional space. 8 such sound information played in sequence depict a three-dimensional contour containing position and distance information. Stereoscopic imaging is achieved by the image processing function in the brain, so that the blind have a certain degree of accurate three-dimensional imaging vision ability. Using this method, the blind can accurately obtain the surrounding environmental information after training, and can freely move through the crowd, cross the road, go down stairs, and accurately pick up cups, computers, mobile phones and other items indoors.

[0076] In an optional embodiment, converting the front-to-back distance information in the three-dimensional object data into sound volume representation dimensional data, so as to represent the front-to-back distance information of the corresponding pixel position of the physical object by the sound volume of each single-frequency sound, may specifically include the following steps:

[0077] Perform a logarithmic calculation on the front-to-back distance information in the object's three-dimensional data to obtain a logarithmic calculation result. Based on the logarithmic calculation result, the front-to-back distance information is converted into a volume value using the following sound volume-to-front-to-back distance correspondence conversion formula:

[0078] A= 30 / X 1 / 3 Where X is the distance information, and the result of the reciprocal function of X's power operation represents the volume value A corresponding to the distance information. The unit of volume value A is dB, and A represents the average distance of the physical objects contained within a pixel window. The volume value of each single-frequency sound represents the distance information of the physical objects corresponding to the pixel position within the pixel window. This data processing method makes the conversion of volume values more accurate.

[0079] The pixel window distance value after the value is taken can be converted into a sound element. Specifically, the distance information is logarithmically calculated and then converted into a volume value, which represents the average distance of obstacles within a pixel window.

[0080] For example, the frequency of the sound can be selected according to the C key note table. These notes are frequency information that has been proven to be clearly recognizable by humans. Of course, any other frequency can be selected, but the recognition of other frequencies has not yet been confirmed by experience. Figure 5 and Figure 6As shown in the figure, the eight pixel positions can be taken as 523Hz, 587Hz, 659Hz, 698Hz, 740Hz, 784Hz, 880Hz, and 988Hz as row tones. When more rows of corresponding frequencies are needed (for example, to make a 32×32 pixel window), they can be taken from the C key note table. The eight tones correspond to the positions of eight consecutive windows from top to bottom, as shown in the figure. Figure 5 and Figure 6 As shown, the distance of each position is represented by the corresponding volume amplitude. 60dB represents a distance of 0.1 meter. The sound at this decibel level is basically the clear sound amplitude that people are accustomed to hearing. The system provides volume adjustment and users can also adjust the volume according to the environment and habits. In the case of 8×8, the minimum distance to identify the outline of the object is calculated to be 1.25 cm. The corresponding relationship between the distance and the increase in sound decibels is as follows: A (dB) = 30 / X 1 / 3 The conversion between distance and decibel is shown in Table 1:

[0081] Table 1

[0082]

[0083] As an optional implementation, the above-mentioned conversion of the upper and lower position information in the upper, lower, left and right position information of the object three-dimensional stereoscopic data into sound frequency representation dimensional data, so as to represent a vertically continuous column of pixel windows from top to bottom in the upper, lower, left and right position information by different single-frequency sounds, specifically includes the following steps:

[0084] The following conversion formula is used to convert the up-down position information in the three-dimensional object data into sound frequency representation dimensional data, so that different single-frequency sounds are used to represent a vertically continuous column of pixel windows from top to bottom in the up-down position information:

[0085] f ( p )= fmax- − × p ;

[0086] in, f represents the frequency of the sound generated; p Indicates the position of the pixel. The higher the pixel position, the larger the value. Assume that the top pixel position is 0. Pmax is the maximum value of the pixel position; fmin Indicates the lowest sound frequency limit corresponding to the highest pixel position (i.e. the top); fmax Indicates the highest sound frequency limit corresponding to the lowest pixel position (i.e. the bottom); f ( p ) is the given pixel position pThe corresponding sound frequency. Through this data processing method, the conversion of sound frequency representation dimensional data is more accurate.

[0087] In a possible implementation, generating a mixed sound signal using a sound wave mixer algorithm based on different single-frequency sounds and the sound volume of each single-frequency sound may specifically include the following steps:

[0088] Based on the different single-frequency sounds and the volume of each single-frequency sound, the sound wave mixer algorithm is used to generate a mixed sound signal using the following formula:

[0089] S ( t )= ;

[0090] in, S ( t ) represents the final generated mixed sound signal, N is the number of single-frequency sounds participating in the mix, Ai is the amplitude of the i-th sound signal, fi is the frequency of the i-th sound signal, ϕi is the initial phase of the i-th sound signal, and t is the time. Through this data processing method, the data of the generated mixed sound signal is made more accurate.

[0091] The following press Figure 5 The calculation results of the frequency and amplitude information of the sound information are given as an example: the volume converted from the closest end surface distance of the three objects is: 0.25m = 44.44dB...1.75m = 23.23dB...2.25m = 21.37dB. As shown in Table 2, the following example shows Figure 3 The sound synthesis component of the information:

[0092] Table 2

[0093]

[0094] As shown in Table 3, the following example shows the note and frequency comparison table for the key of C:

[0095] Table 3

[0096]

[0097] As an optional implementation, the blind person's mobility assist glasses type device corresponds to a plurality of rows and a plurality of columns; the above-mentioned orthographic projection of the environmental information image into a plurality of pixel windows formed by the plurality of rows and columns may specifically include the following steps:

[0098] Testing the user's auditory recognition ability to obtain a test result of the user's auditory recognition ability; wherein the test result of the user's auditory recognition ability includes a test result of the user's auditory recognition ability of sound frequency;

[0099] According to the test results of the user's auditory recognition ability of sound frequency, the number of rows and the number of columns are adjusted to obtain the pixel window fineness adjustment result; according to the pixel window fineness adjustment result, multiple pixel windows formed by multiple rows and multiple columns are determined, and the environmental information image is projected into the multiple pixel windows formed by the multiple rows and multiple columns.

[0100] Through this data processing method, the fineness of the pixel window is more consistent with and fits the personal hearing recognition of blind users, thereby improving the experience of blind users.

[0101] In actual applications, specific recognition capabilities can be adjusted based on the user's hearing ability, for example, various pixel window accuracies such as 32×32, 24×24, 16×16, and 4×4. Blind people who use the system over a long period of time can develop a common sense of spatial environment, such as length, width, height, distance, openness, narrowness, and sidewalk boundaries.

[0102] In an optional embodiment, the user's auditory recognition ability test result further includes the user's auditory recognition ability test result for sound reception; the above-mentioned playing of the multiple mixed sound signals in sequence according to the order in which each column of vertically continuous pixel windows in the upper, lower, left, and right position information is horizontally continuous may specifically include the following steps:

[0103] The sound output length of each mixed sound signal is adjusted according to the test results of the user's auditory recognition ability of sound reception to obtain a sound output length adjustment result; multiple mixed sound signals are played in sequence according to the order of the left and right horizontal continuity of each column of vertically continuous pixel windows in the up, down, left and right position information, and with the sound output length corresponding to the sound output length adjustment result.

[0104] Through this data processing method, the played sounds are more consistent with and fit the personal hearing conditions of blind users, improving the experience of blind users.

[0105] In the embodiment of the present application, the specific recognition parameters of the device can be adjusted according to the user's hearing ability. For example, the output length of each mixed sound can be adjusted according to the user's hearing ability.

[0106] As an optional implementation, the above-mentioned sequentially playing multiple mixed sound signals according to the order in which each column of vertically continuous pixel windows in the up, down, left, and right position information is horizontally continuous can specifically include the following steps:

[0107] During actual playback, a 1000 Hz head positioning signal of the first specified number of seconds is played before the multiple mixed sound signals corresponding to each frame of the environmental information image are played. After the head positioning signal is played, the multiple mixed sound signals are played in sequence according to the order of the left and right horizontal continuity of each column of vertically continuous pixel windows in the upper, lower, left and right position information, with each column of pixel windows having a sound output length of the second specified number of seconds, so as to synthesize and play the multiple mixed sound signals corresponding to the multiple frames of environmental information images through the multiple head positioning signals; wherein the first specified number of seconds is less than the second specified number of seconds.

[0108] For example, during actual playback, a 1 / 64 second 1000Hz head positioning signal is added in front of each image information, and then the mixed sound is output to the left and right earphones in sequence with a length of 1 / 32 second each, thus synthesizing two 8×8 pixel windows of depth of field information. Stereoscopic imaging is realized by the image processing function in the brain, enabling the blind to have a certain degree of accurate stereoscopic imaging vision ability.

[0109] In an optional embodiment, the binocular camera is a structure of two sets of dual camera modules, and the two sets of dual camera modules correspondingly extract two sets of environmental information images, and the two sets of environmental information images represent the user's left field of view environmental information image and the right field of view environmental information image; the above-mentioned playing of multiple mixed sound signals in sequence according to the order in which each column of vertically continuous pixel windows in the upper, lower, left and right position information is horizontally continuous can specifically include the following steps:

[0110] Play multiple mixed sound signals corresponding to the left visual field environment information image in sequence through the left audio channel of the blind person's mobility assist glasses-type device according to the order in which each column of vertically continuous pixel windows in the up, down, left, and right position information is horizontally continuous;

[0111] Through the right channel of the blind person's mobility assist glasses-type device, multiple mixed sound signals corresponding to the right field of view environmental information image are played in sequence according to the order of the left and right horizontal continuity of each column of vertically continuous pixel windows in the up, down, left and right position information.

[0112] In the embodiment of the present application, the two sets of collected results generate sound wave information for the left and right channels, which are played by headphones and provided to the blind. In this way, the blind can obtain more comprehensive and rich environmental information.

[0113] Step S140, based on the target sound frequency representation dimension data, target sound volume representation dimension data and target sound playback time representation dimension data converted into the three spatial dimension data, sound is played through a blind person's mobility assistive glasses type device, so that the user can perceive the outlines of several physical objects in the environment and their relative positions relative to the user through the target sound frequency representation dimension data, target sound volume representation dimension data and target sound playback time representation dimension data.

[0114] In the embodiment of the present application, by representing dimensional data of sound frequency, dimensional data of sound volume and dimensional data of sound playback time, the blind can perceive the real three-dimensional stereoscopic data of each physical object in the environment, thereby enabling the blind to perceive more comprehensive environmental information of the front, back, left, right, up and down, so that the blind can obtain the ability to perceive three-dimensional spatial obstacles, and then present a continuous picture perception of information, that is, it provides the blind with the ability to perceive spatial obstacles, improves the accuracy of the blind's perception of environmental information, and also improves the versatility of application scenarios of blind mobility assistance glasses-type devices.

[0115] In addition to depicting the relative position of the blind person and the outline, the central camera can also perform environmental recognition (specifically, by referencing GPS and electronic maps to identify locations, road locations, traffic lights, road signs, street signs, signboards, and other text-based environmental information), facial recognition (for seeing acquaintances on the street, or finding someone you've agreed to meet at a designated location), and reading and image reading functions (two-dimensional images can also be expressed using sound waves, but this is time-consuming). Blind people can also interact with the system using voice, buttons, and a scroll wheel, requesting adjustments or answering questions.

[0116] In an alternative implementation, the audio converted from spatial information can be shared with other blind people, sharing information about the surrounding environment from the sender's perspective. This technology can even be used to provide special movies for the blind, of course, these movies are viewed from a first-person perspective.

[0117] Figure 7 A schematic diagram of a three-dimensional space obstacle perception data processing device is provided. The device can be applied to a blind person's mobility assistance glasses type device, wherein the blind person's mobility assistance glasses type device is provided with a binocular camera and the blind person's mobility assistance glasses type device is worn on the user's head. Figure 7 As shown, the three-dimensional space obstacle perception data processing device 700 includes:

[0118] An extraction module 701 is configured to extract, through the binocular camera, an environment information image corresponding to the direction in which the user's head is facing; wherein the environment corresponding to the environment information image includes a plurality of physical objects;

[0119] The recognition module 702 is configured to perform three-dimensional spatial recognition of the plurality of physical objects using a pattern recognition algorithm based on the environmental information image, and generate three-dimensional object data corresponding to the plurality of physical objects;

[0120] a conversion module 703 for converting the three spatial dimensions of the object's three-dimensional data into sound representation dimension data; wherein the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, and each of the three spatial dimensions is converted into a corresponding sound representation dimension data;

[0121] The playback module 704 is used to play sound through the blind person's mobility assist glasses type device based on the target sound frequency representation dimension data, target sound volume representation dimension data and target sound playback time representation dimension data converted from the three spatial dimension data, so that the user can perceive the position and outline of the multiple physical objects in the environment through the target sound frequency representation dimension data, the target sound volume representation dimension data and the target sound playback time representation dimension data.

[0122] The three-dimensional space obstacle perception data processing device provided in the embodiment of the present application has the same technical features as the three-dimensional space obstacle perception data processing method provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects.

[0123] An electronic device provided in an embodiment of the present application is Figure 8 As shown, the electronic device 800 includes a processor 802 and a memory 801 , wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps of the method provided in the above embodiment when executing the computer program.

[0124] See also Figure 8 The electronic device further includes: a bus 803 and a communication interface 804, a processor 802, a communication interface 804 and a memory 801 connected via the bus 803; the processor 802 is used to execute executable modules stored in the memory 801, such as computer programs.

[0125] Memory 801 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between the system network element and at least one other network element is achieved via at least one communication interface 804 (which may be wired or wireless), and may utilize the Internet, a wide area network, a local area network, a metropolitan area network, or the like.

[0126] The bus 803 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 8Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0127] Among them, the memory 801 is used to store programs, and the processor 802 executes the program after receiving the execution instruction. The method executed by the device defined by the process disclosed in any embodiment of the present application can be applied to the processor 802 or implemented by the processor 802.

[0128] The processor 802 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 802 or by instructions in the form of software. The above-mentioned processor 802 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 801, and processor 802 reads the information in memory 801 and, in conjunction with its hardware, completes the steps of the above method.

[0129] Corresponding to the above-mentioned three-dimensional space obstacle perception data processing method, an embodiment of the present application also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to execute the steps of the above-mentioned three-dimensional space obstacle perception data processing method.

[0130] The three-dimensional space obstacle perception data processing device provided in the embodiment of the present application can be specific hardware on the device or software or firmware installed on the device. The device provided in the embodiment of the present application, its implementation principle and the technical effect produced are the same as those in the aforementioned method embodiment. For the sake of brief description, for the parts not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiment, and will not be repeated here.

[0131] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0132] For another example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0133] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0134] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0135] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the three-dimensional space obstacle perception data processing method described in each embodiment of this application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0136] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.

[0137] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-mentioned embodiments within the technical scope disclosed in the present application, or perform equivalent replacements for some of the technical features thereof. However, these modifications, changes, or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application. They should all be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A three-dimensional space obstacle perception data processing method, characterized in that: The method is applied to a blind person's mobility assistance glasses-type device, wherein the blind person's mobility assistance glasses-type device is provided with a binocular camera and the blind person's mobility assistance glasses-type device is worn on the user's head; the method comprises: Extracting an environmental information image corresponding to the direction in which the user's head is facing by the binocular camera; the environment corresponding to the environmental information image includes several physical objects closest to the user; orthographically projecting the environmental information image onto a plurality of pixel windows formed by a plurality of rows and columns, determining up-down, down-down, left-down, and right-down position information of the plurality of physical objects in the environmental information image; performing three-dimensional spatial recognition on the plurality of physical objects based on the up-down, down-down, left-down, and right-down position information and the sizes of the plurality of physical objects in the plurality of pixel windows, and determining front-back distance information of the plurality of physical objects in the plurality of pixel windows; and generating three-dimensional object data corresponding to the plurality of physical objects based on the up-down, down-down, left-down, and right-down position information and the front-back distance information; The up and down position information in the up and down, left and right position information in the three-dimensional data of the object is converted into sound frequency representation dimensional data, so as to represent a column of pixel windows vertically continuous from top to bottom in the up and down, left and right position information by different single-frequency sounds; the higher the pixel position in the up and down, left and right position information, the lower the frequency corresponding to the single-frequency sound, and the lower the pixel position in the up and down, left and right position information, the higher the frequency corresponding to the single-frequency sound; the front and back distance information in the three-dimensional data of the object is converted into sound volume representation dimensional data, so as to represent the front and back distance information of the corresponding pixel position of the physical object by the sound volume of each single-frequency sound; the closer the front and back distance information corresponds to, the louder the sound volume, the farther the front and back distance, the smaller the sound volume, the closer the front and back distance, the faster the sound volume decreases, and the farther the front and back distance is, the faster the sound volume decreases. The slower the speed at which the volume of the sound decreases; a mixed sound signal is generated based on the different single-frequency sounds and the sound volume of each single-frequency sound using an acoustic wave mixer algorithm, and the mixed sound signal represents the up and down position information and the front and back distance information corresponding to a column of vertically continuous pixel windows; according to the order in which the vertically continuous pixel windows in each column of the up and down, left and right position information are continuous in the left and right directions, a plurality of the mixed sound signals are played in sequence, so that the plurality of mixed sound signals played in sequence represent the comprehensive three-dimensional information corresponding to all the pixel windows, and the comprehensive three-dimensional information includes the up and down, left and right position information and the front and back distance information; the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data and sound playback time representation dimension data, and each of the three spatial dimension data corresponds to a conversion of one type of sound representation dimension data; Based on the target sound frequency representation dimension data, target sound volume representation dimension data and target sound playback time representation dimension data converted into the three spatial dimension data, sound is played through the blind mobility assistive glasses type device, so that the user can perceive the outlines of the several physical objects in the environment and their relative positions relative to the user through the target sound frequency representation dimension data, the target sound volume representation dimension data and the target sound playback time representation dimension data.

2. The method according to claim 1, characterized in that The converting the front-to-back distance information in the three-dimensional object data into sound volume representation dimensional data, so as to represent the front-to-back distance information of the corresponding pixel position of the physical object through the sound volume of each single-frequency sound, includes: Performing a logarithmic calculation on the front-to-back distance information in the three-dimensional data of the object to obtain a logarithmic calculation result, and based on the logarithmic calculation result, converting the front-to-back distance information into a volume value using the following sound volume-to-front-to-back distance correspondence conversion formula: A=30 / X 1 / 3 ; Wherein, X is the front-to-back distance information, and the calculation result of the reciprocal function of the X power operation represents the volume value A corresponding to the front-to-back distance information. The unit of the volume value A is dB, and A represents the average front-to-back distance corresponding to the physical object contained in a pixel window. The front-to-back distance information of the pixel position corresponding to the physical object contained in the pixel window is represented by the volume value of each single-frequency sound.

3. The method according to claim 1, characterized in that The blind person's mobility assist glasses type device corresponds to a plurality of rows and a plurality of columns; The projecting of the environmental information image into a plurality of pixel windows formed by a plurality of rows and a plurality of columns comprises: Testing the user's auditory recognition ability to obtain a test result of the user's auditory recognition ability; wherein the test result of the user's auditory recognition ability includes a test result of the user's auditory recognition ability of sound frequency; Adjusting the number of the plurality of rows and the number of the plurality of columns according to the test result of the user's auditory recognition ability of sound frequencies to obtain a pixel window fineness adjustment result; A plurality of pixel windows formed by a plurality of rows and a plurality of columns are determined according to the pixel window fineness adjustment result, and the environmental information image is orthographically projected into the plurality of pixel windows formed by the plurality of rows and a plurality of columns.

4. The method according to claim 3, characterized in that The user's auditory recognition ability test result further includes the user's auditory recognition ability test result for sound reception; the sequentially playing the plurality of mixed sound signals in the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are horizontally continuous includes: adjusting the sound output length of each of the mixed sound signals according to the test result of the user's auditory recognition ability of sound reception, to obtain a sound output length adjustment result; According to the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are continuous in the horizontal direction, and according to the sound output length corresponding to the sound output length adjustment result, the multiple mixed sound signals are played in sequence.

5. The method according to claim 1, wherein Playing the plurality of mixed sound signals in sequence according to the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are horizontally continuous includes: During actual playback, a 1000 Hz head positioning signal of a first specified number of seconds is played before the multiple mixed sound signals corresponding to each frame of the environmental information image are played. After the head positioning signal is played, the multiple mixed sound signals are played sequentially according to the order in which the vertically continuous pixel windows in each column of the up, down, left, and right position information are horizontally continuous, with each column of the pixel windows having a sound output length of a second specified number of seconds, so as to synthesize and play the multiple mixed sound signals corresponding to multiple frames of the environmental information image through the multiple head positioning signals. The first specified number of seconds is less than the second specified number of seconds.

6. The method according to claim 1, characterized in that The binocular camera is a structure of two sets of dual camera modules, and the two sets of dual camera modules correspondingly extract two sets of environmental information images, and the two sets of environmental information images represent the left field of view environmental information image and the right field of view environmental information image of the user; Playing the plurality of mixed sound signals in sequence according to the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are horizontally continuous includes: Playing, through the left audio channel of the blind person mobility assist glasses-type device, the plurality of mixed sound signals corresponding to the left visual field environment information image in sequence according to the order in which the vertically continuous pixel windows in each column in the up, down, left, and right position information are horizontally continuous; Through the right channel of the blind person's mobility assistance glasses type device, the multiple mixed sound signals corresponding to the right field of view environmental information image are played in sequence according to the order of the left and right horizontal continuity of the vertically continuous pixel windows in each column in the up, down, left and right position information.

7. A three-dimensional space obstacle perception data processing device, characterized in that: Applicable to a blind person's mobility assistance glasses-type device, the blind person's mobility assistance glasses-type device is provided with a binocular camera, and the blind person's mobility assistance glasses-type device is worn on the user's head; the device includes: an extraction module, configured to extract, through the binocular camera, an environmental information image corresponding to the direction in which the user's head is facing; wherein the environment corresponding to the environmental information image contains a plurality of physical objects; a recognition module, configured to orthographically project the environmental information image onto a plurality of pixel windows formed by a plurality of rows and columns, determine up-down, down-down, left-down, and right-down position information of the plurality of physical objects in the environmental information image; perform three-dimensional spatial recognition of the plurality of physical objects based on the up-down, down-down, left-down, and right-down position information and the sizes of the plurality of physical objects in the plurality of pixel windows, determine front-back distance information of the plurality of physical objects in the plurality of pixel windows; and generate three-dimensional object data corresponding to the plurality of physical objects based on the up-down, down-down, left-down, and right-down position information and the front-back distance information; A conversion module is used to convert the up and down position information in the up and down, left and right position information in the three-dimensional data of the object into sound frequency representation dimensional data, so as to represent a column of pixel windows vertically continuous from top to bottom in the up and down, left and right position information by different single-frequency sounds; the higher the pixel position in the up and down, left and right position information, the lower the frequency corresponding to the single-frequency sound, and the lower the pixel position in the up and down, left and right position information, the higher the frequency corresponding to the single-frequency sound; convert the front and back distance information in the three-dimensional data of the object into sound volume representation dimensional data, so as to represent the front and back distance information of the corresponding pixel position of the physical object by the sound volume of each single-frequency sound; the closer the front and back distance information corresponds to, the louder the sound volume, the farther the front and back distance, the smaller the sound volume, the closer the front and back distance, the faster the sound volume decreases, and the front and back distance The farther away, the slower the sound volume decreases; a mixed sound signal is generated based on the different single-frequency sounds and the sound volume of each single-frequency sound using a sound wave mixer algorithm, and the mixed sound signal represents the up and down position information and the front and back distance information corresponding to a column of vertically continuous pixel windows; according to the order in which the vertically continuous pixel windows in each column of the up and down, left and right position information are continuous in the left and right directions, a plurality of the mixed sound signals are played in sequence, so that the plurality of mixed sound signals played in sequence represent the comprehensive three-dimensional information corresponding to all the pixel windows, and the comprehensive three-dimensional information includes the up and down, left and right position information and the front and back distance information; wherein, the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data and sound playback time representation dimension data, and each of the three spatial dimension data corresponds to the conversion of one type of sound representation dimension data; A playback module is used to play sound through the blind person's mobility assist glasses-type device based on the target sound frequency representation dimension data, target sound volume representation dimension data, and target sound playback time representation dimension data converted into the three spatial dimension data, so that the user can perceive the position and outline of the multiple physical objects in the environment through the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data.

8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Wearable blind assisting system and method for converting image into sound

    CN111862932A

  • Intelligent blind-assisting glasses system capable of realizing stereoscopic perception of environment

    CN113050917A