Three-dimensional space obstacle sensing data processing method and device and electronic equipment

The three-dimensional data is generated through binocular cameras and graph recognition algorithms and converted into sound representations, which solves the problem of low perception accuracy in blind walking glasses, and realizes the precise perception and autonomous action of blind people about the environment.

CN120279495AActive Publication Date: 2025-07-08NANJING TAIRUI HUIZHI ELECTRONIC INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510758295.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The existing blind walking glasses products cannot allow blind people to accurately perceive the specific situation of their surroundings, and their perception accuracy is low.

Method used

The environment information images are extracted through binocular cameras, and three-dimensional spatial recognition algorithms are used to generate three-dimensional three-dimensional data of the object, and converted them into dimensional data for sound frequency, volume and playback time. The sound playback is performed through blind people's actions assisted glasses, so that blind people can perceive the contours and locations of obstacles in the environment.

Benefits of technology

It improves the accuracy of blind people's perception of environmental information, realizes the ability to perceive three-dimensional space obstacles, and enhances the blind people's movement autonomy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279495A_ABST
    Figure CN120279495A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional space obstacle perception data processing method and device and electronic equipment, relates to the technical field of electronics, and solves the technical problem that blind people have low perception accuracy on environment information through existing blind people walking-aid glasses. The method comprises the following steps: extracting an environment information image corresponding to the facing direction of the head of a user through a binocular camera; performing three-dimensional space identification on the plurality of entity objects through a graph identification algorithm based on the environment information image, and generating object three-dimensional data corresponding to the plurality of entity objects; respectively converting three pieces of spatial dimension data in the three-dimensional data of the object into sound representation dimension data; and based on the target sound frequency representation dimension data, the target sound volume representation dimension data and the target sound playing time representation dimension data which are correspondingly converted from the three spatial dimension data, performing sound playing through the blind person action assisting glasses type equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electronic technologies, and in particular, to a method and apparatus for processing three-dimensional space obstacle perception data, and an electronic device. Background Art

[0002] At present, the blind assistive glasses products on the market can perform functions such as environmental reminder, obstacle approach reminder, object recognition, and text reading based on the video recognition ability of cameras or sonic ranging radars, which facilitates the lives of the blind. However, existing devices cannot enable the blind to perceive the specific situation of the surrounding environment. They are all voice reminder-based and non-directional approaching alarm sounds. These blind assistive glasses products in the prior art have limited use and are not comprehensive enough, resulting in a low perception accuracy of environmental information for the blind. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and apparatus for processing three-dimensional space obstacle perception data, and an electronic device, so as to solve the technical problem that the perception accuracy of environmental information by the existing blind assistive glasses for the blind is relatively low.

[0004] In a first aspect, the present application provides a method for processing three-dimensional space obstacle perception data, which is applied to a device of the blind action assistance glasses type. A binocular camera is provided on the device of the blind action assistance glasses type, and the device of the blind action assistance glasses type is worn on the user's head; the method includes: Extracting an environmental information image corresponding to the direction faced by the user's head through the binocular camera; the environment corresponding to the environmental information image contains several entity objects closest to the user; Performing three-dimensional space recognition on the several entity objects through a graphic recognition algorithm based on the environmental information image, and generating object three-dimensional stereoscopic data corresponding to the several entity objects; Converting the three spatial dimension data in the object three-dimensional stereoscopic data into sound representation dimension data respectively; the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, and each spatial dimension data in the three spatial dimension data corresponds to converting one of the sound representation dimension data; Based on the target sound frequency representation dimension data, target sound volume representation dimension data, and target sound playback time representation dimension data corresponding to the conversion of the three spatial dimension data, performing sound playback through the device of the blind action assistance glasses type, so that the user perceives the contour of the several entity objects in the environment and the relative position relative to the user through the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data.

[0005] In a possible implementation, the three-dimensional spatial recognition of the several entity objects is performed on the environmental information image through a graphic recognition algorithm to generate the three-dimensional solid data of the objects corresponding to the several entity objects, including: The environmental information image is orthogonally projected onto a plurality of pixel windows formed by multiple rows and multiple columns to determine the up, down, left, and right position information of the several entity objects in the environmental information image; Based on the up, down, left, and right position information and the sizes presented by the several entity objects in the plurality of pixel windows, the three-dimensional spatial recognition of the several entity objects is performed to determine the front-back distance information of the several entity objects in the plurality of pixel windows; Based on the up, down, left, and right position information and the front-back distance information, the three-dimensional solid data of the objects corresponding to the several entity objects is generated.

[0006] In a possible implementation, the conversion of the three spatial dimension data in the three-dimensional solid data of the object into sound representation dimension data respectively includes: The up-down position information in the up, down, left, and right position information in the three-dimensional solid data of the object is converted into the sound frequency representation dimension data to represent a column of pixel windows that are vertically continuous from top to bottom in the up, down, left, and right position information through different single-frequency sounds; wherein, the higher the pixel position in the up, down, left, and right position information, the lower the frequency corresponding to the single-frequency sound, and the lower the pixel position in the up, down, left, and right position information, the higher the frequency corresponding to the single-frequency sound; The front-back distance information in the three-dimensional solid data of the object is converted into the sound volume representation dimension data to represent the front-back distance information of the pixel position corresponding to the entity object through the sound volume of each single-frequency sound; wherein, the closer the front-back distance corresponding to the front-back distance information, the greater the sound volume, the farther the front-back distance, the smaller the sound volume, the faster the sound volume drops when the front-back distance is closer, and the slower the sound volume drops when the front-back distance is farther; Based on the different single-frequency sounds and the sound volume of each single-frequency sound, a mixed sound signal is generated by using a sound wave mixer algorithm, and the mixed sound signal represents the up-down position information and the front-back distance information corresponding to a column of vertically continuous pixel windows; In the order of the pixel windows that are vertically continuous in each column in the up, down, left, and right position information and are horizontally continuous on the left and right, a plurality of the mixed sound signals are sequentially played, so that the sequentially played plurality of mixed sound signals represent the comprehensive three-dimensional information corresponding to all the pixel windows, and the comprehensive three-dimensional information includes the up, down, left, and right position information and the front-back distance information.

[0007] In a possible implementation, converting the front-back distance information in the three-dimensional solid data of the object into sound volume representation dimensional data, and representing the front-back distance information of the corresponding pixel position of the entity object through the sound volume of each single-frequency sound includes: Performing a logarithmic calculation on the front-back distance information in the three-dimensional solid data of the object to obtain a logarithmic calculation result, and based on the logarithmic calculation result, converting the front-back distance information into a volume value through the following conversion formula for the correspondence between sound volume and front-back distance: A = 30 / X 1 / 3 ; where X is the front-back distance information, the calculation result of the reciprocal function of the X power operation represents the volume value A corresponding to the front-back distance information, the unit of the volume value A is dB, A represents the average front-back distance corresponding to the entity object included in each pixel window, and the front-back distance information of the corresponding pixel position of the entity object included in the pixel window is represented by the volume value of each single-frequency sound.

[0008] In a possible implementation, the blind person action assistance glasses type device corresponds to multiple rows and multiple columns; Projecting the environmental information image orthogonally onto a plurality of pixel windows formed by multiple rows and multiple columns includes: Testing the user's auditory recognition ability to obtain a user auditory recognition ability test result; wherein, the user auditory recognition ability test result includes the test result of the user's auditory recognition ability for sound frequency; Adjusting the number of rows of the multiple rows and the number of columns of the multiple columns according to the test result of the user's auditory recognition ability for sound frequency to obtain a pixel window fineness adjustment result; Determining a plurality of pixel windows formed by multiple rows and multiple columns according to the pixel window fineness adjustment result, and projecting the environmental information image orthogonally onto the plurality of pixel windows formed by multiple rows and multiple columns.

[0009] In a possible implementation, the user auditory recognition ability test result further includes the test result of the user's auditory recognition ability for sound reception; playing a plurality of the mixed sound signals in sequence according to the order of the pixel windows that are longitudinally continuous in each column and laterally continuous in the left and right of the up-down-left-right position information includes: Adjusting the sound output length of each mixed sound signal according to the test result of the user's auditory recognition ability for sound reception to obtain a sound output length adjustment result; Play a plurality of the mixed sound signals in sequence according to the order in which the pixel windows that are vertically consecutive in each column of the up-down, left-right position information are horizontally consecutive from left to right, and with the sound output length corresponding to the sound output length adjustment result.

[0010] In a possible implementation, the step of playing a plurality of the mixed sound signals in sequence according to the order in which the pixel windows that are vertically consecutive in each column of the up-down, left-right position information are horizontally consecutive from left to right includes: During actual playback, before playing a plurality of the mixed sound signals corresponding to each frame of the environmental information image, play a 1000 Hz head positioning signal for a first specified number of seconds, and after the head positioning signal is played, play a plurality of the mixed sound signals in sequence according to the order in which the pixel windows that are vertically consecutive in each column of the up-down, left-right position information are horizontally consecutive from left to right, with the sound output length of each column of the pixel windows being a second specified number of seconds, so as to synthesize and play a plurality of the mixed sound signals corresponding to multiple frames of the environmental information image through a plurality of the head positioning signals; Wherein, the first specified number of seconds is less than the second specified number of seconds.

[0011] In a possible implementation, the binocular camera has a structure of two dual-camera modules, and the two dual-camera modules respectively extract two sets of the environmental information images, and the two sets of the environmental information images represent the left visual field environmental information image and the right visual field environmental information image of the user; The step of playing a plurality of the mixed sound signals in sequence according to the order in which the pixel windows that are vertically consecutive in each column of the up-down, left-right position information are horizontally consecutive from left to right includes: Through the left channel of the blind person action assistance glasses type device, play a plurality of the mixed sound signals corresponding to the left visual field environmental information image in sequence according to the order in which the pixel windows that are vertically consecutive in each column of the up-down, left-right position information are horizontally consecutive from left to right; Through the right channel of the blind person action assistance glasses type device, play a plurality of the mixed sound signals corresponding to the right visual field environmental information image in sequence according to the order in which the pixel windows that are vertically consecutive in each column of the up-down, left-right position information are horizontally consecutive from left to right.

[0012] In a second aspect, the present application provides a three-dimensional space obstacle perception data processing device, which is applied to a blind person action assistance glasses type device, and a binocular camera is provided on the blind person action assistance glasses type device, and the blind person action assistance glasses type device is worn on the head of a user; the device includes: An extraction module, configured to extract an environmental information image corresponding to the direction that the head of the user faces through the binocular camera; wherein, a plurality of entity objects are included in the environment corresponding to the environmental information image. An identification module, configured to perform three-dimensional space identification on the plurality of entity objects through a graphic recognition algorithm based on the environmental information image, and generate three-dimensional solid data of the object corresponding to the plurality of entity objects; A conversion module, configured to convert the three spatial dimension data in the three-dimensional solid data of the object into sound representation dimension data respectively; wherein, the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, and each of the three spatial dimension data corresponds to one of the sound representation dimension data for conversion; A playback module, configured to perform sound playback through the blind action assistance glasses type device based on the target sound frequency representation dimension data, target sound volume representation dimension data, and target sound playback time representation dimension data corresponding to the three spatial dimension data, so that the user can perceive the positions and contours of the plurality of entity objects in the environment through the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data.

[0013] In a third aspect, the present application further provides an electronic device, including a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the method described in the first aspect above is implemented.

[0014] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and run by the processor, the computer-executable instructions cause the processor to run the method described in the first aspect above.

[0015] The present application brings the following beneficial effects: A three-dimensional space obstacle perception data processing method, device, and electronic device provided by the present application can extract an environmental information image corresponding to the direction faced by the user's head through a binocular camera; several entity objects closest to the user are included in the environment corresponding to the environmental information image; three-dimensional space recognition is performed on the several entity objects through a graphic recognition algorithm based on the environmental information image to generate object three-dimensional solid data corresponding to the several entity objects; the three spatial dimension data in the object three-dimensional solid data are respectively converted into sound representation dimension data; the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, and each of the three spatial dimension data corresponds to the conversion of one kind of sound representation dimension data; based on the target sound frequency representation dimension data, target sound volume representation dimension data, and target sound playback time representation dimension data corresponding to the conversion of the three spatial dimension data, sound playback is performed through a blind person action assistance glasses type device, so that the user can perceive the outlines of the several entity objects in the environment and their relative positions relative to the user through the target sound frequency representation dimension data, target sound volume representation dimension data, and target sound playback time representation dimension data. In this solution, through the sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, a blind person can perceive the real object three-dimensional solid data of each entity object in the environment, realizing a more comprehensive perception of the environmental information in all directions (front, back, left, right, up, and down) by the blind person, enabling the blind person to obtain the three-dimensional space obstacle perception ability, and further presenting a continuously information-rich picture perception, that is, providing the blind person with the space obstacle perception ability, improving the perception accuracy of the environmental information by the blind person, and solving the technical problem that the perception accuracy of the environmental information by the blind person through the existing blind person walking assistance glasses is relatively low.

[0016] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flowchart of the three-dimensional space obstacle perception data processing method provided by the embodiment of the present application; Figure 2This is an example of a device of the blind person action assistance glasses type in the 3D space obstacle perception data processing method provided by the embodiments of the present application; Figure 3 This is an example of an environmental information image in the 3D space obstacle perception data processing method provided by the embodiments of the present application; Figure 4 This is an example of a 3D object contour in the 3D space obstacle perception data processing method provided by the embodiments of the present application; Figure 5 This is an example of the up, down, left, right position information and the front and back distance information corresponding to multiple pixel windows in the 3D space obstacle perception data processing method provided by the embodiments of the present application; Figure 6 This is an example of the comparison between the sound frequency and the sound volume corresponding to multiple pixel windows in the 3D space obstacle perception data processing method provided by the embodiments of the present application; Figure 7 This is a schematic structural diagram of a 3D space obstacle perception data processing device provided by the embodiments of the present application; Figure 8 The figure shows a schematic structural diagram of an electronic device provided by the embodiments of the present application. Detailed implementation manners

[0019] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0020] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0021] At present, the unsolved problem has been how to enable blind people to quickly recognize the visualization of the environment in their minds and gain the ability to move freely. For the application requirements, the current devices provided for blind people are as follows: using a camera to recognize the surrounding environment; using an ultrasonic probe to scan for obstacles in the forward direction; relying on voice reminders, such as: "There is an obstacle 1 meter to the left, detour to the right"; there is also a method of graphically itemizing the camera images, then using an electrode matrix on the tongue to convert the image pixels into electrode stimulation voltages, which are sensed by the tongue and the human brain reflects the external picture. Currently, after testing, it is available, proving that the human body can form images in the brain through various perceptions.

[0022] However, the existing devices cannot enable blind people to perceive the specific situation of the surrounding environment, and they are all voice reminder-based and non-directional proximity alarm sounds. Moreover, although the above-mentioned perception on the tongue can be graphically perceived, it is only a two-dimensional image and is restricted in use and inconvenient and basically not available. Therefore, through the existing blind people's walking assistance glasses, the perception accuracy of blind people for environmental information is relatively low.

[0023] Based on this, the embodiments of the present application provide a three-dimensional space obstacle perception data processing method, device, and electronic device, which can solve the technical problem that the perception accuracy of blind people for environmental information is relatively low through the existing blind people's walking assistance glasses.

[0024] The embodiments of the present invention will be further introduced below with reference to the accompanying drawings.

[0025] Figure 1 It is a schematic flowchart of a three-dimensional space obstacle perception data processing method provided by the embodiments of the present application. Among them, this method is applied to a device of the blind people's action assistance glasses type, and a binocular camera is set on the device of the blind people's action assistance glasses type, and the device of the blind people's action assistance glasses type is worn on the user's head. As Figure 1 shown, this method includes: Step S110, extracting an environmental information image corresponding to the direction faced by the user's head through the binocular camera.

[0026] Among them, the environment corresponding to the environmental information image contains several entity objects closest to the user. As an example, using the camera module technology on the device of the blind people's action assistance glasses type as Figure 2 shown, two groups of three-dimensional space scene information extraction are completed by using two groups of binocular cameras, and the position information of obstacles in the environment can be obtained through the double camera modules.

[0027] Step S120, performing three-dimensional space recognition on several entity objects based on the environmental information image through a graphic recognition algorithm, and generating object three-dimensional stereo data corresponding to the several entity objects.

[0028] In an alternative embodiment, the environmental stereo image (environmental information image) can be obtained by a three-dimensional contour extraction algorithm to calculate the distance to the nearest obstacle in each pixel window. For example, as Figure 5 shown, the obstacles in the field of view are decomposed into 8×8 pixel windows, and the distance information of each pixel window is calculated.

[0029] Exemplarily, this step S120 may specifically include the following steps: orthogonally project the environmental information image onto a plurality of pixel windows formed by multiple rows and multiple columns to determine the up, down, left, and right position information of several entity objects in the environmental information image; perform three-dimensional space recognition on the several entity objects according to the up, down, left, and right position information and the sizes presented by the several entity objects in the plurality of pixel windows to determine the front-to-back distance information of the several entity objects in the plurality of pixel windows; generate three-dimensional solid data of the object corresponding to the several entity objects based on the up, down, left, and right position information and the front-to-back distance information. Through this data processing method, the generated three-dimensional solid data of the object is more accurate.

[0030] For example, as Figure 3 and Figure 4 shown, the environmental information image extracted by the binocular camera is processed by a pattern recognition algorithm to form a three-dimensional solid contour for subsequent use of voice to express three-dimensional image information. Exemplarily, as Figure 5 shown, a three-dimensional image is orthogonally projected onto 64 pixel windows with 8 rows and 8 columns, presenting the up, down, left, and right position information, and then marking the distance information of these 8×8 pixel windows to describe the three-dimensional contour information of the environment.

[0031] Step S130, convert the three spatial dimension data in the three-dimensional solid data of the object into voice representation dimension data respectively.

[0032] Among them, the voice representation dimension data includes voice frequency representation dimension data, voice volume representation dimension data, and voice playback time representation dimension data, and each spatial dimension data in the three spatial dimension data corresponds to converting one kind of voice representation dimension data. As an alternative embodiment, this step S130 may specifically include the following steps: Convert the up-and-down position information in the up, down, left, and right position information in the three-dimensional solid data of the object into voice frequency representation dimension data to represent a column of pixel windows that are vertically continuous from top to bottom in the up, down, left, and right position information through different single-frequency voices; among them, the higher the pixel position in the up, down, left, and right position information, the smaller the frequency corresponding to the single-frequency voice, and the lower the pixel position in the up, down, left, and right position information, the larger the frequency corresponding to the single-frequency voice; Convert the front - to - back distance information in the three - dimensional solid data of an object into sound volume - represented dimensional data, so as to represent the front - to - back distance information of the corresponding pixel positions of the entity object through the sound volume of each single - frequency sound; among them, the closer the front - to - back distance information is, the greater the sound volume is, and the farther the front - to - back distance information is, the smaller the sound volume is. Also, the closer the front - to - back distance information is, the faster the frequency of the single - frequency sound decreases from top to bottom, and the farther the front - to - back distance information is, the slower the frequency of the single - frequency sound decreases from top to bottom. Generate a mixed sound signal using the sound wave mixer algorithm based on different single - frequency sounds and the sound volume of each single - frequency sound. The mixed sound signal represents the up - and - down position information and the front - to - back distance information corresponding to a column of longitudinally continuous pixel windows. Play multiple mixed sound signals in sequence according to the left - to - right and top - to - bottom order of each column of longitudinally continuous pixel windows in the up - down and left - right position information, so that the multiple mixed sound signals played in sequence represent the comprehensive three - dimensional information corresponding to all pixel windows. The comprehensive three - dimensional information includes up - down, left - right position information and front - to - back distance information.

[0033] Through this data - processing method, the converted sound - represented dimensional data is more convenient for the blind to perceive, improving the blind's perception efficiency and convenience. In practical applications, most blind people have more developed hearing. In terms of frequency discrimination, ordinary people can distinguish 2 pitches, while blind people can distinguish more than 8 pitches. Using this ability, when transmitting the three - dimensional contour information of the environment to the blind, such as Figure 5 and Figure 6 shown, 8 different single - frequency sounds are used to group - represent a column of pixels from top to bottom. The volume size of each single - frequency sound represents the distance of the corresponding pixel position, with the volume being large for near and small for far, and the frequency decreasing quickly for the near end and slowly for the far end. These 8 single - frequency sounds use the sound wave mixer algorithm to generate a mixed sound signal, and this sound depicts the position and distance information in a longitudinal two - dimensional contour section of three - dimensional space. Sequentially playing 8 such sound information depicts a three - dimensional contour containing position and distance information. Through the picture - processing function in the brain, it is stereoscopically visualized, enabling the blind to have a certain ability of accurate stereoscopic visualization. Using this method, after training, the blind can accurately obtain the surrounding environmental information, and can move freely through the crowd, cross the road, go down the steps, and accurately pick up items such as cups, computers, and mobile phones indoors.

[0034] In an alternative embodiment, the conversion of the front - to - back distance information in the three - dimensional solid data of an object into sound volume - represented dimensional data, so as to represent the front - to - back distance information of the corresponding pixel positions of the entity object through the sound volume of each single - frequency sound, may specifically include the following steps: Perform a logarithmic calculation on the front-back distance information in the three-dimensional solid data of the object to obtain the logarithmic calculation result. Based on the logarithmic calculation result, convert the front-back distance information into a volume value through the following conversion formula for the correspondence between sound volume and front-back distance: A = 30 / X 1 / 3 ; where X is the front-back distance information, the calculation result of the reciprocal function of the X power operation represents the volume value A corresponding to the front-back distance information, the unit of the volume value A is dB, and A represents the average front-back distance corresponding to the entity object contained in a pixel window, so as to represent the front-back distance information of the pixel position corresponding to the entity object contained in the pixel window through the volume value of each single-frequency sound. Through this data processing method, the conversion data of the volume value is made more accurate.

[0035] For the pixel window distance value after taking the value, it can be converted into a sound element. Specifically, perform a logarithmic calculation on the distance information and then convert it into a volume value, which represents the average distance of the obstacles within a pixel window.

[0036] For the configuration of the sound, exemplarily, the frequency of the sound can be selected according to the C major scale note table, and these tones are frequency information that has been proven to be clearly recognizable by humans. Of course, other arbitrary frequencies can also be selected, but the recognition of other frequencies has not been verified by experience. For example Figure 5 and Figure 6 as shown, the 8 pixel positions can respectively take 8 tones of 523Hz, 587Hz, 659Hz, 698Hz, 740Hz, 784Hz, 880Hz, and 988Hz as the row tones. When more rows of corresponding frequencies are needed (such as making a 32×32 pixel window), they can continue to be taken from the C major scale note table. The 8 tones correspond to representing the positions of 8 consecutive windows from top to bottom. As shown in Figure 5 and Figure 6 as shown, the distance at each position is represented by the corresponding volume amplitude. 60dB represents a distance of 0.1 meter. Using this decibel sound is basically the clear sound amplitude that people are used to hearing. The system provides volume adjustment, and the user can also adjust the volume according to the environment and habits. In the case of 8×8, the calculated accuracy of the contour of the nearest recognizable object is 1.25 cm. The relationship between the distance and the sound decibel is as follows: A (dB) = 30 / X 1 / 3 . The conversion of distance and decibel is shown in Table 1 below: Table 1

[0037] As an alternative implementation, the above-mentioned conversion of the vertical position information in the three-dimensional object data into the dimension data represented by the sound frequency, so as to represent a column of pixel windows that are vertically continuous from top to bottom in the left-right and up-down position information by different single-frequency sounds, specifically includes the following steps: Convert the vertical position information in the left-right and up-down position information of the three-dimensional object data into the dimension data represented by the sound frequency through the following conversion formula, so as to represent a column of pixel windows that are vertically continuous from top to bottom in the left-right and up-down position information by different single-frequency sounds: f ( p ) = fmax- − × p ; Wherein, f represents the generated sound frequency; p represents the position of the pixel, the higher the position of the pixel, the larger the value, assuming that the position of the top pixel is 0, Pmax is the maximum value of the pixel position; fmin represents the lowest sound frequency limit corresponding to the highest pixel position (i.e., the top); fmax represents the highest sound frequency limit corresponding to the lowest pixel position (i.e., the bottom); f ( p ) is the sound frequency corresponding to the given pixel position p . Through this data processing method, the conversion of the dimension data represented by the sound frequency is made more accurate.

[0038] In a possible implementation, the above-mentioned generation of the mixed sound signal by using the sound wave mixer algorithm based on different single-frequency sounds and the sound volume of each single-frequency sound may specifically include the following steps: Generate the mixed sound signal by using the sound wave mixer algorithm based on different single-frequency sounds and the sound volume of each single-frequency sound through the following formula: S ( t ) = ; Wherein, S ( t ) represents the finally generated mixed sound signal, N is the number of single-frequency sounds participating in the mixing, Ai is the amplitude of the i-th sound signal, fi is the frequency of the i-th sound signal, ϕi is the initial phase of the i-th sound signal, and t is time. Through this data processing method, the data of the generated mixed sound signal is made more accurate.

[0039] The following is in accordance with Figure 5The calculation results of the frequency and amplitude information of the sound information are given as an example: the volume converted from the closest end surface distance value of the three objects is: 0.25 meters = 44.44dB...1.75 meters = 23.23dB...2.25 meters = 21.37dB. As shown in Table 2, the following example shows Figure 3 The sound synthesis component of the information: Table 2

[0040] As shown in Table 3, the following example shows a C key note and frequency comparison table: Table 3

[0041] As an optional implementation, the blind person's mobility assist glasses type device corresponds to a plurality of rows and a plurality of columns; the above-mentioned positive projection of the environmental information image into a plurality of pixel windows formed by the plurality of rows and the plurality of columns may specifically include the following steps: Testing the user's auditory recognition ability to obtain a test result of the user's auditory recognition ability; wherein the test result of the user's auditory recognition ability includes a test result of the user's auditory recognition ability of sound frequency; According to the test results of the user's auditory recognition ability of sound frequency, the number of rows and the number of columns are adjusted to obtain the pixel window fineness adjustment result; according to the pixel window fineness adjustment result, multiple pixel windows formed by multiple rows and multiple columns are determined, and the environmental information image is projected into the multiple pixel windows formed by the multiple rows and multiple columns.

[0042] Through this data processing method, the fineness of the pixel window is more consistent with and fits the personal hearing recognition of blind users, thereby improving the experience of blind users.

[0043] In actual applications, the specific recognition capabilities can be adjusted according to the user's hearing ability, for example: 32×32, 24×24, 16×16, 4×4 and other pixel window accuracies. Blind people who use it for a long time can establish the length, width, height, distance, openness, narrowness, sidewalk boundaries and other spatial environment concepts of ordinary people.

[0044] In an optional implementation, the user's auditory recognition ability test result further includes the user's auditory recognition ability test result for sound reception; the above-mentioned playing of multiple mixed sound signals in sequence according to the order in which each column of vertically continuous pixel windows in the up, down, left, and right position information is horizontally continuous in the left and right directions may specifically include the following steps: Adjust the sound output length of each mixed sound signal according to the test results of the user's auditory recognition ability for sound reception, and obtain the sound output length adjustment result; play multiple mixed sound signals in sequence according to the order of the vertically consecutive pixel windows in each column of the up-down-left-right position information in the left-right horizontal continuity, and with the sound output length corresponding to the sound output length adjustment result.

[0045] Through this data processing method, the played sound better conforms to and fits the individual hearing conditions of blind users, improving the experience of blind users.

[0046] In the embodiment of the present application, the specific recognition parameters of the device can be adjusted separately according to the user's auditory ability. For example, the output length of each mixed sound is adjusted according to the user's auditory ability.

[0047] As an optional implementation manner, the above-mentioned playing multiple mixed sound signals in sequence according to the order of the vertically consecutive pixel windows in each column of the up-down-left-right position information in the left-right horizontal continuity may specifically include the following steps: When actually playing, before playing the multiple mixed sound signals corresponding to each frame of environmental information image, play a head positioning signal of 1000 Hz for the first specified number of seconds, and after the head positioning signal is played, play multiple mixed sound signals in sequence according to the order of the vertically consecutive pixel windows in each column of the up-down-left-right position information in the left-right horizontal continuity, with the sound output length of each column of pixel windows being the second specified number of seconds, so as to synthesize and play the multiple mixed sound signals corresponding to multiple frames of environmental information images through multiple head positioning signals; wherein, the first specified number of seconds is less than the second specified number of seconds.

[0048] For example, when actually playing, add a 1000 Hz head positioning signal of 1 / 64 second in front of each image information, and then output the mixed sound to the left and right earphones in sequence with a length of 1 / 32 second each, so as to synthesize the depth of field information of two 8×8 pixel windows, and realize three-dimensional imaging through the image processing function in the brain, enabling the blind to have a certain precise three-dimensional imaging visual ability.

[0049] In an optional implementation manner, the binocular camera has a structure of two double camera modules. The two double camera modules respectively extract two sets of environmental information images, and the two sets of environmental information images represent the left visual field environmental information image and the right visual field environmental information image of the user; the above-mentioned playing multiple mixed sound signals in sequence according to the order of the vertically consecutive pixel windows in each column of the up-down-left-right position information in the left-right horizontal continuity may specifically include the following steps: Play multiple mixed sound signals corresponding to the left visual field environmental information image in sequence through the left channel of the blind action assistance glasses type device according to the order of the vertically consecutive pixel windows in each column of the up-down-left-right position information in the left-right horizontal continuity; Through the right channel of the blind person's action assistance glasses type device, in the order of the left-right horizontal continuity of each column of vertically continuous pixel windows in the up-down-left-right position information, a plurality of mixed sound signals corresponding to the right visual field environment information image are sequentially played.

[0050] In the embodiment of the present application, the sound wave information of the left and right channels is generated from the results of the two groups of acquisitions and played by the earphones for the blind to use. In this way, the blind can obtain more comprehensive and rich environmental information.

[0051] Step S140, based on the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data corresponding to the conversion of the three spatial dimension data, perform sound playback through the blind person's action assistance glasses type device, so that the user can perceive the outlines of several entity objects in the environment and their relative positions relative to the user through the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data.

[0052] In the embodiment of the present application, through the sound frequency representation dimension data, the sound volume representation dimension data, and the sound playback time representation dimension data, the blind can perceive the real three-dimensional object data of each entity object in the environment, realizing a more comprehensive perception of the environment in the front, back, left, right, up, and down by the blind, enabling the blind to obtain the three-dimensional space obstacle perception ability, and then presenting a continuous picture perception of information, that is, providing the blind with the space obstacle perception ability, improving the perception accuracy of the blind for environmental information, and also enhancing the application scenario versatility of the blind person's action assistance glasses type device.

[0053] Of course, in addition to depicting the relative position of the outline with the blind, the central camera can also complete environmental recognition (especially recognizing environmental information such as locations, road surface positions, traffic lights, road signs, road plaques, and signboards and other text identifications with reference to GPS and electronic maps), face recognition (seeing acquaintances on the road, or finding a person who has made an appointment to meet at a designated location), and reading and map reading functions (two-dimensional pictures can also be expressed by sound waves, but it will take a long time). The blind can also interact with the system using voice, buttons, and rollers, asking the system to make adjustments or answer questions.

[0054] In an alternative embodiment, the audio converted from the spatial information can also be shared with other blind people, sharing the surrounding environmental information from the perspective of the sender. Even this technology can provide special movies for the blind, of course, this movie is viewed from the first perspective.

[0055] Figure 7A structural schematic diagram of a three-dimensional space obstacle perception data processing device is provided. The device can be applied to a blind person's action assistance glasses type device, and a binocular camera is provided on the blind person's action assistance glasses type device, and the blind person's action assistance glasses type device is worn on the user's head. As Figure 7 shown, the three-dimensional space obstacle perception data processing device 700 includes: An extraction module 701, configured to extract an environmental information image corresponding to the direction faced by the user's head through the binocular camera; wherein, a plurality of entity objects are included in the environment corresponding to the environmental information image; An identification module 702, configured to perform three-dimensional space identification on the plurality of entity objects through a graphic recognition algorithm based on the environmental information image, and generate object three-dimensional stereo data corresponding to the plurality of entity objects; A conversion module 703, configured to respectively convert three spatial dimension data in the object three-dimensional stereo data into sound representation dimension data; wherein, the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, and each of the three spatial dimension data corresponds to converting one of the sound representation dimension data; A playback module 704, configured to perform sound playback through the blind person's action assistance glasses type device based on the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data corresponding to the three spatial dimension data, so that the user can perceive the positions and contours of the plurality of entity objects in the environment through the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data.

[0056] The three-dimensional space obstacle perception data processing device provided by the embodiments of the present application has the same technical features as the three-dimensional space obstacle perception data processing method provided by the above embodiments, so it can also solve the same technical problems and achieve the same technical effects.

[0057] An electronic device provided by an embodiment of the present application, as Figure 8 shown, the electronic device 800 includes a processor 802 and a memory 801. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the steps of the method provided by the above embodiments are implemented.

[0058] See Figure 8 , the electronic device further includes: a bus 803 and a communication interface 804. The processor 802, the communication interface 804, and the memory 801 are connected through the bus 803; the processor 802 is configured to execute an executable module stored in the memory 801, such as a computer program.

[0059] Among them, the memory 801 may include high-speed random access memory (Random Access Memory, referred to as RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between this system network element and at least one other network element is realized through at least one communication interface 804 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0060] The bus 803 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of easy representation, Figure 8 only a single bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0061] Among them, the memory 801 is used to store programs. After receiving an execution instruction, the processor 802 executes the programs. The methods executed by the devices defined by the processes disclosed in any of the foregoing embodiments of the present application can be applied to or implemented by the processor 802.

[0062] The processor 802 may be an integrated circuit chip with the ability to process signals. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 802 or the instructions in the form of software. The above-mentioned processor 802 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 801, and the processor 802 reads the information in the memory 801 and combines its hardware to complete the steps of the above method.

[0063] Corresponding to the above three-dimensional space obstacle perception data processing method, an embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and run by a processor, the computer-executable instructions cause the processor to run the steps of the above three-dimensional space obstacle perception data processing method.

[0064] The three-dimensional space obstacle perception data processing device provided by the embodiments of the present application may be specific hardware on a device or software or firmware installed on the device, etc. For the device provided by the embodiments of the present application, the implementation principle and the technical effects produced are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the above method embodiments, and will not be repeated here.

[0065] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0066] For another example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0067] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0068] In addition, the functional units in the embodiments provided in the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0069] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the three-dimensional space obstacle perception data processing method described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks, or optical discs.

[0070] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.

[0071] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of this application, used to illustrate the technical solution of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in this application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application. All should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for processing three-dimensional space obstacle perception data, characterized in that Applied to a blind person's action assistance glasses type device, a binocular camera is provided on the blind person's action assistance glasses type device, and the blind person's action assistance glasses type device is worn on the user's head; the method includes: Extracting, by the binocular camera, an environmental information image corresponding to the direction faced by the user's head; the environment corresponding to the environmental information image includes several entity objects closest to the user; Performing three-dimensional space recognition on the several entity objects based on the environmental information image through a graphic recognition algorithm to generate object three-dimensional stereo data corresponding to the several entity objects; Converting the three spatial dimension data in the object three-dimensional stereo data into sound representation dimension data respectively; the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, and each of the three spatial dimension data corresponds to converting one of the sound representation dimension data; Based on the target sound frequency representation dimension data, target sound volume representation dimension data, and target sound playback time representation dimension data corresponding to the three spatial dimension data, performing sound playback through the blind person's action assistance glasses type device, so that the user can perceive the contours of the several entity objects in the environment and their relative positions relative to the user through the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data.

2. The method according to claim 1, characterized in that, The performing three-dimensional space recognition on the several entity objects based on the environmental information image through a graphic recognition algorithm to generate object three-dimensional stereo data corresponding to the several entity objects includes: Orthogonally projecting the environmental information image onto a plurality of pixel windows formed by multiple rows and multiple columns, and determining the up, down, left, and right position information of the several entity objects in the environmental information image; Performing three-dimensional space recognition on the several entity objects according to the up, down, left, and right position information and the sizes presented by the several entity objects in the plurality of pixel windows to determine the front-back distance information of the several entity objects in the plurality of pixel windows; Generating object three-dimensional stereo data corresponding to the several entity objects based on the up, down, left, and right position information and the front-back distance information.

3. The method according to claim 2, characterized in that The converting the three spatial dimension data in the object three-dimensional stereo data into sound representation dimension data respectively includes: Converting the up and down position information in the up, down, left, and right position information in the object three-dimensional stereo data into the sound frequency representation dimension data, so as to represent a column of the pixel windows that are vertically continuous from top to bottom in the up, down, left, and right position information through different single-frequency sounds; wherein, the higher the pixel position in the up, down, left, and right position information, the lower the frequency corresponding to the single-frequency sound, and the lower the pixel position in the up, down, left, and right position information, the higher the frequency corresponding to the single-frequency sound; Convert the front - to - back distance information in the three - dimensional solid data of the object into sound volume representation dimension data, so as to represent the front - to - back distance information of the corresponding pixel position of the entity object through the sound volume of each single - frequency sound; wherein, the closer the front - to - back distance corresponding to the front - to - back distance information is, the greater the sound volume is, the farther the front - to - back distance is, the smaller the sound volume is, the closer the front - to - back distance is, the faster the decrease speed of the sound volume is, and the farther the front - to - back distance is, the slower the decrease speed of the sound volume is; Generate a mixed sound signal based on the different single - frequency sounds and the sound volume of each single - frequency sound by using the sound wave mixer algorithm. The mixed sound signal represents the up - and - down position information and the front - to - back distance information corresponding to a column of longitudinally continuous pixel windows; Play multiple mixed sound signals in sequence according to the order of each column of longitudinally continuous pixel windows in the left - to - right and horizontal - continuous order in the up - down - left - right position information, so that the multiple mixed sound signals played in sequence represent the comprehensive three - dimensional information corresponding to all the pixel windows. The comprehensive three - dimensional information includes the up - down - left - right position information and the front - to - back distance information.

4. The method according to claim 3, characterized in that The conversion of the front - to - back distance information in the three - dimensional solid data of the object into sound volume representation dimension data, so as to represent the front - to - back distance information of the corresponding pixel position of the entity object through the sound volume of each single - frequency sound, includes: Perform a logarithmic calculation on the front - to - back distance information in the three - dimensional solid data of the object to obtain a logarithmic calculation result, and based on the logarithmic calculation result, convert the front - to - back distance information into a volume value through the following conversion formula for the correspondence between sound volume and front - to - back distance: A = 30 / X 1 / 3 ; Wherein, X is the front - to - back distance information, the calculation result of the reciprocal function of the X power operation represents the volume value A corresponding to the front - to - back distance information. The unit of the volume value A is dB, and A represents the average front - to - back distance corresponding to the entity object included in a pixel window, so as to represent the front - to - back distance information of the corresponding pixel position of the entity object included in the pixel window through the volume value of each single - frequency sound.

5. The method according to claim 3, characterized in that, The blind - person action - assisting glasses - type device corresponds to multiple rows and multiple columns; The orthographic projection of the environmental information image onto multiple pixel windows formed by multiple rows and multiple columns includes: Test the user's auditory recognition ability to obtain a user auditory recognition ability test result; wherein, the user auditory recognition ability test result includes the test result of the user's auditory recognition ability for sound frequencies; Adjust the number of rows and the number of columns according to the test result of the user's auditory recognition ability for sound frequencies to obtain a pixel window fineness adjustment result; Determine multiple pixel windows formed by multiple rows and multiple columns according to the pixel window fineness adjustment result, and orthographically project the environmental information image onto the multiple pixel windows formed by multiple rows and multiple columns.

6. The method according to claim 5, characterized in that, The test result of the user's auditory recognition ability also includes the test result of the user's auditory recognition ability for sound reception; playing a plurality of the mixed sound signals in sequence according to the order of the pixel windows that are vertically continuous in each column and horizontally continuous on the left and right in the up, down, left, and right position information, including: Adjusting the sound output length of each of the mixed sound signals according to the test result of the user's auditory recognition ability for sound reception to obtain a sound output length adjustment result; Playing a plurality of the mixed sound signals in sequence according to the order of the pixel windows that are vertically continuous in each column and horizontally continuous on the left and right in the up, down, left, and right position information, and with the sound output length corresponding to the sound output length adjustment result.

7. The method according to claim 3, wherein Playing a plurality of the mixed sound signals in sequence according to the order of the pixel windows that are vertically continuous in each column and horizontally continuous on the left and right in the up, down, left, and right position information, including: When actually playing, before playing a plurality of the mixed sound signals corresponding to each frame of the environmental information image, play a head positioning signal of 1000 Hz for a first specified number of seconds, and after the head positioning signal is played, play a plurality of the mixed sound signals in sequence according to the order of the pixel windows that are vertically continuous in each column and horizontally continuous on the left and right in the up, down, left, and right position information, with the sound output length of each column of the pixel windows being the second specified number of seconds, so as to synthesize and play a plurality of the mixed sound signals corresponding to multiple frames of the environmental information image through a plurality of the head positioning signals; Wherein, the first specified number of seconds is less than the second specified number of seconds.

8. The method according to claim 3, wherein The binocular camera has a structure of two double camera modules, and the two double camera modules respectively extract two sets of the environmental information images, and the two sets of the environmental information images represent the left visual field environmental information image and the right visual field environmental information image of the user; Playing a plurality of the mixed sound signals in sequence according to the order of the pixel windows that are vertically continuous in each column and horizontally continuous on the left and right in the up, down, left, and right position information, including: Playing a plurality of the mixed sound signals corresponding to the left visual field environmental information image in sequence through the left channel of the blind person's action assistance glasses type device according to the order of the pixel windows that are vertically continuous in each column and horizontally continuous on the left and right in the up, down, left, and right position information; Playing a plurality of the mixed sound signals corresponding to the right visual field environmental information image in sequence through the right channel of the blind person's action assistance glasses type device according to the order of the pixel windows that are vertically continuous in each column and horizontally continuous on the left and right in the up, down, left, and right position information.

9. A three-dimensional space obstacle perception data processing device, characterized in that Applied to a blind person's action assistance glasses type device, the blind person's action assistance glasses type device is provided with a binocular camera, and the blind person's action assistance glasses type device is worn on the user's head; the device includes: An extraction module, configured to extract the environmental information image corresponding to the direction faced by the user's head through the binocular camera; wherein, there are several entity objects in the environment corresponding to the environmental information image. An identification module, configured to perform three-dimensional space identification on the plurality of entity objects through a graphic recognition algorithm based on the environmental information image, and generate three-dimensional solid data of the object corresponding to the plurality of entity objects; A conversion module, configured to convert the three spatial dimension data in the three-dimensional solid data of the object into sound representation dimension data respectively; wherein, the sound representation dimension data includes sound frequency representation dimension data, sound volume representation dimension data, and sound playback time representation dimension data, and each of the three spatial dimension data corresponds to one of the sound representation dimension data for conversion; A playback module, configured to perform sound playback through the blind person action assistance glasses type device based on the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data corresponding to the three spatial dimension data, so that the user can perceive the positions and contours of the plurality of entity objects in the environment through the target sound frequency representation dimension data, the target sound volume representation dimension data, and the target sound playback time representation dimension data.

10. An electronic device, comprising a memory and a processor, wherein a computer program that can run on the processor is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8 above.

Citation Information

Patent Citations

  • Auxiliary perception method and system based on sensory substitution

    CN110991336A

  • Wearable blind assisting system and method for converting image into sound

    CN111862932A

  • Intelligent blind-assisting glasses system capable of realizing stereoscopic perception of environment

    CN113050917A

  • Blind guiding glasses and interaction method

    CN118370653A