Blind navigation method, device and medium based on visual-tactile cross-modal perception
By acquiring environmental images through smart glasses to generate depth images and drive the smart belt vibration unit, the problem of insufficient voice prompts in existing blind navigation is solved, and a cross-modal navigation experience and higher navigation accuracy are achieved.
Patent Information
- Application Number
- CN202411522648.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-10-29
AI Technical Summary
Existing navigation methods for the blind rely on voice prompts, which make it difficult to fully and three-dimensionally depict the complex and changing surrounding environment, especially in public places, and cannot provide accurate environmental perception.
A method based on visual-tactile cross-modal perception is used to obtain environmental images through smart glasses, generate depth images using dual cameras, divide regional images and convert them into voltage signals, which drive the vibration unit on the smart belt to navigate in a tactile way.
It achieves precise conversion from vision to touch, provides a cross-modal navigation experience for the blind, and improves the accuracy of navigation and obstacle avoidance.
Smart Images

Figure CN119385753B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of intelligent navigation technology, and in particular to a navigation method, device, and medium for the blind based on visual-tactile cross-modal perception. Background Art
[0002] In related technologies, providing navigation prompts to the blind usually relies on voice prompts. However, this one-dimensional language expression method becomes increasingly limited when faced with four-dimensional environmental information composed of three-dimensional space and time. As a linear expression method, language is difficult to fully and three-dimensionally depict the complex and changing surrounding environment, especially in public places such as stations, supermarkets, and roads. These scenes are often full of rich spatial information and time dynamics. The existing voice prompt navigation method for the blind is difficult to ensure that the blind can provide comprehensive, accurate and reliable environmental perception.
[0003] Therefore, a navigation method for the blind based on visual-tactile cross-modal perception is provided to achieve precise conversion from vision to touch through multi-point vibration design, providing users with a cross-modal navigation experience. Summary of the Invention
[0004] The present invention provides a navigation method, device and medium for the blind based on visual-tactile cross-modal perception, which realizes precise conversion from vision to touch through multi-point vibration design, providing users with a cross-modal navigation experience.
[0005] According to one aspect of the present invention, a method for navigation for the blind based on visual-tactile cross-modal perception is provided. The method is executed by a processor of a smart wearable device, the smart wearable device including smart glasses and a smart belt, the smart belt being provided with at least two vibration units. The method comprises:
[0006] Acquire a first environment image captured by a first camera of the smart glasses, acquire a second environment image captured by a second camera of the smart glasses, and obtain a target depth image based on the first environment image and the second environment image;
[0007] Dividing the target depth image into at least two regional images, determining a minimum depth value of each of the regional images, and converting each of the minimum depth values into a voltage signal of corresponding intensity based on a preset correspondence between depth value magnitude and voltage signal intensity, wherein the regional images correspond to the vibration units;
[0008] Each of the voltage signals is transmitted to the smart belt, and a vibration unit corresponding to the voltage signal is driven to vibrate at the intensity of the voltage signal to perform navigation for the subject wearing the smart wearable device.
[0009] According to another aspect of the present invention, a navigation device for the blind based on visual-tactile cross-modal perception is provided. The device comprises:
[0010] An image acquisition module is configured to acquire a first environment image captured by a first camera of the smart glasses, acquire a second environment image captured by a second camera of the smart glasses, and acquire a target depth image based on the first environment image and the second environment image;
[0011] a signal conversion module, configured to divide the target depth image into at least two regional images, determine a minimum depth value of each of the regional images, and convert each of the minimum depth values into a voltage signal of corresponding intensity based on a preset correspondence between depth value magnitude and voltage signal intensity, wherein the regional images correspond to the vibration units;
[0012] The navigation module is used to transmit each of the voltage signals to the smart belt, drive the vibration unit corresponding to the voltage signal to vibrate at the intensity of the voltage signal, and provide navigation for the person wearing the smart wearable device.
[0013] According to another aspect of the present invention, an electronic device is provided, comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the blind navigation method based on visual-tactile cross-modal perception as described in any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the blind navigation method based on visual-tactile cross-modal perception according to any embodiment of the present invention when executed.
[0018] The technical solution of an embodiment of the present invention obtains a first environmental image captured by a first camera based on smart glasses and a second environmental image captured by a second camera based on the smart glasses. This technical solution, equipped with dual cameras, can capture information about the surrounding environment in real time, obtaining a first environmental image and a second environmental image. A target depth image is then obtained based on the first and second environmental images to provide more accurate spatial information. The target depth image is divided into at least two regional images, and a minimum depth value is determined for each of the regional images. Based on a preset correspondence between the depth value and the voltage signal strength, each minimum depth value is converted into a voltage signal of a corresponding strength, wherein the regional images correspond to the vibration units. Each voltage signal is transmitted to a smart belt, which drives the vibration unit corresponding to the voltage signal to vibrate at the intensity of the voltage signal, thereby navigating the person wearing the smart wearable device. The technical solution of an embodiment of the present invention achieves a precise conversion from visual to tactile sensation through a multi-point vibration design, providing users with a cross-modal navigation experience.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 A flowchart of a navigation method for the blind based on visual-tactile cross-modal perception provided by an embodiment of the present invention;
[0022] Figure 2 A flowchart of a navigation method for the blind based on visual-tactile cross-modal perception provided by an embodiment of the present invention;
[0023] Figure 3 An intelligent wearable device suitable for a blind person navigation method based on visual-tactile cross-modal perception provided by an embodiment of the present invention;
[0024] Figure 4 A flowchart of an optional embodiment of a navigation method for the blind based on visual-tactile cross-modal perception provided by an embodiment of the present invention;
[0025] Figure 5A schematic diagram of an image processing flow applicable to a blind navigation method based on visual-tactile cross-modal perception provided by an embodiment of the present invention;
[0026] Figure 6 A schematic structural diagram of a navigation device for the blind based on visual-tactile cross-modal perception provided by an embodiment of the present invention;
[0027] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0030] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0031] Figure 1 A flow chart of a method for blind navigation based on visual-touch cross-modal perception is provided in an embodiment of the present invention. This embodiment is applicable to situations where navigation is provided for the blind. The method can be performed by a blind navigation device based on visual-touch cross-modal perception. The blind navigation device based on visual-touch cross-modal perception can be implemented in the form of hardware and / or software. The blind navigation device based on visual-touch cross-modal perception can be configured in an electronic device such as a computer or server.
[0032] like Figure 1 As shown, the method of this embodiment includes:
[0033] S110: Acquire a first environment image captured by a first camera of the smart glasses, acquire a second environment image captured by a second camera of the smart glasses, and obtain a target depth image based on the first environment image and the second environment image.
[0034] In an embodiment of the present invention, the smart glasses may include a first camera and a second camera. The first camera and the second camera are respectively located at the left and right lenses of the smart glasses. Optionally, the first camera may be located at the center of the left lens of the smart glasses, and the second camera may be located at the center of the right lens of the smart glasses. The first environmental image may be understood as an environmental image captured by the first camera. The second environmental image may be understood as an environmental image captured by the second camera. Optionally, the first camera and the second camera are RGB cameras. Accordingly, the first environmental image may be an RGB color image, and the second environmental image data may be an RGB color image. In an embodiment of the present invention, the target depth image may be understood as a depth image used for region division. Specifically, a first environmental image captured by the first camera of the smart glasses is obtained, and a second environmental image captured by the second camera of the smart glasses is obtained. After obtaining the first and second environmental images, a depth image generation algorithm may be used to generate a depth image based on the first and second environmental images. The generated depth image may then be used as the target depth image.
[0035] In an embodiment of the present invention, before obtaining the target depth image based on the first environmental image and the second environmental image, the method may further include: if the first environmental image has image distortion, performing image distortion correction processing on the first environmental image; and / or, if the second environmental image has image distortion, performing image distortion correction processing on the second environmental image. In an embodiment of the present invention, performing image distortion correction processing on the first environmental image may be performing image distortion correction processing on the first environmental image using the distortion coefficients in the camera intrinsic parameters of the first camera. Performing image distortion correction processing on the second environmental image may be performing distortion correction processing on the second environmental image using the distortion coefficients in the camera intrinsic parameters of the second camera.
[0036] In the embodiment of the present invention, the purpose of using the distortion coefficients in the camera intrinsic parameters to perform distortion correction processing on the first and second environment images is to eliminate the impact of camera distortion on the images. It should be noted that the camera intrinsic parameters of the first camera and the second camera are the same.
[0037] S120. Divide the target depth image into at least two area images, determine the minimum depth value of each area image, and convert each minimum depth value into a voltage signal of corresponding intensity based on a preset correspondence between the depth value size and the voltage signal strength, wherein the area image corresponds to the vibration unit.
[0038] Among them, the regional image can be an image obtained after dividing the target depth image according to actual needs or preset rules. It can be understood that different regional images correspond to different areas. The number of regional images can be two or more. In practical applications, the number of regional images is usually multiple. Optionally, the regional image can be a fixed geometric shape, such as a rectangle, a circle, etc. In an embodiment of the invention, the target depth image can be divided according to the specification of M×N, where M can be expressed as the number of rows of each divided area, and N can be expressed as the number of columns of each divided area. The minimum depth value can be the minimum value of the depth values of all pixels in each divided regional image. The depth value can represent the distance from the surface of an object in the scene to the first camera and / or the second camera. Therefore, the minimum depth value can represent the distance to the surface of the object (such as an obstacle) closest to the first camera and / or the second camera in the area. In an embodiment of the present invention, the minimum depth value of each of the regional images is determined, specifically, for each of the regional images, the minimum depth value of the regional image is determined according to the depth information of the regional image.
[0039] In an embodiment of the present invention, the preset correspondence between the depth value and the voltage signal strength may be a linear relationship or a nonlinear relationship. In an embodiment of the invention, the preset correspondence between the depth value and the voltage signal strength may also determine the most appropriate conversion curve based on the sensitivity of human perception. In an embodiment of the present invention, the area image corresponds to the vibration unit. In other words, the correspondence between the area image and the vibration unit may be one-to-one. In an embodiment of the present invention, the advantage of one area image corresponding to one vibration unit is that the accuracy of navigation can be improved. Optionally, the vibration unit may be a vibration motor.
[0040] Specifically, after obtaining the target depth image, the target depth image can be divided into at least two regional images according to a preset rule. Thus, for each regional image, a minimum depth value can be determined. Furthermore, based on a preset correspondence between depth values and voltage signal strengths, each minimum depth value can be converted into a voltage signal of a corresponding strength.
[0041] S130: Transmit each of the voltage signals to the smart belt, and drive a vibration unit corresponding to the voltage signal to vibrate at the intensity of the voltage signal to perform navigation for the person wearing the smart wearable device.
[0042] In an embodiment of the present invention, the smart belt may be a flexible smart belt. The smart belt includes at least two vibration units. Optionally, the vibration unit may include a vibration motor and / or a vibration plate. In practical applications, the smart belt usually includes multiple vibration units. In an embodiment of the present invention, multiple vibration units are evenly distributed on the smart belt. For example, the smart belt is divided according to the rule of H×W, and a vibration unit can be set for each area, where H can be expressed as the number of rows of each divided area, and W can be expressed as the number of columns of each divided area. In an embodiment of the present invention, the correspondence between the regional image obtained by dividing the target depth image and the area obtained after the smart belt is divided can be a one-to-one correspondence. In an embodiment of the present invention, the smart belt converts the information in the depth image into an electrical signal through the built-in vibration unit, and drives the vibration unit to vibrate with different intensities, thereby providing intuitive navigation information to the wearer.
[0043] Specifically, after obtaining each voltage signal, each voltage signal can be transmitted to the smart belt. For each voltage signal, a vibration unit corresponding to the voltage signal can be determined. The vibration unit can then be driven to vibrate at the intensity of the voltage signal to navigate the person wearing the smart wearable device. Specifically, determining the vibration unit corresponding to the voltage signal can include first determining a regional image containing a minimum depth value corresponding to the voltage signal. Thereafter, the vibration unit corresponding to the regional image can be used as the vibration unit corresponding to the voltage signal.
[0044] In an embodiment of the present invention, the vibration unit vibrates at different intensities based on the minimum depth value of each area image. The closer the object is, the greater the vibration intensity. This navigation method allows the user to perceive the distance and direction of surrounding objects by using different positions of the waist. It is understandable that stronger vibrations can indicate potential danger in that direction. The user can adjust the direction of travel in a timely manner based on the vibration feedback, thereby achieving a cross-modal conversion from visual information to tactile feedback, helping the user to perceive changes in the surrounding environment through touch without voice prompts and respond, thereby improving the accuracy of navigation and obstacle avoidance.
[0045] It is understood that in embodiments of the present invention, the smaller the depth value, the larger the voltage signal converted based on the depth value. A larger voltage signal indicates a stronger vibration intensity of the vibration unit, which can indicate a closer distance to the object. Conversely, a larger depth value may result in a smaller voltage signal. A smaller voltage signal indicates a weaker vibration intensity of the vibration unit, which can indicate a farther distance to the object.
[0046] In an embodiment of the present invention, before obtaining the first environment image captured by the first camera of the smart glasses and obtaining the second environment image captured by the second camera of the smart glasses, the method further includes: performing camera calibration between the first camera and / or the second camera to obtain camera internal parameters.
[0047] In an embodiment of the present invention, the camera internal parameters may include at least focal length, optical center position, and distortion coefficient. It should be noted that the camera internal parameters of the first camera and the second camera are the same. The camera internal parameters may be the camera internal parameters of the first camera or the camera internal parameters of the second camera. In an embodiment of the present invention, a preset calibration algorithm may be used to calibrate the first camera. The preset calibration algorithm may be set according to actual needs and is not specifically limited here. For example, the Zhang Zhengyou calibration algorithm, the Tsai-Lenz (Tsai) calibration algorithm, etc.
[0048] The technical solution of an embodiment of the present invention obtains a first environmental image captured by a first camera based on smart glasses and a second environmental image captured by a second camera based on the smart glasses. This technical solution, equipped with dual cameras, can capture information about the surrounding environment in real time, obtaining a first environmental image and a second environmental image. A target depth image is then obtained based on the first and second environmental images to provide more accurate spatial information. The target depth image is divided into at least two regional images, and a minimum depth value is determined for each of the regional images. Based on a preset correspondence between the depth value and the voltage signal strength, each minimum depth value is converted into a voltage signal of a corresponding strength, wherein the regional images correspond to the vibration units. Each voltage signal is transmitted to a smart belt, which drives the vibration unit corresponding to the voltage signal to vibrate at the intensity of the voltage signal, thereby navigating the person wearing the smart wearable device. The technical solution of an embodiment of the present invention achieves a precise conversion from visual to tactile sensation through a multi-point vibration design, providing users with a cross-modal navigation experience.
[0049] Figure 2A flowchart of a blind navigation method based on visual-tactile cross-modal perception provided by an embodiment of the present invention. Based on the aforementioned embodiment, optionally, obtaining a target depth image based on the first environment image and the second environment image includes: obtaining an initial depth image based on the first environment image and the second environment image; when the depth information of the initial depth image is incomplete, performing depth information completion processing on the initial depth image and using the completed image as the target depth image; when the depth information of the initial depth image is complete, using the initial depth image as the target depth image. Technical features identical or similar to those in the aforementioned embodiment are not repeated here.
[0050] like Figure 2 As shown, the method of this embodiment specifically includes:
[0051] S210: Acquire a first environment image captured by a first camera of the smart glasses, and acquire a second environment image captured by a second camera of the smart glasses.
[0052] S220: Obtain an initial depth image based on the first environment image and the second environment image.
[0053] The initial depth image may be understood as a depth image obtained based on the first environment image and the second environment image.
[0054] In an embodiment of the present invention, there are two ways to obtain an initial depth image based on the first environment image and the second environment image. As an optional implementation of an embodiment of the present invention, a deep learning algorithm (e.g., MonoDepth algorithm, StereoNet algorithm) can be used to obtain an initial depth image based on the first environment image and the second environment image. In an embodiment of the present invention, the use of a deep learning algorithm can preset a more refined depth image from the binocular image through a neural network.
[0055] As another optional implementation of an embodiment of the present invention, a Semi-Global BlockMatching (SGBM) algorithm can be used to generate a depth image based on the first environmental image and the second environmental image. Thereby, the generated depth image can be used as an initial depth image. Specifically, the SGBM algorithm can be used to calculate the view difference between the first environmental image and the second environmental image. The view difference can then be converted into a depth map, that is, an initial depth image is obtained. In an embodiment of the present invention, the benefit of using the SGBM algorithm to generate a depth image is that it can achieve a good balance between real-time performance and accuracy, so that a more accurate depth image can be obtained relatively quickly.
[0056] S230 : When the depth information of the initial depth image is incomplete, perform depth information completion processing on the initial depth image, and use the completed image as the target depth image.
[0057] In an embodiment of the present invention, when the depth information of the initial depth image is incomplete, the depth information of the initial depth image can be supplemented based on a preset depth completion algorithm, and the supplemented image can be used as the target depth image. The preset depth completion algorithm is set according to actual needs and is not specifically limited here. For example, the AGG-Net (Attention Guided Gated-convolutional Network) based on deep learning can be used. In an embodiment of the present invention, the use of the preset depth completion algorithm can generate a depth image with more complete and accurate depth information when the depth information of the initial depth image is incomplete, thereby providing a basis for the generation of subsequent signals, thereby improving the accuracy of navigation.
[0058] On the basis of the above embodiment, when the depth information of the initial depth image is sparse, depth information completion processing may be performed on the initial depth image, and the completed image is used as the target depth image.
[0059] S240 : When the depth information of the initial depth image is complete, use the initial depth image as a target depth image.
[0060] Specifically, when the depth information of the initial depth image is complete, the initial depth image may be used as the target depth image.
[0061] S250. Divide the target depth image into at least two area images, determine the minimum depth value of each area image, and convert each minimum depth value into a voltage signal of corresponding intensity based on a preset correspondence between the depth value size and the voltage signal strength, wherein the area image corresponds to the vibration unit.
[0062] S260: Transmit each of the voltage signals to the smart belt, and drive a vibration unit corresponding to the voltage signal to vibrate at the intensity of the voltage signal to perform navigation for the person wearing the smart wearable device.
[0063] The technical solution of the embodiment of the present invention is to obtain an initial depth image based on the first environmental image and the second environmental image; when the depth information of the initial depth image is incomplete, the depth information of the initial depth image is completed, and the completed image is used as the target depth image; when the depth information of the initial depth image is complete, the initial depth image is used as the target depth image, so as to obtain a depth image with more complete and accurate depth information, thereby improving the accuracy of navigation.
[0064] The embodiment of the present invention provides an optional embodiment of a blind navigation method based on visual-tactile cross-modal perception. The smart wearable device of the embodiment of the present invention may include smart glasses ( Figure 3 1) and smart belt ( Figure 3 2), wherein the smart glasses may include two RGB cameras ( Figure 3 3), the smart belt includes a plurality of vibration plates ( Figure 3 4). Figure 4 The following steps are specifically included for blind navigation based on the smart wearable device:
[0065] First, the two cameras of the smart glasses capture an image of the environment, namely a GRB color map. The captured GRB color map is then processed to generate a target depth image. This target depth image is then segmented to determine the closest point, resulting in a closest point depth map. The closest point depth map includes the minimum depth value for each segmented region. This minimum depth value for each region is then converted into a voltage signal, which drives the corresponding vibrator on the smart belt to vibrate at varying intensities, allowing the user to perceive the distance and direction of surrounding objects based on the position of their waist.
[0066] In the embodiment of the present invention, Figure 5 As shown, the collected GRB color image is processed to obtain a target depth image, and the target depth image is segmented to obtain the closest point. The specific steps may include: first, dual-target positioning and camera internal parameter conversion can be performed. After that, the left and right images, i.e., the first environment image and the second environment image, can be collected in real time by the RGB camera of the smart glasses. After collecting the left and right images, image distortion correction can be performed. After that, the SGBM algorithm can be used to process the distortion-corrected left and right images to generate an initial depth image. After obtaining the initial depth image, the AGG-Net algorithm can be used to complete the depth information of the initial depth image to obtain the target depth image. The target depth image can then be segmented to obtain regional images corresponding to each region. In this way, the distance value of the nearest point can be calculated for each region, i.e., the minimum depth value of each regional image.
[0067] The technical solution of an embodiment of the present invention obtains a first environmental image captured by a first camera based on smart glasses and a second environmental image captured by a second camera based on the smart glasses. This technical solution, equipped with dual cameras, can capture information about the surrounding environment in real time, obtaining a first environmental image and a second environmental image. A target depth image is then obtained based on the first and second environmental images to provide more accurate spatial information. The target depth image is divided into at least two regional images, and a minimum depth value is determined for each of the regional images. Based on a preset correspondence between the depth value and the voltage signal strength, each minimum depth value is converted into a voltage signal of a corresponding strength, wherein the regional images correspond to the vibration units. Each voltage signal is transmitted to a smart belt, which drives the vibration unit corresponding to the voltage signal to vibrate at the intensity of the voltage signal, thereby navigating the person wearing the smart wearable device. The technical solution of an embodiment of the present invention achieves a precise conversion from visual to tactile sensation through a multi-point vibration design, providing users with a cross-modal navigation experience.
[0068] Figure 6 This is a structural diagram of a blind navigation device based on visual-tactile cross-modal perception provided by an embodiment of the present invention. Figure 6 As shown, the device includes: an image obtaining module 310 , a signal conversion module 320 and a navigation module 330 .
[0069] Among them, the image acquisition module 310 is used to obtain a first environmental image taken by a first camera based on the smart glasses, obtain a second environmental image taken by a second camera based on the smart glasses, and obtain a target depth image based on the first environmental image and the second environmental image; the signal conversion module 320 is used to divide the target depth image into at least two regional images, determine the minimum depth value of each of the regional images, and convert each of the minimum depth values into a voltage signal of corresponding intensity based on a preset correspondence between the depth value size and the voltage signal intensity, wherein the regional image corresponds to the vibration unit; the navigation module 330 is used to transmit each of the voltage signals to the smart belt, drive the vibration unit corresponding to the voltage signal to vibrate with the intensity of the voltage signal, so as to navigate the object wearing the smart wearable device.
[0070] The technical solution of an embodiment of the present invention uses an image acquisition module to obtain a first environmental image captured by a first camera of smart glasses and a second environmental image captured by a second camera of the smart glasses. This technical solution uses dual-camera smart glasses to capture information about the surrounding environment in real time, obtaining first and second environmental images. Subsequently, a signal conversion module uses the first and second environmental images to obtain a target depth image, providing more accurate spatial information. The target depth image is divided into at least two regional images, and a minimum depth value is determined for each of the regional images. Based on a preset correspondence between depth value magnitude and voltage signal strength, each minimum depth value is converted into a voltage signal of corresponding strength, wherein the regional images correspond to the vibration units. A navigation module transmits each voltage signal to a smart belt, driving the vibration units corresponding to the voltage signals to vibrate at the same strength as the voltage signal, thereby navigating the wearer of the smart wearable device. The technical solution of an embodiment of the present invention achieves a precise conversion from visual to tactile perception through a multi-point vibration design, providing users with a cross-modal navigation experience.
[0071] Optionally, the signal conversion module includes a depth image generation unit, wherein the depth image generation unit is used to obtain an initial depth image based on the first environmental image and the second environmental image; when the depth information of the initial depth image is incomplete, the initial depth image is completed and the completed image is used as the target depth image; when the depth information of the initial depth image is complete, the initial depth image is used as the target depth image.
[0072] Optionally, the depth image generating unit is configured to obtain an initial depth image based on the first environment image and the second environment image by adopting an SGBM algorithm.
[0073] Optionally, the device also includes a camera calibration module, which is used to perform camera calibration between the first camera and / or the second camera before obtaining the first environment image taken by the first camera of the smart glasses and obtaining the second environment image taken by the second camera of the smart glasses to obtain camera internal parameters.
[0074] Optionally, the first camera and the second camera are RGB cameras respectively.
[0075] Optionally, the device also includes an image distortion correction module, wherein the image distortion correction module is used to perform image distortion correction processing on the first environmental image in the event that image distortion occurs in the first environmental image before obtaining the target depth image based on the first environmental image and the second environmental image; and / or, perform image distortion correction processing on the second environmental image in the event that image distortion occurs in the second environmental image.
[0076] Optionally, the preset corresponding relationship between the depth value and the voltage signal strength may be a linear relationship or a nonlinear relationship.
[0077] The blind navigation device based on visual-tactile cross-modal perception provided by an embodiment of the present invention can execute the blind navigation method based on visual-tactile cross-modal perception provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0078] It is worth noting that the various units and modules included in the above-mentioned blind navigation device based on visual-tactile cross-modal perception are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the embodiments of the present invention.
[0079] Figure 7 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0080] like Figure 7As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0081] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0082] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the navigation method for the blind based on visual-tactile cross-modal perception.
[0083] In some embodiments, the method for blind navigation based on visual-tactile cross-modal perception can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for blind navigation based on visual-tactile cross-modal perception described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the method for blind navigation based on visual-tactile cross-modal perception by any other appropriate means (for example, by means of firmware).
[0084] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0085] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0086] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0087] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0088] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0089] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0090] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0091] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A navigation method for the blind based on visual-tactile cross-modal perception, characterized in that: The method is executed by a processor of a smart wearable device, the smart wearable device including smart glasses and a smart belt, the smart belt being provided with at least two vibration units, and the method comprising: Acquire a first environment image captured by a first camera of the smart glasses, acquire a second environment image captured by a second camera of the smart glasses, and obtain a target depth image based on the first environment image and the second environment image; Dividing the target depth image into at least two regional images, determining a minimum depth value of each of the regional images, and converting each of the minimum depth values into a voltage signal of corresponding intensity based on a preset correspondence between depth value magnitude and voltage signal intensity, wherein the regional images correspond to the vibration units; Each of the voltage signals is transmitted to the smart belt, and a vibration unit corresponding to the voltage signal is driven to vibrate at the intensity of the voltage signal to navigate the subject wearing the smart wearable device.
2. The method according to claim 1, characterized in that The obtaining of a target depth image based on the first environment image and the second environment image includes: Obtaining an initial depth image based on the first environment image and the second environment image; In the case that the depth information of the initial depth image is incomplete, performing depth information completion processing on the initial depth image, and using the completed image as the target depth image; When the depth information of the initial depth image is complete, the initial depth image is used as the target depth image.
3. The method according to claim 2, characterized in that The obtaining of an initial depth image based on the first environment image and the second environment image includes: An initial depth image is obtained based on the first environment image and the second environment image using the SGBM algorithm.
4. The method according to claim 1, wherein Before acquiring the first environment image captured by the first camera of the smart glasses and acquiring the second environment image captured by the second camera of the smart glasses, the method further includes: Perform camera calibration on the first camera and / or the second camera to obtain camera internal parameters.
5. The method according to claim 1, wherein The first camera and the second camera are RGB cameras respectively.
6. The method according to claim 1, characterized in that Before obtaining the target depth image based on the first environment image and the second environment image, the method further includes: In the case where the first environmental image has image distortion, the first environmental image is subjected to image distortion correction processing; and / or in the case where the second environmental image has image distortion, the second environmental image is subjected to image distortion correction processing.
7. The method according to claim 1, characterized in that The preset corresponding relationship between the depth value and the voltage signal strength may be a linear relationship or a nonlinear relationship.
8. A navigation device for the blind based on visual-tactile cross-modal perception, characterized in that: The device comprises: An image acquisition module is configured to acquire a first environment image captured by a first camera of the smart glasses, acquire a second environment image captured by a second camera of the smart glasses, and acquire a target depth image based on the first environment image and the second environment image; a signal conversion module, configured to divide the target depth image into at least two regional images, determine a minimum depth value of each of the regional images, and convert each of the minimum depth values into a voltage signal of corresponding strength based on a preset correspondence between depth value magnitude and voltage signal strength, wherein the regional images correspond to the vibration units; The navigation module is used to transmit each of the voltage signals to the smart belt, drive the vibration unit corresponding to the voltage signal to vibrate at the intensity of the voltage signal, and provide navigation for the subject wearing the smart wearable device.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the blind navigation method based on visual-tactile cross-modal perception according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the blind navigation method based on visual-tactile cross-modal perception according to any one of claims 1 to 7 when executed.
Citation Information
Patent Citations
Binocular stereo vision-based guide eyeglasses for blind person
CN108169927A
Blind person navigation system and method integrating millimeter wave radar and depth vision
CN117224370A