Visual impairment assisting method and system based on multi-modal sensing and low-load feedback and medium
Through multimodal sensing and low-load feedback technology, the problems of limited information bandwidth and insufficient computing resources in traditional visual impairment assistive technologies are solved, and more efficient and robust visual impairment assistive effects are achieved.
Patent Information
- Application Number
- CN202510300879.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional visual impairment assistive technology relies on single sensory feedback, has limited information bandwidth, high cognitive load, and limited computing power and resources, so it is unable to effectively assist people with visual impairment.
Multimodal sensing (including color maps, depth maps and acceleration data) is used for data analysis, combining time and space alignment, pixel-level obstacle detection is performed, computing resource requirements are reduced, and visually impaired people are guided through natural, low-load spatial audio feedback.
Improves the effectiveness of visual impairment assistance, reduces cognitive load and computing resource requirements, and provides more robust and accurate non-visual feedback to help visual impairment navigate the environment more easily.
Smart Images

Figure CN120227235A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vision impairment assistance technologies, and particularly to a vision impairment assistance method, system, and medium based on multi-modal sensing and low-load feedback. Background Art
[0002] Vision impairment assistance senses the environment through various sensors and conveys information to vision-impaired individuals in a non-visual manner to help them cope with various challenges in daily life. Usually, vision impairment assistance also needs to meet the personalized needs of vision-impaired individuals, such as visual recognition ability, obstacle avoidance ability, usage training, etc.
[0003] Currently, traditional vision impairment assistance relies on single-sensory feedback (such as a tactile belt or voice navigation), which has a limited information bandwidth, resulting in a high cognitive load for users as they need to undergo complex training to adapt to the feedback mode. At the same time, since traditional vision impairment assistance usually uses single-modal sensing and the general vision models it relies on are not customized for the behavior patterns of visually impaired individuals, its computing power and computing resources are both limited, and it cannot effectively assist visually impaired individuals. Summary of the Invention
[0004] The objective of this application is to provide a vision impairment assistance method, system, and medium based on multi-modal sensing and low-load feedback, which improves the effectiveness of vision impairment assistance by using multi-modal sensing, and at the same time adopts a lower-load non-visual feedback for vision-impaired individuals, and the entire vision impairment assistance process has lower requirements for computing resources.
[0005] To achieve the above objective, this application provides the following solutions:
[0006] In the first aspect, this application provides a vision impairment assistance method based on multi-modal sensing and low-load feedback, and the vision impairment assistance method based on multi-modal sensing and low-load feedback includes:
[0007] Obtain the environmental image in front of the vision-impaired individual and the original acceleration data of the individual itself; the environmental image includes an original color image and an original depth image;
[0008] Align the environmental image and the original acceleration data in time to obtain a first color image, a first depth image, and first acceleration data;
[0009] Align the first color image, the first depth image, and the first acceleration data in space to obtain a second color image, a second depth image, and second acceleration data;
[0010] Based on the first acceleration data, determine whether the first color image is blurred; if so, return to the step of "acquiring the environmental image in front of the visually impaired person and their own original acceleration data"; if not, determine whether the second color image belongs to a low-light image;
[0011] When the second color image belongs to a low-light image, determine the detection box data of the target object based on the second depth map; the target object is the object of interest to the visually impaired person;
[0012] When the second color image does not belong to a low-light image, determine the detection box data of the target object based on the second color image;
[0013] Based on the detection box data and the second depth map, determine the angle and distance of the target object relative to the visually impaired person;
[0014] Determine the unobstructed block in the second depth map and obtain the unobstructed block information;
[0015] Based on the unobstructed block information and the angle and distance of the target object relative to the visually impaired person, select the center direction of the nearest unobstructed block to the target object as the target direction and obtain the target direction information;
[0016] Perform sound encoding on the target direction information to obtain spatial audio, broadcast the spatial audio, and guide the visually impaired person to approach the target object.
[0017] In a second aspect, the present application further provides a visually impaired assistance system based on multimodal sensing and low-load feedback. The visually impaired assistance system based on multimodal sensing and low-load feedback includes:
[0018] A data acquisition module for acquiring the environmental image in front of the visually impaired person and their own original acceleration data; the environmental image includes an original color image and an original depth image;
[0019] A time alignment module for performing time alignment on the environmental image and the original acceleration data to obtain a first color image, a first depth image, and first acceleration data;
[0020] A spatial alignment module for performing spatial alignment on the first color image, the first depth image, and the first acceleration data to obtain a second color image, a second depth image, and second acceleration data;
[0021] A motion blur determination module for determining whether the first color image is blurred based on the first acceleration data; if so, return to the step of "acquiring the environmental image in front of the visually impaired person and their own original acceleration data"; if not, determine whether the second color image belongs to a low-light image;
[0022] A color map object detection module, configured to determine detection box data of the target object based on the second color map when the second color map does not belong to a low-light map;
[0023] A depth map object detection module, configured to determine detection box data of the target object based on the second depth map when the second color map belongs to a low-light map; the target object is an object of interest to the visually impaired person;
[0024] An angle and distance calculation module, configured to determine the angle and distance of the target object relative to the visually impaired person based on the detection box data and the second depth map;
[0025] An obstacle detection module, configured to determine a barrier-free block in the second depth map and obtain barrier-free block information;
[0026] A path selection module, configured to select the center direction of the block closest to the target object and barrier-free as the target direction based on the barrier-free block information and the angle and distance of the target object relative to the visually impaired person, and obtain target direction information;
[0027] A sound encoding module, configured to perform sound encoding based on the target direction information to obtain spatial audio, broadcast the spatial audio, and guide the visually impaired person to approach the target object.
[0028] In a third aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the visually impaired assistance method based on multimodal sensing and low-load feedback described in the first aspect.
[0029] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0030] The present application performs multimodal data analysis and processing based on the original color map, the original depth map, and the original acceleration data. At the same time, the obstacle detection in the present application does not use a neural network architecture, but according to the depth distribution of the ground, makes full use of the advantage that the depth map can obtain three-dimensional distance information compared with the color map to perform pixel-level obstacle detection. This method greatly reduces the amount of calculation, has lower requirements for computing resources, and is more robust to the environment. In addition, the spatial audio obtained by performing sound encoding based on the target direction information in the present application is natural and low-load. Description of the Drawings
[0031] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0032] Figure 1 It is a flowchart of the visual impairment assistance method based on multi-modal sensing and low-load feedback provided by the embodiments of the present application;
[0033] Figure 2 It is a schematic diagram of the modules of the visual impairment assistance system based on multi-modal sensing and low-load feedback provided by the embodiments of the present application. Detailed implementation manners
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0035] The purpose of the present application is to provide a visual impairment assistance method, system and medium based on multi-modal sensing and low-load feedback. By adopting the multi-modal sensing method, the effectiveness of visual impairment assistance is improved, and at the same time, a non-visual feedback with a lower load is adopted for the visually impaired. The entire visual impairment assistance process has lower requirements for computing resources.
[0036] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0037] As Figure 1 shown, this embodiment provides a visual impairment assistance method based on multi-modal sensing and low-load feedback. The visual impairment assistance method based on multi-modal sensing and low-load feedback includes:
[0038] Step S1: Obtain the environmental image in front of the visually impaired person and the original acceleration data of the person himself; the environmental image includes the original color image and the original depth image.
[0039] In this embodiment, the original color image is an environmental color image within the set visual field in front of the visually impaired person. Its horizontal visual field range is [-35°, 35°], and its vertical visual field range is [-25°, 25°]. An RGB camera can be used to capture the original color image of the environment in front of the visually impaired person; the original depth image is an environmental depth image within the set visual field in front of the visually impaired person. Its horizontal visual field range is [-45°, 45°], and its vertical visual field range is [-30°, 30°]. A Depth camera can be used to capture the original depth image of the environment in front of the visually impaired person; an IMU is used to collect the original acceleration data of the visually impaired person himself / herself, and the acceleration data is the acceleration data in 6-axis directions.
[0040] Step S2: Align the environmental images and the original acceleration data in terms of time to obtain the first color image, the first depth image, and the first acceleration data.
[0041] In this embodiment, the environmental images and the original acceleration data are aligned in terms of time by checking timestamps.
[0042] Step S3: Align the first color image, the first depth image, and the first acceleration data in terms of space to obtain the second color image, the second depth image, and the second acceleration data.
[0043] In this embodiment, an algorithm provided by Intel's librealsense software package is used to align the first color image, the first depth image, and the first acceleration data in terms of space.
[0044] Step S4: Based on the first acceleration data, determine whether the first color image is blurred; if so, return to step S1; if not, determine whether the second color image belongs to a low-light image.
[0045] In this embodiment, step S4 specifically includes:
[0046] Step S41: Perform low-pass filtering on the acceleration data in 6-axis directions in the first acceleration data.
[0047] Step S42: Perform vector addition on the low-pass filtered acceleration data in 6-axis directions to obtain the three-dimensional space acceleration.
[0048] Step S43: When the absolute value of the three-dimensional space acceleration is greater than 0.5, the first color image is blurred.
[0049] Step S44: When the absolute value of the three-dimensional space acceleration is less than or equal to 0.5, the first color image is not blurred.
[0050] Step S5: When the second color image belongs to a low-light image, determine the detection box data of the target object based on the second depth image; the target object is the object of interest to the visually impaired person.
[0051] In this embodiment, the depth map data of the customized environment is collected, and this part of the data is used to adjust the model structure of the YOLOv8 object detection model (changing the input from the original three-channel color map to a one-channel depth map), and then the model weights are adjusted through retraining to obtain a depth map object detection module. Then, the second depth map is input into the depth map object detection module to determine the detection box data of the target object.
[0052] Step S6: When the second color map does not belong to a low-light map, determine the detection box data of the target object based on the second color map.
[0053] In this embodiment, the color map data of the customized environment is collected, and this part of the data is used to adjust the model parameter weights of the YOLOv8 object detection model to obtain a color map object detection module. Then, the second color map is input into the color map object detection module to determine the detection box data of the target object.
[0054] Step S7: Based on the detection box data and the second depth map, determine the angle and distance of the target object relative to the visually impaired person.
[0055] In this embodiment, step S7 specifically includes:
[0056] Step S71: Based on the detection box data and the second depth map, use linear mapping to process the center point coordinates of the target object and calculate the angle of the target object relative to the visually impaired person.
[0057] Among them, extract the abscissa of the center point coordinates of the target object in the detection box data, subtract the abscissa of the second depth map from the abscissa. If the result is greater than 0, the target object is on the right side of the visually impaired person; otherwise, the target object is on the left side of the visually impaired person. Take the absolute value of this result, divide it by the resolution of the RGB camera in the horizontal direction, and then multiply it by the horizontal field of view angle of the RGB camera, so as to obtain the absolute value of the angle of the target object relative to the visually impaired person.
[0058] Step S72: Based on the detection box data and the second depth map, use two-cluster clustering to process the center depth value of the target object in the detection box data and calculate the distance of the target object relative to the visually impaired person.
[0059] Among them, the abscissa and ordinate of the center point coordinates of the target object in the detection box data are extracted, and extended by 10 pixels in the vertical and horizontal directions to obtain a 10×10 area. The depth values of the pixels within the 10×10 area in the second depth map are clustered into two categories. Determine the difference between the depth values of the two cluster centers. If the difference < 0.1 meter, directly take the average of the two cluster centers as the depth value of the target object; otherwise, take the distance from the class center of the class with more pixels in the two classes as the depth value of the target object, and use this as the distance of the target object relative to the visually impaired person.
[0060] Step S8: Determine the unobstructed blocks in the second depth map and obtain the unobstructed block information.
[0061] In this embodiment, step S8 specifically includes:
[0062] Step S81: Detect obstacles in the second depth map by fitting the depth values with a binomial exponential function.
[0063] Calculate the average value of the depth values of all columns in the second depth map, and take the column with the largest average value as the candidate column. Use the row index of the candidate column as x and the row depth value of the candidate column as y to fit the following function:
[0064] f(x) = ae bx +ce dx ;
[0065] First, determine the values of a, b, c, and d in the above formula, and then judge the difference between the measured depth value and the fitted depth value of the part where the row index of each column > the maximum row index - 180. If the difference < 0.18, it is determined as the ground; otherwise, it is determined as an obstacle. Then, evenly divide the horizontal direction of the second depth map into 7 blocks. For each block, if the proportion of its ground pixels > 95%, it is determined as an unobstructed area; otherwise, it is determined as an area with obstacles. Thus, it can be determined whether each of the 7 blocks belongs to an unobstructed block.
[0066] Step S82: Divide unobstructed blocks based on the obstacles detected in the second depth map to obtain unobstructed block information.
[0067] Step S9: Based on the unobstructed block information, as well as the angle and distance of the target object relative to the visually impaired person, select the center direction of the nearest unobstructed block to the target object as the target direction to obtain the target direction information.
[0068] In this embodiment, the central direction of the block closest to the target object and without obstacles is selected as the target direction. Combining the minimum rotation cost, the optimal passable direction is determined to obtain the target direction information. The priority of obstacles is the highest, that is, if there are obstacles in the direction where the object is located, obstacle avoidance will be performed first (that is, the central direction of the block closest to the middle (the fourth block from left to right) and without obstacles among the 7 blocks is selected as the final walking direction), and then walk towards the target object; if there are no obstacles in the direction where the target object is located, the central direction of the block closest to the central direction of the block where the target object is located is directly selected as the final walking direction.
[0069] Step S10: Based on the target direction information, perform sound encoding to obtain spatial audio, broadcast the spatial audio, and guide the visually impaired person to approach the target object.
[0070] In this embodiment, step S10 specifically includes:
[0071] Step S101: Determine two sounds with different durations and spectrograms, which are the first sound and the second sound respectively. Among them, both sounds are water droplet sounds, the duration of the first sound is 50 milliseconds, and the duration of the second sound is 100 milliseconds.
[0072] Step S102: When the angle corresponding to the target direction information is less than -5°, use the head-related transfer function to encode the first sound into spatial audio located at 90° on the left.
[0073] Step S103: When the angle corresponding to the target direction information is greater than 5°, use the head-related transfer function to encode the first sound into spatial audio located at 90° on the right.
[0074] Step S104: When the angle corresponding to the target direction information is greater than or equal to -5° and less than or equal to 5°, use the head-related transfer function to encode the second sound into spatial audio located directly in front.
[0075] As Figure 2 shown, this embodiment also provides a visually impaired assistance system based on multi-modal sensing and low-load feedback. The visually impaired assistance system based on multi-modal sensing and low-load feedback includes:
[0076] A data acquisition module 100, configured to acquire the environmental image in front of the visually impaired person and the original acceleration data of itself; the environmental image includes an original color map and an original depth map.
[0077] A time alignment module 200, configured to perform time alignment on the environmental image and the original acceleration data to obtain a first color map, a first depth map, and first acceleration data.
[0078] The spatial alignment module 300 is used to spatially align the first color image, the first depth image, and the first acceleration data to obtain the second color image, the second depth image, and the second acceleration data.
[0079] The motion blur determination module 400 is used to determine whether the first color image is blurred based on the first acceleration data; if so, return to the step of "acquiring the environmental image in front of the visually impaired person and their own original acceleration data"; if not, determine whether the second color image belongs to a low-light image.
[0080] The color image target detection module 500 is used to determine the detection box data of the target object based on the second color image when the second color image does not belong to a low-light image.
[0081] The depth image target detection module 600 is used to determine the detection box data of the target object based on the second depth image when the second color image belongs to a low-light image; the target object is the object of interest to the visually impaired person.
[0082] The angle and distance calculation module 700 is used to determine the angle and distance of the target object relative to the visually impaired person based on the detection box data and the second depth image.
[0083] The obstacle detection module 800 is used to determine the obstacle-free blocks in the second depth image and obtain the obstacle-free block information.
[0084] The path selection module 900 is used to select the center direction of the block closest to the target object and obstacle-free based on the obstacle-free block information and the angle and distance of the target object relative to the visually impaired person as the target direction, and obtain the target direction information.
[0085] The sound encoding module 1000 is used to perform sound encoding based on the target direction information to obtain spatial audio, broadcast the spatial audio, and guide the visually impaired person to approach the target object.
[0086] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0087] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0088] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0089] All actions of obtaining signals, information, or data in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and with the authorization given by the owner of the corresponding device.
[0090] In summary, the present application mainly has the following technical features:
[0091] (1) Visual recognition requires lower computing resources and is more robust to the environment compared to the prior art.
[0092] The present application performs pixel-level obstacle detection. Instead of using a neural network architecture, it uses a binomial exponential function to fit the depth values according to the depth distribution of the ground to make judgments. This fully utilizes the advantage that the Depth camera can obtain three-dimensional distance information compared to the RGB camera and combines the characteristic that the depth value distribution will show abnormal protrusions due to raised obstacles on the ground plane. Therefore, the amount of calculation is greatly reduced, enabling the algorithm to run in real time on wearable devices. The present application considers the more essential feature of obstacles, that is, obstacles will cause changes in the terrain, and does not require the determination of the type of obstacles, so it is more robust.
[0093] (2) Non-visual feedback with lower load.
[0094] Existing sounds include voice feedback and 3D sounds, etc. Compared with such methods, the present application is more accurate and has a lower mental load during use. First, the natural binaural effect of humans enables them to perceive the source of sound, and this source perception is of low load. However, the accuracy of the human ear in judging the direction of the sound source is not high enough. The present application uses the head-related transfer function to simulate the binaural effect, making our direction feedback natural and of low load. In terms of accuracy, the present application uses different sounds and spatial effects to help users quickly locate the direction of the navigation guidance, which is easier to accurately locate the specific direction compared with 3D sounds. Compared with voice, the sound of the present application is short and has no specific language information, which is faster and does not affect the user's perception and understanding of the ambient sound when receiving the sound feedback.
[0095] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0096] Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, based on the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for assisting visual impairment based on multimodal sensing and low-load feedback, characterized in that: The visual impairment assistance method based on multimodal sensing and low-load feedback includes: Acquire an environment image in front of the visually impaired person and the original acceleration data of the person himself; the environment image includes an original color image and an original depth image; Time-aligning the environment image and the original acceleration data to obtain a first color image, a first depth image, and first acceleration data; Spatially aligning the first color image, the first depth image, and the first acceleration data to obtain a second color image, a second depth image, and second acceleration data; Determine whether the first color image is blurred based on the first acceleration data; if so, return to the step of "obtaining the environment image in front of the visually impaired person and the original acceleration data of the device itself"; if not, determine whether the second color image is a dark light image; When the second color image is a dark-light image, determining detection frame data of a target object based on the second depth image; the target object is an object of interest to the visually impaired person; When the second color image is not a dark-light image, determining the detection frame data of the target object based on the second color image; Determining an angle and a distance of the target object relative to the visually impaired person based on the detection frame data and the second depth map; Determine an unobstructed block in the second depth map, and obtain unobstructed block information; Based on the unobstructed block information and the angle and distance of the target object relative to the visually impaired person, select the center direction of the block closest to the target object and unobstructed as the target direction, and obtain the target direction information; Perform sound encoding based on the target direction information to obtain spatial audio, broadcast the spatial audio and guide the visually impaired person to approach the target object.
2. The visual impairment assistance method based on multimodal sensing and low-load feedback according to claim 1 is characterized in that: The original color image is an environmental color image within a set visual field in front of the visually impaired person; wherein the horizontal visual field range is [-35°, 35°], and the vertical visual field range is [-25°, 25°].
3. The visual impairment assistance method based on multimodal sensing and low-load feedback according to claim 1 is characterized in that: The original depth map is an environmental depth map within a set field of view in front of the visually impaired person; wherein the horizontal field of view is [-45°, 45°], and the vertical field of view is [-30°, 30°].
4. The visual impairment assistance method based on multimodal sensing and low-load feedback according to claim 1, characterized in that: Determining whether the first color image is blurred based on the first acceleration data specifically includes: Performing low-pass filtering on the acceleration data in the six axis directions in the first acceleration data; The acceleration data in the six axis directions after low-pass filtering are vector-added to obtain the three-dimensional spatial acceleration; When the absolute value of the three-dimensional space acceleration is greater than 0.5, the first color image is blurred; When the absolute value of the three-dimensional space acceleration is less than or equal to 0.5, the first color image is not blurred.
5. The visual impairment assistance method based on multimodal sensing and low-load feedback according to claim 1, characterized in that: Determining whether the second color image is a dark-light image specifically includes: converting the second color image into a grayscale image; Calculate the average brightness value of the pixels in the grayscale image; When the average brightness value is less than 50, the second color image is a dark-light image; When the average brightness value is greater than or equal to 50, the second color image is not a dark-light image.
6. The visual impairment assistance method based on multimodal sensing and low-load feedback according to claim 1, characterized in that: Determining the angle and distance of the target object relative to the visually impaired person based on the detection frame data and the second depth map specifically includes: Based on the detection frame data and the second depth map, use linear mapping to process the center point coordinates of the target object to calculate the angle of the target object relative to the visually impaired person; Based on the detection frame data and the second depth map, two-cluster clustering is used to process the center depth value of the target object in the detection frame data to calculate the distance of the target object relative to the visually impaired person.
7. The visual impairment assistance method based on multimodal sensing and low-load feedback according to claim 1, characterized in that: Determining an unobstructed block in the second depth map and obtaining unobstructed block information specifically includes: Detecting obstacles in the second depth map by fitting the depth values with a binomial exponential function; Obstacle-free blocks are divided based on the obstacles detected in the second depth map to obtain obstacle-free block information.
8. The visual impairment assistance method based on multimodal sensing and low-load feedback according to claim 1, characterized in that: Performing sound encoding based on the target direction information to obtain spatial audio specifically includes: Determine two sounds with different durations and spectrograms, which are the first sound and the second sound; When the angle corresponding to the target direction information is less than -5°, the first sound is encoded into a spatial audio located 90° to the left using a head-related transformation function; When the angle corresponding to the target direction information is greater than 5°, using a head-related transfer function to encode the first sound into a spatial audio located 90° to the right; When the angle corresponding to the target direction information is greater than or equal to -5° and less than or equal to 5°, the second sound is encoded into spatial audio located directly in front using a head-related transfer function.
9. A visually impaired assistance system based on multimodal sensing and low-load feedback, characterized in that: The visual impairment assistance system based on multimodal sensing and low-load feedback includes: A data acquisition module, used to obtain an environmental image in front of the visually impaired person and its own original acceleration data; the environmental image includes an original color image and an original depth image; A time alignment module, used for performing time alignment on the environment image and the original acceleration data to obtain a first color image, a first depth image and first acceleration data; A spatial alignment module, used for spatially aligning the first color image, the first depth image and the first acceleration data to obtain a second color image, a second depth image and second acceleration data; a motion blur determination module, configured to determine whether the first color image is blurred based on the first acceleration data; if so, return to the step of "obtaining the environment image in front of the visually impaired person and the original acceleration data of the device itself"; if not, determine whether the second color image is a dark light image; a color image target detection module, configured to determine detection frame data of the target object based on the second color image when the second color image is not a dark-light image; a depth map target detection module, configured to determine detection frame data of a target object based on the second depth map when the second color map is a dark-light map; the target object is an object of interest to the visually impaired person; an angle and distance calculation module, configured to determine an angle and a distance of the target object relative to the visually impaired person based on the detection frame data and the second depth map; An obstacle detection module is used to determine an unobstructed block in the second depth map and obtain unobstructed block information; A path selection module, configured to select the center direction of the block closest to the target object and without obstacles as the target direction based on the barrier-free block information and the angle and distance of the target object relative to the visually impaired person, and obtain target direction information; The sound encoding module is used to perform sound encoding based on the target direction information to obtain spatial audio, broadcast the spatial audio and guide the visually impaired person to approach the target object.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for assisting visual impairment based on multimodal sensing and low-load feedback described in any one of claims 1 to 8 is implemented.