Blind person intelligent glasses based on binocular vision

Through blind glasses combined with binocular vision and GPS, three-dimensional depth perception and precise navigation are achieved, which solves the problem of insufficient perception range and recognition accuracy in the existing technology, and enhances the independent action and sense of security of blind people.

CN120045075APending Publication Date: 2025-05-27LINKER
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510379572.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing blind glasses cannot comprehensively utilize visual and positioning information, resulting in insufficient perceptual range and recognition accuracy, and lack of global navigation functions, limiting the blind's free activities.

Method used

The binocular vision system is used for three-dimensional depth perception, combined with GPS positioning information, to achieve accurate obstacle detection and avoidance, and to provide real-time navigation prompts. At the same time, the cloud-based visual language model is used to assist blind people in identifying texts, scenes, items and obstacles and conduct obstacle avoidance navigation.

Benefits of technology

It enhances the blind's environmental perception ability, provides accurate obstacle avoidance navigation support, and improves their independent movement ability and safety in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045075A_ABST
    Figure CN120045075A_ABST
Patent Text Reader

Abstract

The invention discloses a pair of blind person intelligent glasses based on binocular vision. The glasses comprise a glasses terminal and a cloud end. A data acquisition and transmission module, a voice interaction and feedback module and a power supply system are arranged in the glasses terminal, and the glasses terminal is used for acquiring and processing data and receiving and feeding back voice instructions; the cloud integrates a sensing module, a task planning module, an action module and a path planning and navigation module, the sensing module analyzes instructions and environment information, the task planning module extracts natural language intentions and generates task structures, the action module dynamically generates task templates to execute tasks, and the path planning and navigation module achieves fine-grained obstacle avoidance navigation. The environmental perception precision is improved through binocular vision and GPS fusion, the use threshold is reduced through natural language interaction, the complex task processing capacity of the system is improved through dynamic task decomposition, and the obstacle avoidance accuracy is improved through hierarchical path planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a pair of smart glasses for the blind, and more particularly to a pair of smart glasses for the blind based on binocular vision. Background Art

[0002] At present, the auxiliary equipment for blind people to travel generally adopts a single sensor or navigation technology, which cannot comprehensively utilize visual and positioning information to provide efficient obstacle avoidance solutions. Although the existing blind glasses can detect obstacles, there are problems with insufficient perception range and recognition accuracy, and there is no global navigation combined with GPS, which restricts the freedom of movement of blind people. Summary of the invention

[0003] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a smart obstacle avoidance glasses for the blind that combines binocular vision with GPS. The glasses can perform stereoscopic depth perception through the binocular vision system, and combine GPS positioning information to accurately detect and avoid obstacles, while providing real-time navigation prompts to help the blind walk safely. And based on the cloud-based visual language large model capabilities, it assists the blind to realize the ability to recognize text, scenes, objects, obstacles, and obstacle avoidance navigation.

[0004] To achieve the above object, the present invention provides the following technical solution: a pair of smart glasses for the blind based on binocular vision, comprising:

[0005] The eyeglass terminal has a built-in data collection and transmission module, a voice interaction and feedback module, and a battery and power management system. The data collection and transmission module collects data, processes the data and outputs it, and the voice interaction and feedback module receives the voice commands of the blind and simultaneously feeds back prompt commands to the blind.

[0006] The cloud has built-in perception module, task planning module, action module and path planning and navigation module. The perception module has the ability to perceive, understand and assist in decision-making for blind users and the surrounding environment. The task planning module has the ability to extract intent from natural language instructions and dynamically generate task decomposition structure, so as to convert complex natural language input into a series of executable subtasks and clarify their dependencies. The action module is responsible for receiving the task decomposition structure generated by the task planning module and dynamically constructing task templates to complete the actual execution. The path planning and navigation module implements the three sub-modules of visual processing and depth estimation, map generation and path planning and dynamic obstacle avoidance to achieve fine-grained obstacle avoidance navigation.

[0007] As a further improvement of the present invention, the data acquisition and transmission module includes:

[0008] A binocular camera is used to collect environmental images and video data to provide stereoscopic vision, helping the blind to identify the depth and distance information of obstacles;

[0009] GPS module, used to provide real-time geographic location data of the blind;

[0010] Wi-Fi module, used to connect to the cloud via wireless network;

[0011] External RTC module, used to provide accurate time tracking function, ensuring that the system can maintain consistent time data during long-term operation;

[0012] Local storage device, temporarily storing collected data through eMMC or MicroSD card;

[0013] A lightweight processor used to extract video key frames, process image data, calculate timestamps, mark camera IDs, and perform local tasks.

[0014] As a further improvement of the present invention, the voice interaction and feedback module includes:

[0015] A microphone array is used to collect the blind person’s voice commands and send them to the cloud for processing;

[0016] A speaker or bone conduction headset is used to output information about the location of obstacles and steering prompts according to the instructions provided by the voice feedback module;

[0017] The audio processing chip is connected to the microphone array and the speaker or the bone conduction headset to perform noise reduction and echo suppression on the audio.

[0018] As a further improvement of the present invention, the perception module unifies multimodal data such as images and audio into a processable internal representation based on the Decode-Only Transformer architecture, specifically in the following manner:

[0019] Step 1: Receive the multimodal raw data such as images and audio uploaded by the glasses, extract the GPS location information from the metadata of the image, and update the current GPS location information in the system in real time;

[0020] Step 2: Preprocess the image data, including resizing, color standardization, and background removal to ensure uniform image format and extract effective scene information; and preprocess the audio data, including removing environmental noise, sampling, and framing to optimize audio quality and provide a clearer data basis for audio analysis.

[0021] Step 3: extract features from the data preprocessed in step 2 and map the multimodal data into a unified vector representation;

[0022] Step 4: Standardize the feature vectors of multiple modalities such as images and audio into a unified representation space by sharing the encoding dimension or performing linear mapping;

[0023] Step 5: Based on feature alignment, the Decode-Only Transformer architecture is used to further process these multimodal feature vectors, and the mutual influence between different modalities is captured through the self-attention mechanism to generate a unified internal representation. The personalized background information of the portrait configuration is then fused with the multimodal data.

[0024] Step six, output the unified multimodal semantic representation vector fused in step five.

[0025] As a further improvement of the present invention, the specific method of feature extraction in step 3 is as follows:

[0026] Image feature extraction is performed. Based on ViT (Vision Transformer), the image is cut into fixed-size tiles. Each tile is then mapped to a high-dimensional vector through a linear transformation, and special tags and position encodings are added to ensure that the model perceives the spatial information of the tile. All tile embedding vectors are then arranged in order to generate an input sequence. Finally, the image embedding feature sequence is provided as input to Decode-Only Transformer, which is processed together with text or other modal data to support joint decoding of multimodal tasks.

[0027] Audio feature extraction,First, the audio signal is sampled, framed and feature extracted to generate a multi-dimensional feature matrix such as MFCC or Log-Mel Spectrogram. Then, a pre-trained audio model is used to extract a high-dimensional embedding representation and align it with the text features through a dimensionality reduction mechanism. Then, the ASR module is used to transcribe the audio signal into text while retaining the original audio features such as pitch and speaking rate. Finally, the audio embedding and text embedding are serialized and specific tags and position information are added to construct a joint input sequence input Decode-Only Transformer.

[0028] As a further improvement of the present invention, the specific steps of the task planning module converting complex natural language input into a series of executable subtasks are as follows:

[0029] Step 1: Identify the task type and key information from the user's instructions, including intent recognition and entity recognition: Intent recognition: Use the pre-trained big model to perform semantic analysis on the input instructions to determine the user's core task objectives. For the instruction "navigate to the bus stop", the big model can identify the intent as "navigation";

[0030] Entity recognition: Combine the big model and named entity recognition technology to extract the core entity "bus stop" in the instruction and mark it as the target location;

[0031] Step 2: Divide the task into multiple subtasks based on the intent and entity recognition results;

[0032] Step 3: Use a priority queue to sort tasks based on their dependencies and urgency, and support real-time adjustment of priorities in response to user mid-process changes in requirements or emergencies.

[0033] As a further improvement of the present invention, the action module receives the task decomposition structure generated by the task planning module and dynamically constructs a task template to complete the specific steps of actual execution as follows:

[0034] Step 4: The action module receives the task decomposition structure output by the task planning module, generates the corresponding execution template, and is responsible for scheduling and executing subtasks;

[0035] Step 5: The system maintains an interface registry to associate task descriptions with implementation methods. The interface registry associates the requirements of each task with the corresponding implementation interface. The system uses the registry to find the appropriate interface to execute the task.

[0036] Among them, the relevant plug-in functions are called through the interface in the registry. At the same time, the interface registry supports hot updates, provides dynamic addition or modification of interface definitions, and facilitates the integration of new functions;

[0037] Step 6: Through the reflection mechanism, the system dynamically loads classes and methods at runtime, and then the task scheduler is responsible for calling subtasks step by step and processing the returned results according to the dependencies of the task decomposition structure.

[0038] As a further improvement of the present invention, the specific manner of generating the corresponding execution template in step 4 and being responsible for scheduling and executing the subtasks is as follows:

[0039] First, perform subtask description mapping: According to the subtask description in the task decomposition, query the interface registry to find the corresponding method or interface, and then construct the execution template: Combine the subtask requirements and the matching interface to dynamically generate the subtask execution template.

[0040] As a further improvement of the present invention, the specific steps of gradually calling subtasks and processing the returned results in step 6 are as follows:

[0041] ① Sort subtasks according to priority and dependencies;

[0042] ②Call the methods in the task template in sequence to complete the actual execution;

[0043] ③Perform data verification on the returned results and use them as input for subsequent tasks.

[0044] As a further improvement of the present invention, the path planning and navigation module includes:

[0045] The visual processing and depth estimation module is used to process the images collected by the binocular camera, eliminate distortion and generate high-precision disparity maps. The acquisition time and left-right correspondence of the image are determined by the timestamp and 0, 1 flags in the image metadata, and then the depth estimation algorithm is used to convert the disparity data into depth information to generate a three-dimensional point cloud, providing accurate spatial depth support.

[0046] The map generation module uses semantic segmentation to mark pixel categories in the image, combines object detection to extract obstacle locations, and projects the 3D point cloud into a 2D grid map to identify walkable areas and obstacles, providing map support for path planning.

[0047] The path planning and dynamic obstacle avoidance module uses depth information, three-dimensional point cloud information and grid maps to iteratively calculate the shortest path based on the temporary destination, and finally completes the task of navigating to the end point by continuously navigating to the new temporary destination.

[0048] Beneficial effects of the present invention:

[0049] (1) Enhanced environmental perception: Through the visual language big model technology, the system can recognize text, scenes, objects and obstacles in real time, providing the blind with more comprehensive environmental information, effectively improving their perception of the surrounding environment and their ability to protect themselves.

[0050] (2) Accurate obstacle avoidance and navigation support: Combining binocular vision, GPS positioning and shortest path algorithm, the system can provide accurate obstacle avoidance and navigation for the blind. It not only helps them avoid obstacles effectively in complex environments, but also provides real-time voice navigation guidance to ensure safe passage.

[0051] (3) Intelligent task scheduling and personalized services: Based on dynamic task templates and reflection mechanisms, the system can flexibly adjust task scheduling, intelligently match the current needs of the blind, provide personalized auxiliary functions, and optimize their user experience.

[0052] (4) Global path planning optimization: Through the dynamic combination of regional path planning and global path planning, the system can continuously optimize the navigation route according to real-time environmental changes, providing the blind with the safest and most convenient travel route.

[0053] (5) Improving independence and quality of life: The combined application of the above technologies not only helps blind people move more independently and safely in their daily lives, but also enhances their ability to participate in social activities, greatly improving their quality of life and independence.

[0054] The application of these technologies will greatly enhance the confidence and sense of security of the blind in complex environments and help them better integrate into society. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a flowchart of the cloud data processing architecture;

[0056] Figure 2 Schematic diagram of the multimodal feature fusion architecture based on Decoder-Only Transformer;

[0057] Figure 3 It is a schematic diagram of a flat map projected from three dimensions to two dimensions;

[0058] Figure 4 This is a schematic diagram of a grid map;

[0059] Figure 5 A schematic diagram of the temporary path planning area;

[0060] Figure 6 Schematic diagram of an undirected map and candidate temporary destinations;

[0061] Figure 7 Schematic diagram of the temporary optimal navigation path. DETAILED DESCRIPTION

[0062] The present invention will be further described below in detail with reference to the embodiments shown in the accompanying drawings.

[0063] Reference Figure 1 As shown, the binocular vision-based smart glasses for the blind in this embodiment are mainly divided into a glasses terminal and a cloud. In actual use, through the communication between the glasses terminal and the cloud, the blind are assisted to realize the ability of recognizing text, scenes, objects, obstacles, and obstacle avoidance navigation. Therefore, this embodiment further describes the glasses terminal and the cloud.

[0064] The glasses terminal of this embodiment mainly uses a binocular camera as a basic component, collects environmental data through the binocular camera, and obtains real-time location information in combination with a GPS module. The camera data is processed by a lightweight processor, compressed and stored locally to ensure that the data is stably uploaded to the cloud. The cloud uses powerful computing resources to perform complex image processing, obstacle detection, and path planning. The cloud will generate real-time feedback information and return the processing results to the user in the form of voice or vibration. Specifically:

[0065] Binocular camera: collects environmental images and video data through high-resolution cameras. Binocular cameras can provide stereoscopic vision to help blind people identify the depth and distance information of obstacles.

[0066] GPS module: provides real-time geographic location data for the blind, and works with the cloud-based path planning system to help the blind obtain location information of the surrounding environment and navigate.

[0067] Wi-Fi module: Connect to the cloud via wireless network to achieve fast and low-latency data upload. If necessary, a 5G module can be integrated to ensure a stable connection even in an environment without Wi-Fi network.

[0068] External RTC module: Provides accurate time tracking function to ensure that the system can maintain consistent time data during long-term operation. For example, DS3231, a very common and accurate RTC module, has built-in temperature compensation and extremely high accuracy. It communicates with the main control chip (such as ESP32, Arduino, etc.) through the I2C interface.

[0069] Local storage device: Use eMMC or MicroSD card to temporarily store the collected data to avoid data loss due to network fluctuations and ensure that the data can be uploaded in time when the network is restored. This storage device can also handle network disconnection to ensure system stability.

[0070] Lightweight processor: used to extract video keyframes, process image data, calculate timestamps, mark camera IDs, and perform local tasks (such as image compression, data synchronization, etc.). The processor also needs to manage data uploads to the cloud. You can use ESP32 (with Wi-Fi and Bluetooth connectivity) or Raspberry Pi Zero W (with USB interface, Wi-Fi, strong processing power and small size).

[0071] In this embodiment, voice is used to communicate with the blind, so the following voice interaction and feedback module is provided:

[0072] The system receives the blind person's voice commands through a microphone array and sends them to the cloud for analysis, and the cloud returns the corresponding commands, such as obstacle locations, turn prompts, etc. Feedback is transmitted through speakers or bone conduction headphones to ensure that users can hear clearly.

[0073] Microphone array: used to collect the blind person's voice commands and send them to the cloud for processing. The array design can help improve the accuracy of speech recognition and avoid interference from environmental noise.

[0074] Speaker or bone conduction headset: Output information about the location of obstacles, turn prompts, etc. according to the instructions provided by the voice feedback module. Bone conduction headset is suitable for this application because it can avoid interference from external noise while keeping the ear canal open, suitable for blind people to wear.

[0075] Audio processing chip: supports noise reduction and echo suppression to ensure clear and natural voice feedback. Audio quality is crucial for voice interaction systems, especially in complex environments, where clear voice feedback can help blind people respond quickly.

[0076] In order to ensure that each module of the above-mentioned eyewear terminal is powered, a battery and power management system is provided in this embodiment.

[0077] Through efficient battery and power management systems, the device can be ensured to run for a long time, and the fast charging module provides users with a convenient charging method.

[0078] Lithium battery: provides enough power to support the long-term operation of the equipment, and can ensure that the system maintains sufficient battery life when performing tasks such as image acquisition, data uploading, and voice feedback.

[0079] Fast charging module: provides users with a convenient charging method, reduces charging time and increases flexibility of use.

[0080] Power management chip: optimizes overall power consumption and extends battery life, especially during long periods of data uploading and image processing, ensuring efficient system operation.

[0081] From this we can obtain the following data processing method on the glasses side:

[0082] Image data processing

[0083] (1) Image acquisition

[0084] The binocular cameras collect image data in real time. Each camera (left and right) can capture images through its built-in sensor and send the image data to ESP32 for subsequent processing. Based on an external RTC module (such as DS3231), ESP32 reads the external clock through the I2C or SPI interface to ensure the accuracy of the timestamp.

[0085] (2) Metadata Embedding

[0086] Timestamp: Since the image or video data is in binary stream format, ESP32 will embed timestamp information in the metadata part of the image or video frame such as JPEG, PNG or H.264. For static images, the timestamp can be embedded in the EXIF ​​metadata of the image file.

[0087] Device Flags: In the data of each image or video frame, the ESP32 also needs to embed a simple flag indicating whether the image was acquired by the left camera "0" or the right camera "1".

[0088] GPS location information: The geographical location information obtained from the GPS module is embedded in the metadata.

[0089] (3) Image transmission and synchronization

[0090] After the image is captured and marked with a timestamp and camera tag, the processor will upload the image data to the cloud via Wi-Fi or other communication methods (such as Bluetooth). During the transmission process, the image data will be compressed, such as JPEG or H.264, to reduce the transmission bandwidth requirements.

[0091] Voice command processing

[0092] After the voice command is recorded through the microphone, it is pre-processed by noise suppression, encoding, and framing. Then, it is transmitted to the server through the processor based on protocols such as HTTP, WebSocket, or gRPC. The server performs voice recognition and processing.

[0093] Reference Figure 1 As shown, the cloud of this embodiment is mainly composed of a perception module, a task planning module, an action module, and a path planning and navigation module, as follows:

[0094] Perception Module

[0095] The perception module of the blind glasses of the present invention is based on the Decode-Only Transformer architecture, focusing on the efficient parsing of multimodal inputs and the internal representation transformation, and can unify multimodal data such as images and audio into processable internal representations, thereby realizing the perception, understanding and auxiliary decision-making capabilities of the blind user's instructions and the surrounding environment. The specific technical route is as follows:

[0096] (1) Receiving data

[0097] Receive multimodal raw data such as images and audio uploaded by the glasses, extract GPS location information from the metadata of the image, and update the current GPS location information in the system in real time.

[0098] (2) Data preprocessing

[0099] Image data preprocessing: including size adjustment, color standardization, and background removal to ensure uniform image format and extract valid scene information.

[0100] Preprocessing of audio data: including removing environmental noise, sampling and framing, optimizing audio quality, and providing a clearer data basis for audio analysis.

[0101] Through these preprocessing steps, image and audio data are converted into normalized input forms, which improves the processing efficiency and accuracy of the model.

[0102] (3) Feature Coding

[0103] The preprocessed data needs to go through a feature extraction process to map the multimodal data into a unified vector representation:

[0104] Image feature extraction: Based on ViT (Vision Transformer), the image is cut into fixed-size patches (such as 16×16). Each patch is mapped to a high-dimensional vector through linear transformation and a special marker is added ( ) and position encoding to ensure that the model perceives the spatial information of the tile. All tile embedding vectors are arranged in order to generate an input sequence. Finally, the image embedding feature sequence is provided as input to the Decode-Only Transformer and processed together with text or other modal data to support joint decoding of multimodal tasks.

[0105] Audio feature extraction: First, the audio signal is sampled, framed, and feature extracted to generate a multi-dimensional feature matrix such as MFCC or Log-Mel Spectrogram. Subsequently, a pre-trained audio model (such as Wav2Vec 2.0 or HuBERT) is used to extract a high-dimensional embedding representation and align it with the text features through a dimensionality reduction mechanism. The ASR module is used to transcribe the audio signal into text. After the audio embedding and text embedding are serialized, specific tags and position information are added to construct a joint input sequence input Decode-Only Transformer.

[0106] (4) Multimodal feature alignment

[0107] In multimodal data processing, feature alignment is a key step to ensure that data from different modalities can be represented in the same vector space. By sharing encoding dimensions or performing linear mapping, feature vectors of multiple modalities such as images and audio are standardized into a unified representation space. The goal of feature alignment is to enable the perception module of blind glasses to effectively understand the association between different modalities and enhance the comprehensive understanding of the environment by fusing the features of these data forms.

[0108] (5) Multimodal feature fusion

[0109] Based on feature alignment, the Decode-Only Transformer architecture is used to further process these multimodal feature vectors. The Transformer decoding layer effectively models and parses cross-modal information by gradually extracting context-related deep features. The self-attention mechanism can capture the mutual influence between different modalities to generate a unified internal representation. This representation not only contains contextual information, but also integrates the semantic content of each modality. It has a high degree of versatility and interpretability, ensuring the accurate understanding and effective integration of cross-modal data, such as Figure 2 shown.

[0110] By fusing the personalized background information of the portrait configuration with multimodal data, the perception module can optimize task processing according to the specific needs of blind users.

[0111] (6) Output data

[0112] It outputs a unified multimodal semantic representation vector that integrates text, image, audio and other features, and has consistency, context awareness and multimodal semantic fusion capabilities.

[0113] Mission Planning Module

[0114] The task planning module is fine-tuned based on pre-trained large language models (such as BERT, GPT) and blind domain data, combined with a knowledge base in the blind domain, so that it has the ability to extract intent from natural language instructions and dynamically generate task decomposition structures. Its core goal is to convert complex natural language input into a series of executable subtasks and clarify their dependencies.

[0115] (1) Task understanding

[0116] The goal of task understanding is to identify the task type and key information from user instructions, which mainly includes intent recognition and entity recognition:

[0117] Intent recognition: Use the pre-trained big model to perform semantic analysis on input commands to determine the user’s core task objectives. For example, for the command “navigate to the bus stop”, the big model can identify the intent as “navigation”.

[0118] Entity recognition: Combine the large model and named entity recognition (NER) technology to extract the core entity "bus stop" in the instruction and mark it as the target location.

[0119] (2) Task decomposition

[0120] The task decomposition module divides the task into multiple subtasks based on the intent and entity recognition results.

[0121] ① Dynamic decomposition logic: By fine-tuning the pre-trained large model, we can obtain a task decomposition model, complete task decomposition, and generate a fine-grained task structure. For example, "navigating to the bus stop" can be decomposed into the following subtasks:

[0122] Get current location: Call the device GPS interface to obtain the current location longitude and latitude.

[0123] Get destination information: query the latitude and longitude of the target location "bus stop" through public map APIs (such as Amap and Google Maps).

[0124] Path planning and navigation: Take the current location and destination as input and call the custom navigation interface to complete path planning.

[0125] ② Output structure: The task decomposition results are output in the form of structured data, including subtask descriptions, input requirements, output requirements, and dependencies between subtasks. For example:

[0126]

[0127]

[0128] (3) Task Priority Arrangement

[0129] Dynamic priority evaluation: Use priority queues to sort tasks based on their dependencies and urgency. For example, "get current location" is a prerequisite for subsequent tasks and is executed first.

[0130] Real-time adjustment: In response to users' mid-process changes in requirements or emergencies, real-time adjustment of priorities is supported to ensure the dynamism and flexibility of tasks.

[0131] Action Module

[0132] The action module is responsible for receiving the task decomposition structure generated by the task planning module and dynamically constructing the task template to complete the actual execution. The generation of the task template is completely dependent on the matching of the subtask description in the task decomposition structure and the interface registry.

[0133] Dynamic task template generation

[0134] The action module receives the task decomposition structure output by the task planning module, generates the corresponding execution template, and is responsible for scheduling and executing subtasks.

[0135] Subtask description mapping: According to the subtask description in the task decomposition (such as "get the current position"), query the interface registry to find the corresponding method or interface, such as:

[0136] PositionService.getCurrentPosition()

[0137] Execution template construction: Combine the subtask requirements (input and output) and the matching interface to dynamically generate the subtask execution template. For example:

[0138]

[0139] Interface Registry

[0140] Interface definition and registration: The system maintains an interface registry to associate task descriptions with implementation methods. The structure of the interface registry can be a dictionary or a mapping, which includes the following key fields:

[0141] Task description: the description of the corresponding task or the task name (for example: get the current location). Method name: The method or interface called when the task is executed (for example:

[0142] PositionService.getCurrentPosition).

[0143] Input parameters: The input parameters and types required by the method.

[0144] Output result: the output result and type after the method is executed.

[0145] Parameter validation: validation rules for input parameters and output results, such as whether it is a required field, type validation, etc.

[0146] For example:

[0147] Taking "get destination location" as an example, the interface registry entry may be as follows:

[0148]

[0149] }

[0150] ],

[0151] "Parameter Validation":{

[0152] "device_id":{"required":true,"type":"String"}

[0153] }

[0154] }

[0155] The interface registry associates each task requirement with the corresponding implementation interface (method or class). The system uses the registry to find the appropriate interface to perform the task. For example, the "get current position" requirement may be associated with the getCurrentPosition() method in the PositionService class, while the "plan route" requirement may be associated with the calculateRoute() method in the RoutePlanner class.

[0156] Extensibility support: Through the interface in the registry, you can call related plug-in functions. At the same time, the interface registry supports hot updates, which can dynamically add or modify interface definitions to facilitate the integration of new functions.

[0157] Reflection Mechanism

[0158] Dynamic loading: Through the reflection mechanism, the system dynamically loads classes and methods at runtime. Taking Java as an example, Class.forName() and Method.invoke() are used to dynamically call interfaces.

[0159] Interface security verification: During the loading process, the interface input and output parameters are verified to avoid data format or security issues.

[0160] Task Scheduler

[0161] The task scheduler is responsible for calling subtasks step by step and processing the returned results according to the dependencies of the task decomposition structure. Scheduling logic:

[0162] ① Sort subtasks according to priority and dependencies.

[0163] ②Call the methods in the task template in sequence to complete the actual execution.

[0164] ③Perform data verification on the returned results and use them as input for subsequent tasks.

[0165] Data returned to smart glasses: Based on the execution results and the generation capability of the visual language large model, the system can generate relevant text and convert it into audio and instructions and return it to the glasses to guide the blind to complete relevant tasks, including knowledge questions, scene recognition, text recognition, obstacle recognition, object recognition, etc.

[0166] Exception handling: Provides a complete exception capture mechanism to retry or roll back failed tasks and provide real-time error feedback.

[0167] Path planning and navigation module

[0168] When the task manager performs a navigation task, this module will receive the current longitude and latitude information and the longitude and latitude information of the destination, and obtain the collected image information in real time. Through the three sub-modules of visual processing and depth estimation, map generation and path planning and dynamic obstacle avoidance, fine-grained obstacle avoidance navigation is achieved.

[0169] Visual processing and depth estimation algorithm: Process the images collected by the binocular camera, eliminate distortion and generate high-precision disparity maps. The timestamp and 0, 1 flags in the image metadata are used to determine the acquisition time and left-right correspondence of the image. The depth estimation algorithm is used to convert disparity data into depth information, generate a three-dimensional point cloud, and provide accurate spatial depth support.

[0170] Map generation module: Through semantic segmentation, the pixel categories in the labeled image are extracted by combining target detection, and the 3D point cloud is projected into a 2D grid map to identify the walkable area and obstacles, providing map support for path planning.

[0171] Path planning and dynamic obstacle avoidance module: Utilizes depth information and grid maps to calculate the shortest path and generate a complete navigation path. Through real-time image analysis, the path is dynamically adjusted to avoid obstacles, ensuring a safe and efficient navigation process.

[0172] Visual processing and depth estimation algorithm module

[0173] (1) Stereo calibration

[0174] Based on the image data from the data receiving and preprocessing module, in order to ensure the alignment of the images generated by the binocular camera system and eliminate distortion, the camera is calibrated using a chessboard diagram to obtain the intrinsic matrix (focal length, optical center coordinates, etc.) and the extrinsic matrix (rotation and translation matrix).

[0175] Correct distortion: Eliminate barrel or pincushion distortion to ensure consistent field of view for left and right images.

[0176] Epipolar constraint: Through binocular geometric correction, the two images are corrected so that their corresponding pixels are aligned in the horizontal direction. Specifically, the pixel rows in the corrected image will be parallel, ensuring that the pixels in each row have the same depth information during stereo matching, which is convenient for calculating the depth map through disparity.

[0177] (2) Stereo Matching

[0178] The SGBM algorithm (Semi-Global Block Matching) is used to generate a disparity map by pixel block matching to infer depth information.

[0179] Input: left-right calibrated image.

[0180] Output: The disparity value d for each pixel, which represents the difference in the position of the same object in the image captured by the two cameras, usually in pixels.

[0181] Parameter optimization: adjust the window size and penalty coefficient to balance accuracy and computational efficiency.

[0182] (3) Depth Estimation

[0183] The depth here refers to the distance from the object surface to the camera, usually along the optical axis of the camera. For each pixel, the depth D is calculated based on the disparity value d and the camera baseline distance B and focal length f:

[0184]

[0185] D: Depth of pixel (unit: meter).

[0186] f: Focal length of the camera, usually in millimeters or meters.

[0187] B: The baseline distance of the binocular camera, that is, the physical distance between the two cameras.

[0188] d: disparity value.

[0189] Through this formula, we can convert the disparity value of each pixel in the two-dimensional image into the actual depth information in the scene, thereby generating a depth map that represents the real-world distance corresponding to each pixel.

[0190] (4) 3D point cloud generation

[0191] According to the depth map and pixel position, each pixel point is converted to the actual three-dimensional space coordinates. Specifically, the generation of a three-dimensional point cloud requires converting the depth value (the depth information corresponding to each pixel in the depth map) and the two-dimensional position of the pixel into a point in three-dimensional space through a projection model.

[0192] Based on the pinhole camera model, assume that the depth value of a given pixel (x, y) is D (x,y) , then the three-dimensional coordinates (X, Y, Z) corresponding to the pixel point can be calculated using the following camera projection model:

[0193]

[0194] x,y: The position of the pixel in the image, representing the horizontal and vertical coordinates in the image coordinate system.

[0195] C x ,C y:Image optical center coordinates, which are the positions of the image optical center in the image coordinate system. This is usually obtained during camera calibration and represents the position of the camera center.

[0196] f: is the focal length of the camera, usually in pixels. The focal length determines the size of the field of view and how objects are mapped in the image.

[0197] D (x,y) : The depth value of the pixel (x, y) in the depth map, indicating the actual distance from the pixel to the camera (unit: meter).

[0198] (X,Y,Z): is the coordinate of the pixel in three-dimensional space, in meters.

[0199] Through the above formula, the depth information D of each pixel can be (x,y) The corresponding image coordinates (x, y) are converted to three-dimensional coordinates (X, Y, Z). For each pixel in the image, a corresponding three-dimensional point can be obtained. These three-dimensional point sets constitute the three-dimensional point cloud of the entire image.

[0200] Map generation module

[0201] (1) Obstacle recognition and walkable area extraction method based on semantic segmentation and target detection

[0202] This method uses semantic segmentation and object detection models in parallel, combined with the post-processing stage of image processing, to extract obstacle information and walkable areas in the image. The specific steps are as follows:

[0203] ① Image input: The input image is passed to two models at the same time - the semantic segmentation model and the object detection model, and the image is processed in parallel. The semantic segmentation model will give the category label of each pixel, while the object detection model will identify the specific location of the object in the image.

[0204] ②Semantic segmentation: Use advanced semantic segmentation models (such as DeepLabV3+ or FCN) to perform semantic segmentation on the input image and generate category information for each pixel in the image. The goal of semantic segmentation is to classify the scene pixel by pixel to obtain segmentation masks for walkable areas and obstacle areas.

[0205] Walkable area: Pixels in the image are marked as “0”.

[0206] Obstacle area: Pixels in the image are marked as "1", such as buildings, obstacles on the road, etc.

[0207] ③ Target detection: Use a target detection model (such as YOLOv5) to detect the specific location of obstacles and obtain the bounding boxes and category labels (such as pedestrians, vehicles, etc.) of each obstacle.

[0208] ④ Post-processing and fusion: After obtaining the results of semantic segmentation and target detection, the two are effectively fused in the post-processing stage to ensure the accuracy of obstacle recognition and the integrity of the walkable area, as follows:

[0209] Merge segmentation mask and object detection results: Based on the mask obtained by semantic segmentation, the obstacle area is enhanced or corrected through the obstacle bounding box information detected by the object detection model. Specifically, if the object detection identifies an obstacle, the pixels within its bounding box area will be considered as an obstacle, and even if the semantic segmentation does not completely mark the area as an obstacle, it will be corrected through the bounding box.

[0210] Obstacle filtering: If obstacles detected by object detection overlap or are close to obstacles in semantic segmentation, these areas are considered obstacles and are treated as obstacle avoidance areas in subsequent path planning.

[0211] Corrected bounding box: Based on the bounding box obtained by target detection, the obstacle bounding box is further corrected using the semantic segmentation results. Specifically, if there is a walkable area outside the bounding box and an unmarked obstacle area inside, the semantic segmentation mask can be used to expand or adjust the bounding box to ensure that the target detection box accurately covers the complete area of ​​the obstacle to avoid missed detection.

[0212] ⑤ Parameter tuning: The model parameters of semantic segmentation and object detection (such as learning rate, training dataset size, network architecture, etc.) need to be tuned through experiments to obtain the best accuracy and efficiency in different scenarios.

[0213] (2) 3D to 2D projection

[0214] ① Project the 3D point cloud onto a 2D plane (bird’s eye view)

[0215] Only X,Y information is kept, ignoring Z (height) to produce a flat map.

[0216] In the two-dimensional plane, each point is marked as a walkable area or obstacle based on the depth value and semantic segmentation results, see Figure 3 shown.

[0217] ②Generate a grid map

[0218] The point projection result on the two-dimensional plane is converted into a grid map (two-dimensional matrix), as follows:

[0219] Grid division: Divide the two-dimensional plane into grids at a certain resolution (e.g., k pixels). Here, "k" represents the size of the grid unit, that is, the actual space area covered by each grid (e.g., 1 grid = 0.5m x 0.5m, this value can be set dynamically). By controlling the size of the grid (k pixels), you can balance accuracy and computational efficiency.

[0220] Obstacle classification: For each grid cell, check whether there are any obstacles (such as buildings, vehicles, etc.) overlapping with it. If so, the grid is marked as "non-walkable area" with a value of 1. If not, it is marked as "walkable area" with a value of 0. Each grid cell records the corresponding area type and obstacle-related information. Based on this information, the distance between the area and the obstacle can be obtained.

[0221] Calculation of longitude and latitude of the center of the grid unit: Use the depth map and the internal and external parameters of the camera to calculate the position of each pixel in the world coordinate system. Then, based on the world coordinates (X, Y, Z) of each pixel, use the GPS position of the camera and the coordinate conversion method to calculate the longitude and latitude of the point. Finally, based on the grid size (for example, k pixels), the longitude and latitude of the center point of each grid unit can be inferred. Figure 4 shown.

[0222] Path planning and dynamic obstacle avoidance module

[0223] (1) Set the path planning area

[0224] Set the depth M meters in front of the camera as the single path planning depth. M is recommended to be 3-6 meters, which can be obtained based on statistical analysis results and can be set as needed through the background. Through the preset ratio, you can get the pixel length m corresponding to M meters in the grid map. Construct the area with a vertical distance of less than m pixels from the horizontal line where the camera is located as a temporary path planning area, refer to Figure 5 shown.

[0225] (2) Construct an undirected graph of the walkable area and set candidate temporary destinations

[0226] The feasible region grid is regarded as a node, and edges are built for grids with common edges. Let it be a graph G(E,V), where E represents the edge set and V represents the node set. The following graph structure can be obtained. The walkable nodes on the edge line are set as candidate temporary destinations, as shown in the dark nodes in the figure (b) below, and are denoted as set TP. In this example, TP = {v 1 ,v 2 ,v 3 ,v4 ,v 5 ,v 6 ,v 7 ,v 14 ,v 15 ,v 25 ,v 26 ,v 27 ,v 28 ,v 29 ,v 30 ,v 31 ,v 32 ,v 33}, refer to Figure 6 shown.

[0227] (3) Calculate the shortest path from S to the candidate temporary destination

[0228] Since the scaling ratio is consistent, the side length can be recorded as 1, and the shortest path S is obtained min After that, multiply it by k to get the actual pixel distance, and then multiply it by the length M / m represented by each pixel to get the actual length value. i The distance is denoted as dis(S,v i ), which can be obtained through the shortest path algorithm (such as A* algorithm, Spfa, Dijkstra algorithm, etc.).

[0229] (4) Calculate the shortest path from the candidate temporary destination to the actual destination point T

[0230] Based on the longitude and latitude of the candidate temporary destination and the longitude and latitude of the actual destination T, the shortest distance from each candidate temporary destination to the target location T is obtained through the map API. i The distance to T is denoted as DIS(v i ,T).

[0231] (5) Calculate the temporary destination and the optimal temporary navigation path

[0232] Traverse TP, add the distance from the candidate temporary destination temp (temp∈TP) to the starting point S and the distance from temp to the end point T, and get the shortest distance from S to T through temp. Select the point with the shortest total distance as the temporary target point P. The determination and update process of the temporary target point can be implemented by the following formula:

[0233]

[0234] Record S min The temporary destination when is the minimum is P. Assume that node v 28 is a temporary destination, then S->v 4 ->v 10 ->v19 ->v 18 ->v 17 ->v 28 and S->v 4 ->v 10 ->v 19 ->v 18 ->v 29 ->v 28 All of them are optional paths. You can choose any one of them as the final temporary navigation path. Figure 7 shown.

[0235] (6) Direction calculation

[0236] By calculating the direction vectors of two consecutive points in the path, using the inverse tangent function to find the azimuth, and correcting it to the range of 0°-360°, the angle is updated in real time to guide the walking direction, ensuring the accuracy of path planning and dynamic adjustment capabilities. The blind can be reminded to adjust the direction through voice broadcasts, vibration reminders, etc.

[0237] (7) Dynamic obstacle avoidance path planning

[0238] The similarity of continuously collected pictures is compared through the visual big model. If the similarity is lower than a certain threshold, such as 95%, the calculation is re-expanded and the path is planned to achieve the final navigation from the starting point to the end point.

[0239] In addition, large models can be combined to detect moving obstacles and predict their trajectories, thereby achieving more accurate and complete obstacle avoidance navigation.

[0240] In summary, the binocular vision-based smart glasses for the blind in this embodiment:

[0241] (1) Blind-assisted recognition system based on visual language large model technology: By integrating vision and language processing technology, the system can help blind people recognize surrounding text, scenes, objects and obstacles in real time. Combined with multimodal feedback such as voice and vibration, it provides real-time environmental perception and obstacle avoidance support, thereby greatly improving the blind people's independent action ability and safety in complex environments.

[0242] (2) Task scheduling and interface matching technology based on dynamic task templates and reflection mechanisms: Through the design of dynamic task templates and reflection mechanisms, this technology can flexibly schedule and match the current task requirements and environmental changes of blind users. The system can automatically select the optimal functional module or service interface to ensure efficient and accurate auxiliary support in different situations.

[0243] (3) Temporary regional path planning and obstacle avoidance navigation technology for blind glasses based on binocular vision + GPS and map API: This technology combines binocular vision and GPS positioning data to achieve real-time temporary regional path planning and obstacle avoidance navigation for blind glasses in complex environments. Through continuous iteration of regional path planning, it can dynamically update and optimize the global path. With the help of map API and shortest path algorithm, the system can dynamically calculate the optimal path and provide users with voice navigation guidance to ensure their safe passage in complex or dynamically changing environments.

[0244] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A pair of smart glasses for the blind based on binocular vision, characterized in that: include: The eyeglass terminal has a built-in data collection and transmission module, a voice interaction and feedback module, and a battery and power management system. The data collection and transmission module collects data, processes the data and outputs it, and the voice interaction and feedback module receives the voice commands of the blind and simultaneously feeds back prompt commands to the blind. The cloud has built-in perception module, task planning module, action module and path planning and navigation module. The perception module has the ability to perceive, understand and assist in decision-making for blind users and the surrounding environment. The task planning module has the ability to extract intent from natural language instructions and dynamically generate task decomposition structure, so as to convert complex natural language input into a series of executable subtasks and clarify their dependencies. The action module is responsible for receiving the task decomposition structure generated by the task planning module and dynamically constructing task templates to complete the actual execution. The path planning and navigation module implements the three sub-modules of visual processing and depth estimation, map generation and path planning and dynamic obstacle avoidance to achieve fine-grained obstacle avoidance navigation.

2. The binocular vision-based smart glasses for the blind according to claim 1, characterized in that: The data acquisition and transmission module includes: A binocular camera is used to collect environmental images and video data to provide stereoscopic vision, helping the blind to identify the depth and distance information of obstacles; GPS module, used to provide real-time geographic location data of the blind; Wi-Fi module, used to connect to the cloud via wireless network; External RTC module, used to provide accurate time tracking function, ensuring that the system can maintain consistent time data during long-term operation; Local storage device, temporarily storing collected data through eMMC or MicroSD card; A lightweight processor used to extract video key frames, process image data, calculate timestamps, mark camera IDs, and perform local tasks.

3. The binocular vision-based smart glasses for the blind according to claim 1 or 2, characterized in that: The voice interaction and feedback module includes: A microphone array is used to collect the blind person’s voice commands and send them to the cloud for processing; A speaker or bone conduction headset is used to output information about the location of obstacles and steering prompts according to the instructions provided by the voice feedback module; The audio processing chip is connected to the microphone array and the speaker or the bone conduction headset to perform noise reduction and echo suppression on the audio.

4. The binocular vision-based smart glasses for the blind according to claim 1 or 2, characterized in that: The perception module unifies multimodal data such as images and audio into a processable internal representation based on the Decode-Only Transformer architecture, as follows: Step 1: Receive the multimodal raw data such as images and audio uploaded by the glasses, extract the GPS location information from the metadata of the image, and update the current GPS location information in the system in real time; Step 2: Preprocess the image data, including resizing, color standardization, and background removal to ensure uniform image format and extract effective scene information; and preprocess the audio data, including removing environmental noise, sampling, and framing to optimize audio quality and provide a clearer data basis for audio analysis. Step 3: extract features from the data preprocessed in step 2 and map the multimodal data into a unified vector representation; Step 4: Standardize the feature vectors of multiple modalities such as images and audio into a unified representation space by sharing the encoding dimension or performing linear mapping; Step 5: Based on feature alignment, the Decode-Only Transformer architecture is used to further process these multimodal feature vectors, and the mutual influence between different modalities is captured through the self-attention mechanism to generate a unified internal representation. The personalized background information of the portrait configuration is then fused with the multimodal data. Step six, output the unified multimodal semantic representation vector fused in step five.

5. The binocular vision-based smart glasses for the blind according to claim 4, characterized in that: The specific method of feature extraction in step 3 is as follows: Image feature extraction is performed. Based on ViT, the image is cut into fixed-size tiles. Each tile is then mapped to a high-dimensional vector through a linear transformation, and special tags and position encodings are added to ensure that the model perceives the spatial information of the tile. All tile embedding vectors are then arranged in order to generate an input sequence. Finally, the image embedding feature sequence is provided as input to the Decode-Only Transformer, which is processed together with text or other modal data to support joint decoding of multimodal tasks. Audio feature extraction,First, the audio signal is sampled, framed and feature extracted to generate a multi-dimensional feature matrix such as MFCC or Log-MelSpectrogram. Then, a pre-trained audio model is used to extract a high-dimensional embedding representation and align it with the text features through a dimensionality reduction mechanism. Then, the ASR module is used to transcribe the audio signal into text. Finally, the audio embedding and text embedding are serialized and specific tags and position information are added to construct a joint input sequence input Decode-OnlyTransformer.

6. The binocular vision-based smart glasses for the blind according to claim 1 or 2, characterized in that: The specific steps of the task planning module to convert complex natural language input into a series of executable subtasks are as follows: Step 1: Identify the task type and key information from the user's instructions, including intent recognition and entity recognition: Intent recognition: Use the pre-trained big model to perform semantic analysis on the input instructions to determine the user's core task objectives. For the instruction "navigate to the bus stop", the big model can identify the intent as "navigation"; Entity recognition: Combine the big model and named entity recognition technology to extract the core entity "bus stop" in the instruction and mark it as the target location; Step 2: Divide the task into multiple subtasks based on the intent and entity recognition results; Step 3: Use a priority queue to sort tasks based on their dependencies and urgency, and support real-time adjustment of priorities in response to user mid-process changes in requirements or emergencies.

7. The binocular vision-based smart glasses for the blind according to claim 6, characterized in that: The action module receives the task decomposition structure generated by the task planning module and dynamically constructs a task template to complete the specific steps of actual execution as follows: Step 4: The action module receives the task decomposition structure output by the task planning module, generates the corresponding execution template, and is responsible for scheduling and executing subtasks; Step 5: The system maintains an interface registry to associate task descriptions with implementation methods. The interface registry associates the requirements of each task with the corresponding implementation interface. The system uses the registry to find the appropriate interface to execute the task. Among them, the relevant plug-in functions are called through the interface in the registry. At the same time, the interface registry supports hot updates, provides dynamic addition or modification of interface definitions, and facilitates the integration of new functions; Step 6: Through the reflection mechanism, the system dynamically loads classes and methods at runtime, and then the task scheduler is responsible for calling subtasks step by step and processing the returned results according to the dependencies of the task decomposition structure.

8. The binocular vision-based smart glasses for the blind according to claim 7, characterized in that: The specific method of generating the corresponding execution template in step 4 and being responsible for scheduling and executing subtasks is as follows: First, perform subtask description mapping: According to the subtask description in the task decomposition, query the interface registry to find the corresponding method or interface, and then construct the execution template: Combine the subtask requirements and the matching interface to dynamically generate the subtask execution template.

9. The binocular vision-based smart glasses for the blind according to claim 7, characterized in that: The specific steps of calling subtasks step by step and processing the returned results in step 6 are as follows: ① Sort subtasks according to priority and dependencies; ②Call the methods in the task template in sequence to complete the actual execution; ③Perform data verification on the returned results and use them as input for subsequent tasks.

10. The binocular vision-based smart glasses for the blind according to claim 1 or 2, characterized in that: The path planning and navigation module includes: The visual processing and depth estimation module is used to process the images collected by the binocular camera, eliminate distortion and generate high-precision disparity maps. The acquisition time and left-right correspondence of the image are determined by the timestamp and 0, 1 flags in the image metadata, and then the depth estimation algorithm is used to convert the disparity data into depth information to generate a three-dimensional point cloud, providing accurate spatial depth support. The map generation module uses semantic segmentation to mark pixel categories in the image, combines object detection to extract obstacle locations, and projects the 3D point cloud into a 2D grid map to identify walkable areas and obstacles, providing map support for path planning. The path planning and dynamic obstacle avoidance module uses depth information, three-dimensional point cloud information and grid maps to iteratively calculate the shortest path based on the temporary destination, and finally completes the task of navigating to the end point by continuously navigating to the new temporary destination.

Citation Information

Cited By

  • Blind guiding robot interactive navigation system combining vision, inertial navigation and voice

    CN121346774A

  • A blind guiding robot interactive navigation system combining vision, inertial navigation and voice

    CN121346774B

  • AI intelligent glasses and interaction method thereof

    CN121478127A

  • Multi-mode blind-assisting indoor navigation method and system

    CN121612280A

  • Method for assisting blind person in daily life through intelligent glasses voice

    CN121861317A