Intelligent glasses based on multi-modal vision-language model and environment perception method
Smart glasses using a multimodal vision-language model enable accurate recognition of complex environments and real-time voice prompts, solving the problems of insufficient environmental perception and real-time performance of existing devices, and improving the travel safety and self-care ability of visually impaired people.
Patent Information
- Application Number
- CN202510974944.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-31
AI Technical Summary
Existing assistive devices for the visually impaired have limited environmental perception capabilities, cannot provide rich environmental semantic descriptions, lack real-time performance, have poor user experience, and are not very versatile, making it difficult to meet the travel needs of visually impaired individuals in complex environments.
The smart glasses, based on a multimodal vision-language model, combine a camera, edge computing unit, bone conduction audio unit, and touch interaction unit to achieve accurate recognition of obstacles, signs, and text and provide real-time voice prompts. The lightweight model and hardware acceleration technology improve processing speed and dynamically optimize parameters to adapt to different environments.
It provides high-precision environmental descriptions and real-time navigation guidance, improving the travel safety and self-care ability of visually impaired people, adapting to multiple application scenarios, reducing information fatigue, extending battery life, and improving emergency avoidance capabilities.
Smart Images

Figure CN120859816A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of assistive technology for the visually impaired, specifically to a smart glasses and environmental perception method based on a multimodal vision-language model. Background Technology
[0002] As a special group, blind people face numerous difficulties in daily life, including access to information, due to the lack of visual senses. It is estimated that there are approximately 285 million visually impaired people worldwide, of whom 36 million are completely blind. Visually impaired individuals need to overcome multiple challenges in their daily lives, such as environmental perception, obstacle recognition, and path planning, which severely impacts their ability to live independently and their quality of life.
[0003] Traditional assistive devices for the visually impaired mainly include white canes and guide dogs. While widely used, white canes have a limited detection range (generally no more than 1.5 meters), making them unable to detect obstacles at head height in a timely manner and providing comprehensive information about complex environments. Guide dogs are expensive to train and maintain, and their use is restricted in certain public places. In recent years, with the development of electronic technology and artificial intelligence, some intelligent assistive devices have begun to enter the market.
[0004] Currently, the main types of intelligent assistive devices on the market are as follows: First, obstacle avoidance glasses based on ultrasonic sensors, which detect obstacles by emitting and receiving ultrasonic waves, but cannot identify obstacle types or environmental semantic information; second, navigation devices based on simple computer vision, such as smart glasses equipped with a monocular camera, which can perform basic image recognition, but the accuracy is not high in complex environments; and third, remote assistive products, which transmit real-time images to a remote human operator through a camera, allowing the operator to provide voice guidance, but such services rely on human labor, are costly, and have prominent privacy protection issues.
[0005] Based on comprehensive market analysis, existing technologies generally suffer from the following shortcomings: First, limited environmental perception capabilities; most devices can only identify simple obstacles and cannot provide rich environmental semantic descriptions. Second, insufficient real-time performance; processing delays lead to delayed obstacle avoidance commands. Third, poor user experience; mechanical alarm prompts are not natural enough and can easily cause fatigue after prolonged use. Fourth, limited versatility; most devices are only suitable for specific scenarios, such as outdoor roads or indoor environments, lacking all-scenario adaptability. Therefore, developing smart glasses that can provide accurate, real-time, and natural environmental perception services to assist visually impaired individuals in improving their travel safety and self-care abilities has significant practical value and social significance. Summary of the Invention
[0006] To address one or more shortcomings of the existing technologies, this invention provides a smart glasses and environmental perception method based on a multimodal vision-language model. This method can accurately identify multiple elements such as obstacles, signs, and text in complex environments and generate voice prompts that conform to human language habits. It can provide visually impaired people with high-precision environmental descriptions and real-time navigation guidance, effectively improving their travel safety and self-care ability.
[0007] To achieve the above objectives, the present invention adopts one or more of the following technical solutions:
[0008] Firstly, a smart glasses based on a multimodal vision-language model is provided, comprising:
[0009] The main body of the glasses;
[0010] A camera unit is disposed on the front frame of the glasses body for real-time acquisition of images of the environment in front of the wearer;
[0011] An edge computing unit is electrically connected to the camera unit. The edge computing unit has a built-in lightweight multimodal vision-language model, which is used to preprocess the acquired environmental images, perform semantic analysis, and generate environmental description information.
[0012] A bone conduction audio unit, electrically connected to the edge computing unit, is used to broadcast the environmental description information to the wearer in the form of voice.
[0013] A touch interaction unit is disposed on the main body of the glasses and is used to receive touch or gesture commands from the wearer;
[0014] The power management unit is used to power each functional unit and perform dynamic power consumption optimization.
[0015] The wireless communication unit is used to enable bidirectional data interaction and model updates between the edge computing unit and an external mobile terminal or cloud.
[0016] As a further implementation, the edge computing unit includes:
[0017] The preprocessing module is used to perform denoising, enhancement, and cropping on the environmental image;
[0018] A lightweight multimodal vision-language model module is developed, employing quantization and knowledge distillation techniques to compress parameter size.
[0019] The scene analysis module is used to output structured scene information, including object category, location, and hazard level information.
[0020] The priority judgment module is used to calculate weights and filter high-priority events based on the risk index R, distance D, and user attention factor U.
[0021] The semantic generation module is used to convert the structured scene information of the high-priority events into semantic description text;
[0022] A speech synthesis module is used to synthesize the semantic description text into a speech stream.
[0023] As a further implementation, the horizontal field of view of the camera unit is not less than 120°, and the resolution is not less than 1600×1200@30fps.
[0024] As a further implementation, the touch interaction unit supports click, double click, swipe and long press gestures to achieve mode switching, voice repeat, speech speed adjustment and emergency help.
[0025] As a further implementation, the edge computing unit is also connected to an FPGA coprocessor for hardware acceleration of JPEG-XS encoding and decoding.
[0026] As a further implementation method, it also includes:
[0027] A 60GHz millimeter-wave radar unit is used to detect rapidly approaching objects, including vehicles and pedestrians.
[0028] A vibration motor unit is disposed in the main body of the glasses and electrically connected to the 60GHz millimeter-wave radar unit, used to vibrate rapidly approaching objects at high frequency.
[0029] As a further implementation method, it also includes:
[0030] An ultrasonic sensor unit, electrically connected to an edge computing unit, is used to detect transparent obstacles and cover the camera's blind spot;
[0031] The ToF sensor unit is electrically connected to the edge computing unit and is used to detect the difference in ground elevation.
[0032] On the other hand, an environment perception method based on a multimodal vision-language model is provided. This method is executed using smart glasses based on the multimodal vision-language model as described above, and includes the following steps:
[0033] S1. Real-time acquisition of environmental images using a camera unit;
[0034] S2. Preprocessing: Denoising, enhancing, and cropping the environmental image;
[0035] S3. Based on a lightweight multimodal vision-language model, perform cross-modal reasoning on the preprocessed image to obtain scene semantic vectors;
[0036] S4. Based on the scene semantic vector, identify and output the object category, spatial location, and hazard level;
[0037] S5. Hazard level processing: Select n high-priority information items according to the weight formula P(i)=w1R+w2D+w3U;
[0038] S6. Convert the priority information into semantic description text;
[0039] S7. The semantic description text is converted into a speech stream and played through a bone conduction audio unit.
[0040] As a further implementation, step S7 also includes:
[0041] S8. Collect user feedback through touch interaction unit or voice command;
[0042] S9. When the wireless communication unit is connected to the network, the user feedback is uploaded to the cloud for incremental fine-tuning.
[0043] S10. The updated model is then distributed to the edge computing unit via OTA to achieve online adaptive learning.
[0044] As a further implementation, in step S3, the multimodal vision-language model (Qwen-VL model) is cropped to remove irrelevant modal branches, the processed image is input into the Qwen-VL model, a multidimensional semantic vector is output, and spatial attention weights are added; the ToF sensor unit is used to identify elevation differences and mark key regions.
[0045] As a further implementation method, in step S4, the criteria for judging the hazard level are defined according to category and distance, specifically from high to low as follows:
[0046] If the category is vehicles and the distance is less than 200cm, or if the category is steps and the distance is less than 50cm, the danger level is level 3;
[0047] If the category is pedestrian and the distance is less than 100cm, the danger level is level 2;
[0048] If the obstacle is classified as static and the distance is less than 30cm, the hazard level is Level 1.
[0049] In other cases, the danger level is 0, meaning it is safe.
[0050] As a further implementation method, step S5 specifically includes:
[0051] Input the parsed scene data, including object categories, spatial orientation, and hazard level;
[0052] Extract parameters such as risk index R, distance D, and user attention factor U from the scene data;
[0053] The priority score P is calculated using the following weighting formula:
[0054] P(i) = w1R + w2D + w3U;
[0055] Compare the priority score P with the set threshold T:
[0056] If P≥T, the data in this scenario has high priority; the data information is retained and sorted.
[0057] If P < T, the data for this scenario has low priority and is filtered.
[0058] As a further implementation, in step S5, when the risk index R is higher than the first threshold, the weight calculation is skipped directly and an emergency broadcast is triggered.
[0059] As a further implementation, while executing steps S1-S7, the power management unit dynamically switches between three power consumption states—"standby," "identification," and "network update"—based on the system load to extend battery life.
[0060] By adopting the above technical solution, the beneficial effects of the present invention are as follows:
[0061] 1. In this invention, the edge computing unit processes visual and linguistic features simultaneously based on a multimodal vision-language model, effectively identifying obstacles, traffic signals, signs, and scene semantics. This provides a more comprehensive perception dimension and rich environmental semantic descriptions, enabling accurate identification of multiple elements such as obstacles, signs, and text in complex environments, thus improving the safety of navigation guidance for visually impaired individuals. Furthermore, it can generate voice prompts that conform to human language habits from the environmental description information, which are then output by the bone conduction audio unit for the wearer. This natural interaction avoids the information fatigue caused by noisy, mechanical alarm prompts, effectively improving the wearer's comfort and ensuring that the wearer can receive timely environmental information over a long period, further guaranteeing the travel safety of visually impaired individuals.
[0062] 2. The smart glasses of the present invention are equipped with a 60GHz millimeter-wave radar unit and a vibration motor unit, which can alert the wearer through high-frequency vibration when an object is rapidly approaching, thereby improving the ability to avoid danger in emergency situations.
[0063] 3. In the model application stage, this invention uses model compression and hardware acceleration to achieve an average end-to-end latency of 380ms (95th < 520ms), which provides strong real-time performance and meets the requirements for safe obstacle avoidance.
[0064] 4. This invention utilizes online learning and multi-device collaboration mechanisms to dynamically optimize parameters based on user environment and habits, making it highly adaptable and widely applicable.
[0065] 5. This invention balances integration and comfort. The smart glasses weigh no more than 48g, have a battery life of no less than 16 hours, and are waterproof and dustproof, making them suitable for long-term daily wear.
[0066] 6. This invention, through the integration of hardware and software design and multimodal deep learning technology, provides smart glasses that can significantly improve the real-time, comprehensive, and accurate perception of complex environments for visually impaired individuals. This, in turn, helps improve their ability to travel safely and live independently, thus improving their quality of life. It has broad prospects for industrial applications and, in turn, promotes the development and construction of a service-friendly society in my country, thus having profound social significance. Attached Figure Description
[0067] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0068] Figure 1 This is a schematic diagram of the smart glasses structure in one or more embodiments of the present invention;
[0069] Figure 2 These are three views of the smart glasses in one or more embodiments of the present invention;
[0070] Figure 3 This is a hardware module block diagram of smart glasses in one or more embodiments of the present invention;
[0071] Figure 4 This is a schematic diagram of the connection circuit between the camera unit and the edge computing unit in one or more embodiments of the present invention;
[0072] Figure 5 This is a block diagram of an edge computing unit in one or more embodiments of the present invention;
[0073] Figure 6 This is a flowchart of an environment perception method based on a multimodal vision-language model in one or more embodiments of the present invention;
[0074] Figure 7 This is a sub-flowchart of step S5 in the environment perception method based on a multimodal vision-language model in one or more embodiments of the present invention;
[0075] Figure 8 This is a flowchart of the adaptive learning mechanism in one or more embodiments of the present invention;
[0076] Figure 9This is a schematic diagram of current consumption in one or more embodiments of the present invention;
[0077] Figure 10 This is a schematic diagram of multi-device collaborative application in one or more embodiments of the present invention;
[0078] Figure 11 Radar chart of experimental test data in one or more embodiments of the present invention;
[0079] Figure 12 This is a bar chart of experimental test data from one or more embodiments of the present invention.
[0080] In the diagram: 1. Main body of glasses; 2. Camera unit; 3. Edge computing unit; 4. Bone conduction audio unit; 5. Touch interaction unit; 6. Power management unit; 7. Wireless communication unit; 8. Infrared camera unit; 9. Vibration motor unit; 10. 60GHz millimeter-wave radar unit; 11. Ultrasonic sensor unit; 12. ToF sensor unit. Detailed Implementation
[0081] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0082] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0083] Example 1
[0084] In one typical embodiment of this application, smart glasses based on a multimodal vision-language model are provided, with reference to... Figures 1-3 The device includes a glasses body 1, a camera unit 2, an edge computing unit 3, a bone conduction audio unit 4, a touch interaction unit 5, a power management unit 6, and a wireless communication unit 7. The camera unit 2, the edge computing unit 3, the bone conduction audio unit 4, the touch interaction unit 5, the power management unit 6, and the wireless communication unit 7 are all located on the glasses body 1.
[0085] Camera unit 2, disposed on the main body of the glasses 1, is used to acquire real-time images of the environment in front of the wearer. Specifically, in conjunction with... Figure 1 As shown, in this embodiment, camera unit 2 is installed on the front frame of the glasses body 1. To ensure the image acquisition range and clarity, camera unit 2 adopts a wide-angle dual fisheye camera with a horizontal field of view of not less than 120° and a resolution of not less than 1600×1200@30fps. In a preferred embodiment of this application, the camera unit has a built-in OV2640 CMOS sensor with a horizontal field of view of 140° and a resolution of 1600×1200@30fps. The acquired environmental images can be sent to the edge computing unit via the MIPI-CSI2 bus.
[0086] Edge computing unit 3 is electrically connected to camera unit 2. Edge computing unit 3 incorporates a lightweight multimodal vision-language model for preprocessing and semantic analysis of acquired environmental images and generating environmental description information. Specifically, refer to... Figure 4 The edge computing unit utilizes the XIAO-ESP32-S3 development board, whose ESP32-S3 chip employs a dual-core XTexa LX7 processor with a clock speed of 240MHz, and features 8MB of PSRAM and 8MB of Flash memory. The edge computing unit also includes an external FPGA coprocessor (LatticeiCE40UP5K) for JPEG-XS encoding / decoding and zero-copy transmission, providing hardware acceleration for JPEG-XS encoding / decoding. The FPGA communicates with the MCU via an 8-bit SPIM channel, reducing video link latency by approximately 30%. The edge computing unit receives image signals from the camera unit, executes a lightweight Qwen-VL model, and finally outputs audio signals to the bone conduction audio unit.
[0087] The bone conduction audio unit 4, electrically connected to the edge computing unit 3, is used to broadcast the environmental description information to the wearer in voice form. In this embodiment, refer to... Figure 1 and Figure 2 The bone conduction audio unit 4 uses a bone conduction speaker, which is installed on the side of the main body of the glasses 1. The bone conduction speaker consists of two medical-grade vibrating pads, a power amplifier TPA2016 and a digital filter, with a bandwidth of 250Hz-8kHz and a maximum sound pressure level of 115dBSPL. It can output voice descriptions of the environment and navigation commands.
[0088] A touch interaction unit 5 is disposed on the glasses body 1 and is used to receive touch or gesture commands from the wearer. In this embodiment, refer to... Figure 1The touch interaction unit 5 uses a capacitive touch strip (35mm in length) and is installed on the main body of the glasses 1. It can recognize four types of gestures: click, double click, swipe, and long press. It is used to realize mode switching, voice repeat, speech speed adjustment and emergency help. Its detection interrupt is connected to GPIO19 of ESP32-S3.
[0089] Power Management Unit 6 is used to power various functional units and perform dynamic power consumption optimization. Specifically, the power management unit uses a 3.7V / 650mAh polymer lithium battery, in conjunction with an AEM10941 buck-boost PMIC, to achieve dynamic power consumption grading: 28mW in standby, 520mW for identification, and 790mW for OTA updates; battery life ≥16 hours. (Refer to...) Figure 1 A power button is provided on the main body of the glasses 1, which makes it easy for the wearer to operate and control the power management unit.
[0090] The wireless communication unit 7 is used to enable bidirectional data interaction and model updates between the edge computing unit and an external mobile terminal or the cloud. Specifically, refer to... Figure 2 The wireless communication unit 7 is installed on the front side of the glasses body 1. The wireless communication unit supports 802.11b / g / n and Bluetooth 5.0 (Mesh) and has a reserved U.FL antenna socket for easy connection of external UWB module. It also supports 2.4GHz Wi-Fi and Bluetooth 5.0 for model / firmware updates and mobile terminal pairing.
[0091] The 60GHz millimeter-wave radar unit 10 is used to detect rapidly approaching objects. Specifically, the 60GHz millimeter-wave radar has a update rate of 50Hz and a latency of <5ms, and can detect rapidly approaching vehicles or pedestrians to respond to emergencies.
[0092] A vibration motor unit 9, disposed on the main body 1 of the glasses, is electrically connected to the 60GHz millimeter-wave radar unit 10 and is used to provide high-frequency vibration to rapidly approaching objects. Specifically, the vibration motor unit provides high-frequency vibration when an object approaches rapidly, and the vibration intensity varies with the distance of the approaching object, allowing the wearer to intuitively perceive the distance of moving objects around them, facilitating avoidance and improving safety. Meanwhile, referring to... Figure 1 Vibration motors are installed on the left and right sides of the main body of the glasses 1. The position of the moving object can be determined according to the different vibration states on the left and right: vibration on the left side indicates that the moving object is closer to the left, and vice versa.
[0093] The ultrasonic sensor unit 11 is electrically connected to the edge computing unit 3 and is used to detect transparent obstacles such as glass doors, covering the blind spots of the camera.
[0094] The ToF sensor unit 12, electrically connected to the edge computing unit 3, is used to detect ground elevation differences. Specifically, the ToF sensor can detect ground elevation differences of 1-4 meters in front, further enriching the semantic description of the environment and providing the wearer with more comprehensive environmental information to improve safety during travel.
[0095] Infrared camera unit 8 is powered by power management unit 6 and electrically connected to camera unit 2. Infrared camera unit 8 and camera unit 2 adopt a hybrid architecture. Under default low light conditions, the two cameras work together to fuse key data. In extreme darkness, they automatically switch to pure infrared acquisition mode. Through dynamic power consumption optimization and multimodal feedback, a balance between safety and energy efficiency is achieved.
[0096] Reference Figure 3 Each functional unit is electrically integrated via a 4-layer HDIPCB, with power layer and signal layer separated to reduce electromagnetic interference.
[0097] Reference Figure 5 In this embodiment, the functional modules of the edge computing unit specifically include:
[0098] The preprocessing module is used to perform denoising, enhancement, and cropping on the environmental image.
[0099] The lightweight multimodal vision-language model module uses quantization and knowledge distillation techniques to compress parameter size.
[0100] The scene analysis module is used to output structured scene information, including object category, location, and hazard level information.
[0101] The priority judgment module is used to calculate weights and filter high-priority events based on the risk index R, distance D, and user attention factor U.
[0102] The semantic generation module is used to convert the structured scene information of the high-priority events into semantic description text.
[0103] A speech synthesis module is used to synthesize the semantic description text into a speech stream.
[0104] Example 2
[0105] In another typical embodiment of this application, an environment perception method based on a multimodal vision-language model is provided. This method is executed on smart glasses based on the multimodal vision-language model of Embodiment 1, referring to... Figure 5 It includes the following steps:
[0106] S1. Real-time acquisition of environmental images using a camera unit.
[0107] Specifically, the camera unit captures the field of view at 30fps and transmits it to the edge computing unit; in low light, it uses a hardware HDR mode (such as the 12-bit double exposure output of the OV9282); and it uses a double fisheye lens with a diagonal of 140°.
[0108] S2. Preprocessing: Denoising, enhancing, and cropping the environmental image.
[0109] Specifically, the preprocessing module performs CLAHE (clipLimit=0.03, tileGridSize=8×8) and adaptive hybrid noise reduction (median kernel 5×5 + Gaussian σ=0.8) with a processing delay of 33ms. It performs noise reduction, histogram equalization, and image enhancement on the acquired environmental images to improve the quality of low-light and high-contrast scenes.
[0110] S3. Input the preprocessed image into the lightweight multimodal vision-language model module for cross-modal reasoning to obtain the scene semantic vector.
[0111] Furthermore, in step S3, the multimodal vision-language model (Qwen-VL model) is cropped to remove irrelevant modal branches. The processed image is then input into the multimodal vision-language model to output a multidimensional scene semantic vector, with spatial attention weights added. The ToF sensor unit is used to identify elevation differences and mark key regions.
[0112] Furthermore, in step S3, the lightweight multimodal vision-language model is compressed to 350 million parameters through knowledge distillation and hybrid precision quantization (INT8 / FP16); the inference latency of TensorRT8.6 is ≤120ms, improving the real-time performance of feedback.
[0113] S4. Based on the scene semantic vector, identify and output the target category, spatial orientation and hazard level.
[0114] Furthermore, in step S4, target category identification includes:
[0115] Use a camera to capture images of the environment. Run a lightweight AI model (such as TensorFlow LiteMicro) to identify objects and output category labels (such as "pedestrian", "vehicle", "steps", "text" etc.).
[0116] OCR processing (text targets only): If the "text" category is detected, OCR (such as a simplified version of Tesseract) is invoked to extract the text content and append it to the output.
[0117] Distance recognition includes:
[0118] The ultrasonic sensor (HC-SR04) emits sound waves and calculates the echo time, which is then converted into distance. Laser ranging is performed using a ToF sensor (VL53L0X), which offers higher accuracy and is suitable for short-distance detection.
[0119] Calculation logic: The sensor returns the raw distance value, which is directly used for subsequent hazard level determination.
[0120] Orientation recognition includes:
[0121] Using camera units and image analysis: Determine the orientation of an object based on its position (left / center / right) in the frame.
[0122] Combining a multi-directional ultrasonic array: Multiple sensors detect different directions respectively, and the distance values are compared to determine the location (e.g., the object is closer on the left → the object is on the left).
[0123] Output:
[0124] Possible values: "front", "left", "right", "front left", "front right" (currently, it recognizes 180 degrees in front).
[0125] Hazard level identification includes:
[0126] The criteria for determining the hazard level are based on a dynamic calculation of category and distance, specifically:
[0127] If the category is a vehicle, the danger threshold will be set within the range of <200cm, and the danger level is the highest (Level 3); if the category is a step, the danger threshold will be set within the range of <50cm, and the danger level is also the highest (Level 3); if the category is a pedestrian, the danger threshold will be set within the range of <100cm, and the danger level is moderate (Level 2); if the category is a static obstacle, the danger threshold will be set within the range of <30cm, and the danger level is the lowest (Level 1); all others are safe (Level 0).
[0128] S5. Hazard level processing: Select n high-priority information items according to the weight formula P(i)=w1R+w2D+w3U.
[0129] Furthermore, it receives input from a 60GHz millimeter-wave radar unit, urgently detects rapidly approaching objects, determines the risk index R, and then immediately triggers vibration and voice broadcast; ordinary obstacles are triggered by distance D, and in static environments, map data is saved to the cloud platform.
[0130] Further, the specific steps are as follows:
[0131] S51. Input the parsed scene data, including object categories, spatial orientation, and hazard level.
[0132] S52. Extract parameters such as risk index R, distance D, and user attention factor U from the scene data.
[0133] Furthermore, when the risk index R is higher than the first threshold, i.e. R≥Re, the weight calculation is skipped and an emergency broadcast is triggered, directly outputting the "Stop immediately" command; otherwise, the normal sorting is entered and the next step is executed.
[0134] S53. Calculate the priority score P, using the weighting formula:
[0135] P(i) = w1R + w2D + w3U.
[0136] Where R is the risk index, ranging from 0 to 3, which can be directly substituted into the calculation; D is the distance, with higher weight for closer objects, and needs to be normalized, such as 1 - (distance / maximum detection distance); U is the user's attention factor or urgency level, which can be set manually or calculated based on the speed of movement (e.g., increasing the U value when an object approaches quickly).
[0137] In a specific embodiment, the weighting coefficients are w1 = 0.5 (most important for danger level), w2 = 0.3 (second most important for distance), and w3 = 0.2 (second most important for urgency level).
[0138] S54. Compare the priority score P with the set threshold T.
[0139] If P≥T, the data in this scenario has high priority; the data information is retained and sorted.
[0140] If P < T, the data for this scenario has low priority and is filtered.
[0141] In one specific embodiment, T = 0.6.
[0142] Furthermore, under outdoor street scene conditions of 20℃ and 300lx, the end-to-end average response time is 380ms, and the 95th percentile is 520ms, which can meet the real-time requirements for visually impaired travel.
[0143] S6. Convert the priority information into semantic description text.
[0144] Furthermore, based on a template-fine-tuning hybrid strategy, guidance texts conforming to spoken language habits are generated in the form of natural language short sentences with an average length of 14±3 Chinese characters.
[0145] S7. The semantic description text is converted into a speech stream and played through a bone conduction audio unit.
[0146] Furthermore, by utilizing the neural vocoder WaveNet / Tacotron2, a 32kHz PCM stream can be generated within 200ms, meeting real-time requirements.
[0147] Example 3
[0148] In another embodiment of the present invention, after step S7 in embodiment two is completed, the following steps are further performed:
[0149] S8: Collect user feedback through touch interaction unit or voice commands.
[0150] S9. When the wireless communication unit is connected to the network, the user feedback is uploaded to the cloud for incremental fine-tuning.
[0151] S10. The updated model is then distributed to the edge computing unit via OTA to achieve online adaptive learning.
[0152] Furthermore, refer to Figure 8 After the voice prompt, the system waits for user feedback within 4 seconds: a positive response (double-tap touch) or a negative response (long press). Feedback is cached in a CircularBuffer (256 entries). When the glasses are charging and connected to Wi-Fi, the feedback is automatically uploaded to the cloud for incremental fine-tuning during model training. The updated weights are then distributed via OTA. Real-world testing shows that after a 7-day learning cycle, the system can improve the accuracy of commonly used scenarios by an average of 9.1%.
[0153] Example 4
[0154] In another embodiment of the present invention, while performing steps S1-S7 in Embodiment 2, the power management unit can dynamically switch between three power consumption states—"standby," "identification," and "network update"—according to the system load, thereby extending battery life. Specifically, refer to... Figure 9 The display shows the current curves of the camera unit, edge computing unit, and wireless communication unit under three power consumption states: "standby-recognition-network update". When the FPGA is off, the recognition power consumption is 520mW; when the FPGA is on, the total power consumption drops to 370mW due to the acceleration of encoding and decoding. Based on this, dynamic power consumption optimization can be performed to extend the battery life of the entire smart glasses system.
[0155] Additionally, refer to Figure 10 In this embodiment, the smart glasses can form a Bluetooth mesh network (BLEMesh) with Internet of Things (IoT) devices such as smart bracelets and guide canes, and connect to a mobile terminal. When the accelerometer detects a fall, or the GPS module identifies a complex intersection, the system automatically pushes an SOS message to the guardian's mobile terminal app, enabling multi-device collaborative applications.
[0156] Example 5
[0157] In this embodiment, three typical scenarios were selected for experimental testing: outdoor street scene, indoor shopping mall, and low-light nighttime scene. Figure 11 and Figure 12 The test data is as follows:
[0158] The smart glasses provided in this application achieve an average recognition accuracy of 92.3% across multiple scenarios; the accuracy rate remains at 88.7% even in low-light conditions (<50lx); user collision incidents are reduced by 71%, and independent travel distance is increased by 63%. It is evident that, compared to existing smart glasses such as OrCam MyEye and Envision Glasses, the smart glasses in this application significantly improve recognition accuracy in complex scenarios, greatly enhancing the travel distance and safety of visually impaired individuals, and contributing to improving the quality of life for this special group.
[0159] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Those skilled in the art should understand that the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A smart glasses based on a multimodal vision-language model, characterized in that, include: The main body of the glasses; A camera unit is disposed on the main body of the glasses and is used to capture images of the environment in front of the wearer in real time; An edge computing unit is electrically connected to the camera unit. The edge computing unit has a built-in lightweight multimodal vision-language model, which is used to preprocess the acquired environmental images, perform semantic analysis, and generate environmental description information. A bone conduction audio unit, electrically connected to the edge computing unit, is used to broadcast the environmental description information to the wearer in the form of voice. A touch interaction unit is disposed on the main body of the glasses and is used to receive touch or gesture commands from the wearer; The power management unit is used to power each functional unit and perform dynamic power consumption optimization. The wireless communication unit is used to enable bidirectional data interaction and model updates between the edge computing unit and an external mobile terminal or cloud.
2. The smart glasses based on a multimodal vision-language model as described in claim 1, characterized in that, The edge computing unit includes: The preprocessing module is used to perform denoising, enhancement, and cropping on the environmental image; A lightweight multimodal vision-language model module is developed, employing quantization and knowledge distillation techniques to compress parameter size. The scene analysis module is used to output structured scene information, including object category, location, and hazard level information. The priority determination module is used to calculate weights and filter high-priority events based on the risk index R, distance D, and user attention factor U; the weight calculation formula is as follows: P(i) = w1R + w2D + w3U; The semantic generation module is used to convert the structured scene information of the high-priority events into semantic description text; A speech synthesis module is used to synthesize the semantic description text into a speech stream.
3. The smart glasses based on a multimodal vision-language model as described in claim 1, characterized in that, The horizontal field of view of the camera unit is not less than 120°, and the resolution is not less than 1600×1200@30fps.
4. The smart glasses based on a multimodal vision-language model as described in claim 1, characterized in that, The touch interaction unit supports click, double click, swipe and long press gestures, which can be used to switch modes, repeat voice, adjust speech speed and request emergency help.
5. The smart glasses based on a multimodal vision-language model as described in claim 1, characterized in that, The edge computing unit is also connected to an FPGA coprocessor for hardware acceleration of JPEG-XS encoding and decoding.
6. The smart glasses based on a multimodal vision-language model as described in claim 1, characterized in that, Also includes: A 60GHz millimeter-wave radar unit is used to detect rapidly approaching objects; A vibration motor unit is disposed in the main body of the glasses and electrically connected to the 60GHz millimeter-wave radar unit, used to vibrate rapidly approaching objects at high frequency.
7. The smart glasses based on a multimodal vision-language model as described in claim 1, characterized in that, Also includes: An ultrasonic sensor unit, electrically connected to an edge computing unit, is used to detect transparent obstacles and cover the camera's blind spot; The ToF sensor unit is electrically connected to the edge computing unit and is used to detect the difference in ground elevation.
8. An environment perception method based on a multimodal vision-language model, characterized in that, Performed using smart glasses based on a multimodal vision-language model as described in any one of claims 1-7, the process includes the following steps: S1. Real-time acquisition of environmental images using a camera unit; S2. Preprocessing: Denoising, enhancing, and cropping the environmental image; S3. Based on a lightweight multimodal vision-language model, perform cross-modal reasoning on the preprocessed image to obtain scene semantic vectors; S4. Identify object categories, spatial locations, and hazard levels based on the scene semantic vectors; S5. Hazard level processing: Select n high-priority information items according to the weight formula P(i)=w1R+w2D+w3U; S6. Convert the priority information into semantic description text; S7. The semantic description text is converted into a speech stream and played through a bone conduction audio unit.
9. The environment perception method based on a multimodal vision-language model as described in claim 8, characterized in that, Step S7 is followed by: S8. Collect user feedback through touch interaction unit or voice command; S9. When the wireless communication unit is connected to the network, the user feedback is uploaded to the cloud for incremental fine-tuning. S10. The updated model is then distributed to the edge computing unit via OTA to achieve online adaptive learning.
10. The environment perception method based on a multimodal vision-language model as described in claim 8, characterized in that, In step S3, the multimodal vision-language model is cropped to remove irrelevant modal branches. The processed image is then input into the multimodal vision-language model to output a multidimensional semantic vector, with spatial attention weights added. The ToF sensor unit is used to identify elevation differences and mark key regions.
11. The environment perception method based on a multimodal vision-language model as described in claim 8, characterized in that, In step S4, the criteria for determining the hazard level are defined based on category and distance, specifically from high to low as follows: If the category is vehicles and the distance is less than 200cm, or if the category is steps and the distance is less than 50cm, the danger level is level 3; If the category is pedestrian and the distance is less than 100cm, the danger level is level 2; If the obstacle is classified as static and the distance is less than 30cm, the hazard level is Level 1. In other cases, the danger level is 0, meaning it is safe.
12. The environment perception method based on a multimodal vision-language model as described in claim 8, characterized in that, Step S5 is as follows: Input the parsed scene data, including object categories, spatial orientation, and hazard level; Extract parameters such as risk index R, distance D, and user attention factor U from the scene data; The priority score P is calculated using the following weighting formula: P(i) = w1R + w2D + w3U; Compare the priority score P with the set threshold T: If P≥T, the data in this scenario has high priority; the data information is retained and sorted. If P < T, the data for this scenario has low priority and is filtered.
13. The environment perception method based on a multimodal vision-language model as described in claim 8, characterized in that, In step S5, when the priority judgment module detects that the risk index R is higher than the first threshold, it directly skips the weight calculation and triggers an emergency broadcast.
14. The environment perception method based on a multimodal vision-language model as described in claim 8, characterized in that, While executing steps S1-S7, the power management unit dynamically switches between three power consumption states: "standby - identification - network update" based on the system load.