Intelligent blind guiding cap system based on multi-mode perception and operation method thereof

By using multimodal perception fusion technology, the problems of perception blind spots and rigid prompts in existing guide systems have been solved. This enables high-precision obstacle recognition and personalized prompts in complex environments, adapting to users with different sensory abilities and improving the safety and coordination of guide systems.

CN121818232APending Publication Date: 2026-04-10陈子怡
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing guide systems for the blind have a single perception dimension and lack multi-dimensional data fusion. They are prone to perception blind spots in complex scenarios, fail to fully cover the differentiated needs of users with different sensory abilities, and their prompts are easily affected by external interference. Their interconnection and collaboration capabilities are limited, and their path re-identification accuracy is insufficient in low-light or occluded scenarios, making it difficult to meet the diverse needs of blind people in complex travel scenarios.

Method used

Employing multimodal perception fusion technology, environmental information is collected through a binocular TOF camera, a 4-microphone pickup array, and temperature and humidity sensors. Combined with image super-resolution technology and audio noise reduction algorithms, data processing is performed to generate comprehensive decisions. Bone conduction hearing, micro-visual display, and tactile vibration are used for prompts, enabling data sharing and emotion perception optimization between the device and associated terminals.

Benefits of technology

It improves the accuracy of obstacle recognition, enhances the safety of the guide process, adapts to users with different sensory abilities, increases the reception rate of prompts, realizes cross-scenario path recognition and personalized experience, breaks down data barriers between devices and users, and enhances the collaboration of guides.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121818232A_ABST
    Figure CN121818232A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent blind guiding, and discloses an intelligent blind guiding cap system based on multi-modal perception, which comprises a multi-source data acquisition module, a data processing module, a multi-terminal interaction and prompt module and a collaboration and optimization module. Through multi-modal perception fusion, visual information of the obstacle is captured, sound source and temperature and humidity data are also included, and correlation reasoning is carried out, so that the problem of a single-modal perception blind area is solved, the accuracy of obstacle recognition in complex environments such as rainy days is improved, and the safety of the blind guiding process is greatly enhanced; meanwhile, through layered adaptive interaction and integration of bone conduction hearing, micro-vision display, tactile vibration and other multi-sensory prompt modes, users with residual vision and mild hearing impairment are covered, the prompt information receiving rate is increased, and adaptive coverage of different user groups is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent guide blind technology, and more particularly discloses an intelligent guide blind hat system based on multi-modal perception and an operating method thereof. BACKGROUND

[0002] Blind people refer to a group who cannot or cannot completely perceive environmental information through vision due to visual impairment or loss. Blind people will be seriously affected in their independent mobility, environmental perception and personal safety in complex environments (such as streets and public places) due to visual loss. Therefore, a guide blind system is needed to guide blind people.

[0003] The prior art patent document with the authorization announcement number CN116549267A discloses a "distributed voice guide blind system and method", which includes a guide blind host, a guide blind hat and a guide blind stick. The guide blind host is provided with an embedded main controller, a Beidou module, a 5G module, a voice recognition and voice synthesis broadcast module and a wireless communication module one. The guide blind hat is provided with a single-chip microcomputer one, a front intelligent camera, a rear intelligent camera, a microphone, a headset, a sound and light alarm and a wireless communication module two. The guide blind stick is provided with a single-chip microcomputer two, an intelligent camera, an obstacle avoidance sensor array and a wireless communication module three.

[0004] The prior art patent document with the authorization announcement number CN110368272A discloses an "intelligent navigation device and navigation method", which includes a processor and a positioning module, an image acquisition module, a ranging module, a voice system, a battery and a memory electrically connected with the processor. The positioning module is used for positioning. The image acquisition module is used for acquiring road video and image information. The image acquisition module includes a camera, a lens and an illumination light source.

[0005] Although the prior art can acquire road environment information, realize positioning navigation and voice prompt and other functions, assist blind people to avoid obstacles, independently go to the destination and reduce the dependence on family companionship, it does not need to use multiple devices separately, and to a certain extent, reduces the guide blind cost. However, the perception dimension of the prior art is relatively single, and it mainly depends on visual and voice single mode or limited combination, lacks fusion analysis of multi-dimensional data such as environmental temperature and humidity, and is easy to have a perception blind area in a complex scene. The differentiated needs of users with different sensory abilities such as residual vision and mild hearing impairment are not fully covered. The prompt information is easy to be disturbed by the outside world. There is a lack of personalized prompt adjustment mechanism based on user emotional state, the use experience is relatively rigid, the interconnection and cooperation ability is limited, there is no community cooperation function such as sharing of dangerous road sections, and the path re-identification mainly depends on single visual feature, which has insufficient accuracy in scenes such as dim light and shielding, and it is difficult to meet the diversified needs of blind people in complex travel scenes. SUMMARY

[0006] The application mainly provides an intelligent guide blind cap system based on multi-modal perception and an operation method thereof, which can solve the problems in the background art.

[0007] To solve the above technical problems, the application provides the following technical solutions, more specifically an operation method of an intelligent guide blind cap based on multi-modal perception, comprising: S1, acquiring peripheral environment related information, object related information and sound related information to ensure that the collected information can cover various key situations around the user; S2, sorting, screening, strengthening processing and integrating analysis of the collected information, converting the dispersed information into unified decision basis with correlation and usability; S3, delivering corresponding prompt information to the user according to the integrated decision basis to ensure that the prompt content can be accurately perceived by the user; S4, realizing information transmission and interactive sharing between the device and the associated terminal, recording relevant data in the use process, and dynamically adjusting and optimizing the perception, processing and prompt process according to the data.

[0008] Further, in S1, the distance, volume and motion trajectory information of the object are collected by a binocular TOF camera, the sound source position, type and motion speed information are collected by a 4-microphone pickup array, and the temperature and humidity environment data are collected by a temperature and humidity sensor.

[0009] Further, in S2, the visual, auditory and environmental data collected are denoised and standardized in format, and are simultaneously strengthened by image super-resolution technology and audio noise reduction algorithm, and the path and obstacle are recognized by a re-identification model fusing visual, auditory and environmental features, to generate a comprehensive decision including obstacle information, optimal avoidance direction and re-identification result.

[0010] Further, in S3, the voice prompt synthesized in real time by TTS is played by bone conduction, the voice speed corresponds to the danger level, the red / green indicator light and simple text prompt are displayed on the inner side of the hat brim, and the vibration motor in the corresponding direction vibrates with different intensity and frequency, the vibration direction is bound to the avoidance direction, and the intensity and frequency are linked to the danger level.

[0011] Further, in S4, the MQTT protocol is used to realize data transmission between the device and the associated terminal, to support bidirectional instruction interaction and danger section marking sharing between the device and the associated terminal, to store trajectory data, multi-modal data and user reaction related information, to capture user micro-expression by a camera and to collect tone by a pickup array, to identify user emotional state, and to adjust the prompt strategy based on reinforcement learning algorithm.

[0012] According to another aspect of the present invention, a smart guide hat system based on multimodal perception is provided. This system is implemented based on the above-mentioned smart guide hat operation method based on multimodal perception, specifically including: a multi-source data acquisition module, a data processing module, a multi-terminal interaction and prompting module, and a collaboration and optimization module. The multi-source data acquisition module acquires multi-dimensional data of the user's surrounding environment, related objects, and sounds. The data processing module processes and analyzes the acquired multi-dimensional data to form effective information that can support decision-making. The multi-terminal interaction and prompting module delivers prompts to the user. The collaboration and optimization module realizes information transmission and interactive sharing between the device and associated terminals, and optimizes prompts based on relevant data during use.

[0013] Furthermore, the multi-source data acquisition module includes: a binocular camera module, a sound source acquisition module, and an environmental perception module; Binocular camera module: Employs a binocular TOF camera to acquire information on the distance, volume, and motion trajectory of objects at 30 frames per second; Sound source acquisition module: It consists of a 4-microphone pickup array and uses beamforming technology to acquire information on the location, type and speed of sound sources, and filter out ambient noise; Environmental sensing module: Employs temperature and humidity sensors to collect real-time temperature and humidity data of the surrounding environment.

[0014] Furthermore, the data processing module includes: a preprocessing module, a data enhancement module, and a data fusion module; Preprocessing module: performs noise reduction and format standardization on the acquired visual, auditory, and environmental data; Data augmentation module: Employs image super-resolution technology and audio noise reduction algorithms to enhance the preprocessed data; Data fusion module: Based on the processed data, the re-identification model integrates visual, auditory and environmental features to identify paths and obstacles, and generates a comprehensive decision that includes obstacle information, optimal avoidance direction and re-identification results.

[0015] Furthermore, the multi-terminal interaction and prompting module includes: a bone conduction auditory module, a visual interaction module, and a vibration prompting module; Bone conduction hearing module: Uses bone conduction technology to play real-time synthesized TTS voice prompts, with the speech rate corresponding to the danger level; Visual interaction module: A 0.8-inch OLED micro-screen is set on the inside of the brim to display red / green indicator lights and simple text prompts; Vibration alert module: Six miniature linear vibration motors are distributed on the inside of the brim. The vibration direction is linked to the avoidance direction, and the vibration intensity and frequency are linked to the hazard level.

[0016] Furthermore, the collaboration and optimization module includes: a data communication module, an interaction and sharing module, a data storage module, and a perception optimization module; Data communication module: Receives data output from the multi-terminal interaction and prompting module, and uses the MQTT protocol to realize data transmission between the device and associated terminals; Interaction and Sharing Module: Supports two-way command interaction between devices and associated terminals, allowing users to mark dangerous road sections and share them within the area; Data storage module: Stores trajectory data, multimodal data, and related information about user responses; Perception optimization module: It captures the user's micro-expressions through the camera, combines the tone of voice collected by the microphone array to identify the user's emotional state, and uses the data in the data storage module to adjust the prompting strategy through reinforcement learning algorithm.

[0017] The beneficial effects of this invention, a smart guide hat system for the blind based on multimodal perception and its operation method, are as follows: Through multimodal perception fusion, it not only captures visual information about obstacles but also incorporates sound source, temperature, and humidity data for correlation and reasoning, solving the blind spots of single-modal perception and improving the accuracy of obstacle recognition in complex environments such as rainy days, significantly enhancing the safety of the guide process; simultaneously, through layered adaptive interaction, it integrates multi-sensory prompts such as bone conduction hearing, micro-visual display, and tactile vibration, covering users with residual vision and mild hearing impairment, improving the reception rate of prompt information and achieving adaptive coverage for different user groups; furthermore, through cloud-edge collaborative interconnection, it builds a two-way interactive and route-sharing platform, breaking down data barriers between devices, family members, and regional users, achieving real-time data synchronization and command interaction, and improving the coordination of guides; finally, through emotion perception and re-identification technology, it associates emotional states and dynamically adjusts prompt strategies, achieving cross-scenario path recognition and optimizing personalized experiences. Attached Figure Description

[0018] The present invention will now be described in further detail with reference to the accompanying drawings and specific implementation methods.

[0019] Fig. 1 This is a schematic diagram of the system framework; Fig. 2 This is a flowchart illustrating the method. Detailed Implementation

[0020] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present application can be combined with each other.

[0021] According to one aspect of the invention, such as Figs. 1-2 As shown, a smart guide hat system based on multimodal perception and its operation method are provided, including: Step 1: Multi-source data acquisition Acquire information related to the surrounding environment, objects, and sound to ensure that the collected information covers all key situations around the user. Specifically, it uses a binocular TOF camera to collect information on the distance, volume, and motion trajectory of objects; a four-microphone pickup array to collect information on the location, type, and speed of sound sources; and a temperature and humidity sensor to collect environmental data on temperature and humidity.

[0022] First, by working together with a binocular lens (8 megapixels) and a TOF depth sensor, the system continuously captures images of the scene in front at a rate of 30 frames per second. The binocular lens captures the two-dimensional contour information of objects through the principle of parallax, while the TOF depth sensor calculates the time difference of light signal propagation by emitting infrared modulated light and receiving reflected signals to accurately obtain the actual distance between the object and the user. At the same time, by combining the inter-frame comparison analysis of continuous frame images, the system tracks the motion trajectory of the object in space. Then, by comprehensively calculating the pixel distribution density and parallax data, the system accurately calculates the size of the object, making the core information about the object both comprehensive and accurate, providing basic data for subsequent judgment of whether the object constitutes an obstacle. The pickup structure, consisting of four microphones, uses beamforming technology to capture and focus surrounding sounds in a directional manner. It can accurately locate the sound source within a ±10° range. By analyzing the frequency characteristics and amplitude variation of the sound, it can distinguish different types of sound sources such as car horns, pedestrian conversations, and moving vehicles. Based on the time difference and intensity attenuation of the same sound received by the four microphones, it can calculate the speed of the sound source. During the entire acquisition process, wind noise, environmental noise, and other irrelevant interference signals are filtered to ensure that the acquired sound information can truly reflect the actual state of the surrounding sound sources and help users detect potential sound-related risks in a timely manner. In addition, the built-in temperature and humidity sensor senses changes in the temperature and humidity of the surrounding environment through its built-in sensitive elements. It converts changes in environmental physical quantities into quantifiable electrical signals, thereby generating accurate temperature and humidity data and outputting it in real time. When the humidity value exceeds the threshold (e.g., 80%), it will determine that the current environment is rainy and synchronize this key environmental information to the subsequent data processing stage. This provides a reliable basis for dynamically adjusting obstacle recognition and allows subsequent decision analysis to adapt to different environmental conditions. Finally, through the collaborative work of a binocular TOF camera, a 4-microphone pickup array, and a high-precision temperature and humidity sensor, the system achieves simultaneous acquisition of multi-dimensional data in vision (object distance, volume, and motion trajectory), hearing (sound source location, type, and motion speed), and environment (temperature and humidity). After unified time-series calibration, the data collected by each component forms a comprehensive information set covering a 360° range around the user, including static obstacles and dynamic sound sources. This solves the problem of one-sided information in single-modal perception. At the same time, data time-series alignment ensures the correlation of information in different dimensions, providing complete and reliable raw data support for subsequent integrated analysis.

[0023] Step 2: Data Processing and Integration The collected information is sorted, filtered, enhanced, and integrated for analysis, transforming scattered information into a unified decision-making basis that is relevant and usable. Specifically, the collected visual, auditory, and environmental data are denoised and standardized in format. At the same time, image super-resolution technology and audio noise reduction algorithms are used for enhancement processing. The path and obstacles are identified by a re-identification model that integrates visual, auditory, and environmental features, and a comprehensive decision is generated that includes obstacle information, optimal avoidance direction, and re-identification results.

[0024] First, preprocessing operations were performed on the three types of raw data collected: visual, auditory, and environmental. For visual data, a combination of median filtering and Gaussian filtering was used to remove salt-and-pepper noise and Gaussian noise from the images. At the same time, different frames of images were uniformly adjusted to 1080P resolution and RGB color format to ensure that the pixel arrangement and data dimensions of each frame were completely consistent. For auditory data, an adaptive noise cancellation algorithm was used to filter out residual wind noise and background noise during the sound pickup process. The audio data was uniformly converted to PCM format with a sampling rate of 16kHz and a depth of 16 bits to standardize the audio signal. For environmental data, the units of the raw data collected by the temperature and humidity sensors were calibrated (temperature was uniformly converted to degrees Celsius and humidity to percentage). An abnormal data with instantaneous changes (such as invalid values ​​with instantaneous humidity fluctuations exceeding 20%) was removed using a sliding window algorithm. Temporal interpolation was used to fill in occasional missing data points. Finally, all three types of data met the processing standards of being free of redundancy, noise-free, and uniform in format, laying the data foundation for subsequent enhancement processing. Then, data augmentation is used to further improve data quality: For the preprocessed visual data, a lightweight ESPCN image super-resolution model is used to reconstruct the pixels of the image. By enlarging image details and enhancing the clarity of edge contours, the problem of blurred obstacle images in low-light and long-distance scenes is solved (for example, the blurred edges of steps and the contours of guardrails are accurately restored). For the auditory data, on the basis of the previous denoising, spectral subtraction is used to analyze the audio frequency domain spectrum, separating the spectral characteristics of effective sound sources (such as car horns and pedestrian warning sounds) from residual noise, further filtering environmental noise, ensuring the accuracy of sound source type identification, and enabling subsequent fusion analysis to be carried out based on higher quality and more valuable data. Furthermore, the end-to-end fusion model, integrated through the data fusion module, enables multi-feature association reasoning. Simultaneously, a re-identification model, incorporating visual, auditory, and environmental features, is embedded to achieve accurate path and obstacle recognition. First, the preprocessed and enhanced data are aligned using a unified time stamp. After being input into the fusion model, the model dynamically allocates the weights of each feature through an attention mechanism (e.g., increasing the weight of visual features in bright light, amplifying the influence of auditory features in noisy environments, and increasing the correction of recognition results by temperature and humidity data in special environments such as rainy days). The re-identification model then extracts visual features of path markers in the current scene (e.g., the outline of shop signs, streetlight layout) and auditory features of fixed noise along the road segment (e.g., the background of supermarkets). The system compares the background music, traffic light sounds at intersections, and environmental temperature and humidity patterns (such as the annual humidity range of a certain road section) with the stored historical path feature database to achieve accurate cross-scene path recognition. Based on this, the fusion model combines the distance, volume, and movement trajectory of obstacles (visual data), sound source type and movement speed (auditory data), and temperature and humidity status (environmental data) to perform multi-dimensional comprehensive reasoning, accurately determining the specific type of obstacle (such as steps, vehicles, pedestrians) and the danger level (low / medium / high). The optimal avoidance direction is calculated through path planning algorithms, and finally, decision data containing obstacle information, distance, danger level, optimal avoidance direction, and re-identification results is generated.

[0025] Step 3: Transmission of prompt information Based on the integrated decision-making criteria, corresponding prompts are delivered to users to ensure that the prompts are accurately perceived by them. Specifically, the system plays real-time synthesized TTS voice prompts via bone conduction, with the voice speed corresponding to the danger level. Red / green indicator lights and concise text prompts are displayed on a micro-screen inside the brim, and vibration prompts are provided by vibration motors in the corresponding directions at different intensities and frequencies. The vibration direction is linked to the avoidance direction, and the intensity and frequency are linked to the danger level.

[0026] Bone conduction uses a bone conduction sound-generating device that fits close to the side of the user's skull. Sound waves are transmitted directly to the auditory nerve through skull vibrations, avoiding the external auditory canal and tympanic membrane, reducing interference from external environmental noise, and ensuring that the prompts are clear and identifiable. At the same time, it is connected to a TTS real-time synthesis engine, which dynamically adjusts the speech rate according to the danger level output by the data processing module: for example, when the danger level is determined to be emergency (such as a fast-moving vehicle at close range), the engine increases the speech rate to 180 words per minute to quickly convey the core avoidance information in concise and refined sentences. When the danger level is normal (such as a stationary step at a distance), the speech rate is maintained at 120 words per minute, and the sentences are more relaxed and easy to understand. Moreover, the voice content strictly corresponds to the comprehensive decision-making result, accurately including key information such as obstacle type, distance, and avoidance direction, to achieve a flawless conversion of "decision data → voice prompts". The micro-screen display is a 0.8-inch OLED micro-screen embedded inside the brim. The screen brightness can be automatically adjusted according to the ambient light (for example, the brightness is reduced by 30% in low light to avoid glare, and the brightness is increased by 50% in strong light to ensure visibility). The screen display is deeply linked to the decision-making results: for example, the green indicator light corresponds to the low danger level or safe path prompt, and the red indicator light corresponds to the medium / high danger level. The flashing frequency of the indicator light is linked to the danger level (for example, it flashes twice per second when it is high danger, and once per second when it is medium danger). The text prompts use a large and bold font, retaining only the core information (such as "1.2-meter step to the left" and "avoid to the right"), avoiding redundant content. At the same time, the text color is consistent with the indicator light color, forming visual coordination. This allows users with residual vision to capture key guidance information by quickly scanning without focusing on complex screens, perfectly matching their visual perception ability. Finally, miniature linear vibration motors are evenly distributed on the inside of the hat brim (e.g., in the directions of front, back, left, right, left front, and right front). Each motor precisely corresponds to a spatial direction and is directly bound to the optimal avoidance direction output by the data processing stage. For example, when the decision result is "right rear avoidance", the vibration motor at the corresponding position is triggered. The vibration intensity is divided into three levels: weak, medium, and strong, which match the low, medium, and high danger levels, respectively. The motor outputs the maximum vibration amplitude at the high danger level and the minimum vibration amplitude at the low danger level. The vibration frequency is differentiated between normal and emergency scenarios. In normal scenarios, it is an intermittent vibration of 1 time / 2 seconds, while in emergency scenarios, it switches to a continuous high-frequency vibration of 2 times / second. Through the vibration of "direction + intensity + frequency", even if the user cannot rely on hearing or vision, they can accurately perceive the avoidance direction and danger level through touch, realizing the synergistic complementarity of multi-sensory prompts.

[0027] Step 4: Interaction, Storage, and Optimization It enables information transmission and interactive sharing between devices and associated terminals, while recording relevant data during use, and dynamically adjusting and optimizing the sensing, processing and prompting processes based on this data; Specifically, the device uses the MQTT protocol to transmit data between the device and associated terminals, supports bidirectional command interaction and sharing of dangerous road segment markers between the device and associated terminals, stores trajectory data, multimodal data, and related information of user reactions, captures user micro-expressions through cameras and collects tone of voice through a microphone array, identifies user emotional state, and adjusts prompting strategies based on reinforcement learning algorithms.

[0028] Firstly, a transmission link is established based on the data communication module. The lightweight MQTT protocol is used to realize real-time data interaction between the device and associated terminals (such as family members' mobile APP). The device transmits the collected multimodal raw data (images, audio, temperature and humidity), the comprehensive decision results after data processing, and the current prompt status. The family members can receive and view this information in real time, clearly understand the user's travel environment and the device's working status. It also supports two-way command interaction. The family members can input voice commands through the terminal APP (such as "turn left at the intersection ahead" or "beware of construction on the left"). The commands are encoded by the APP and sent to the device through the MQTT protocol. The device receives and decodes the commands and plays them through the bone conduction hearing device, realizing remote guidance from the family members to the user. If the user finds a dangerous road section during the trip (such as "there is a temporary guardrail at the XX intersection" or "severe water accumulation on the XX section"), they can mark the location through the simple trigger button on the hat. The marking information includes the current location, environmental data, and scene image. After being uploaded to the interaction and sharing module, the platform's backend will manually review and confirm that the dangerous road section information is synchronized to all user terminals using the same device in the area, realizing the community-based sharing of dangerous information. Then, the key data throughout the entire trip is stored in a structured manner: for example, the database is indexed by "timestamp + location coordinates" and stores the multimodal data collected at the corresponding time (image frames, audio clips, temperature and humidity values), the decision results after data processing (obstacle type, avoidance direction, danger level), user reaction data (avoidance action time, operation feedback after triggering prompts), and emotional state tags, forming a complete data link of "collection-processing-prompt-feedback" (for example, a record is "10:00:30 + intersection of XX Road and XX Street + image of 1.5 meters step to the left front + "avoid to the right rear" decision + user completes avoidance in 3 seconds + tension emotion tag"). All data is stored in an encrypted format, which not only ensures data integrity and relevance, but also provides comprehensive and traceable data source support for subsequent perception optimization. Finally, the perception optimization module completes the emotion recognition and strategy adjustment and optimization: It continuously captures the user's facial micro-expressions through the secondary camera of the binocular TOF camera, extracting key features such as eyebrow shape changes (e.g., frowning), eyelid state (e.g., squinting), and mouth corner position (e.g., drooping). Simultaneously, the sound pickup array collects the user's tone data (e.g., increased speech rate, increased volume, and hurried tone). These two types of features are input into the emotion recognition model to determine the user's current emotional state (tension, relaxation, irritability, etc.). Then, it calls upon historical data from the data storage module, combines it with the real-time emotional state, and uses a reinforcement learning algorithm to optimize the prompting strategy. The algorithm uses "avoidance success rate" and "emotional comfort" as reward functions. When a user repeatedly reacts slowly and is emotionally stressed in a certain scenario (such as encountering stairs), the algorithm adjusts the timing of prompts (e.g., triggering a prompt 2 meters in advance) and the intensity of prompts (e.g., increasing vibration intensity by one level and slowing down the speech rate to 100 words per minute). When the algorithm detects that the user prefers a certain travel route (e.g., repeatedly choosing a park side road instead of a main road), it will prioritize recommending this route in subsequent navigation and optimize the re-identification feature weights of this route, making the prompt strategy more in line with the user's usage habits and emotional tolerance, thus achieving personalized dynamic adaptation.

[0029] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention are also within the protection scope of the present invention.

Claims

1. A method for operating an intelligent guide hat based on multimodal perception, characterized in that, The method includes: S1. Acquire information related to the surrounding environment, objects, and sound to ensure that the collected information covers all key situations around the user. S2. Organize, screen, process, and integrate the collected information to transform scattered information into a unified decision-making basis that is relevant and usable. S3. Based on the integrated decision-making criteria, deliver corresponding prompts to users to ensure that the prompts are accurately perceived by users. S4. Enable information transmission and interactive sharing between the device and associated terminals, record relevant data during use, and dynamically adjust and optimize the sensing, processing, and prompting processes based on this data.

2. The method for operating a smart guide hat based on multimodal perception according to claim 1, characterized in that: In S1, the distance, volume, and motion trajectory information of the object are collected by a binocular TOF camera, the location, type, and speed of the sound source are collected by a 4-microphone pickup array, and the temperature and humidity environmental data are collected by a temperature and humidity sensor.

3. The method for operating a smart guide hat based on multimodal perception according to claim 1, characterized in that: In step S2, the collected visual, auditory, and environmental data are denoised and standardized in format. At the same time, image super-resolution technology and audio noise reduction algorithm are used for enhancement processing. The path and obstacles are identified by a re-identification model that integrates visual, auditory, and environmental features, and a comprehensive decision is generated that includes obstacle information, optimal avoidance direction, and re-identification results.

4. The method for operating a smart guide hat based on multimodal perception according to claim 1, characterized in that: In S3, TTS real-time synthesized voice prompts are played via bone conduction, with the voice speed corresponding to the danger level. Red / green indicator lights and concise text prompts are displayed on a micro-screen inside the brim. Vibration prompts are provided by vibration motors in the corresponding directions at different intensities and frequencies, with the vibration direction linked to the avoidance direction and the intensity and frequency linked to the danger level.

5. The method for operating a smart guide hat based on multimodal perception according to claim 1, characterized in that: In S4, the MQTT protocol is used to realize data transmission between the device and the associated terminal, support bidirectional command interaction and dangerous road segment marking sharing between the device and the associated terminal, store trajectory data, multimodal data, and related information of user reactions, capture user micro-expressions through camera and collect tone of voice through microphone array, identify user emotional state, and adjust prompting strategy based on reinforcement learning algorithm.

6. A smart guide hat system based on multimodal perception, characterized in that, This system is implemented based on the multimodal perception-based intelligent guide hat operation method described in any one of claims 1-5, specifically including: a multi-source data acquisition module, a data processing module, a multi-terminal interaction and prompting module, and a collaboration and optimization module; the multi-source data acquisition module acquires multi-dimensional data of the user's surrounding environment, related objects, and sounds; the data processing module processes and analyzes the acquired multi-dimensional data to form effective information that can support decision-making; the multi-terminal interaction and prompting module delivers prompts to the user; the collaboration and optimization module realizes information transmission and interactive sharing between the device and associated terminals, and optimizes prompts based on relevant data during use.

7. The intelligent guide hat system based on multimodal perception according to claim 6, characterized in that: The multi-source data acquisition module includes: a binocular camera module, a sound source acquisition module, and an environmental perception module; Binocular camera module: Employs a binocular TOF camera to acquire information on the distance, volume, and motion trajectory of objects at 30 frames per second; Sound source acquisition module: It consists of a 4-microphone pickup array and uses beamforming technology to acquire information on the location, type and speed of sound sources, and filter out ambient noise; Environmental sensing module: Employs temperature and humidity sensors to collect real-time temperature and humidity data of the surrounding environment.

8. The intelligent guide hat system based on multimodal perception according to claim 6, characterized in that: The data processing module includes: a preprocessing module, a data enhancement module, and a data fusion module; Preprocessing module: performs noise reduction and format standardization on the acquired visual, auditory, and environmental data; Data augmentation module: Employs image super-resolution technology and audio noise reduction algorithms to enhance the preprocessed data; Data fusion module: Based on the processed data, the re-identification model integrates visual, auditory and environmental features to identify paths and obstacles, and generates a comprehensive decision that includes obstacle information, optimal avoidance direction and re-identification results.

9. A smart guide hat system based on multimodal perception according to claim 6, characterized in that: The multi-terminal interaction and prompting module includes: a bone conduction auditory module, a visual interaction module, and a vibration prompting module; Bone conduction hearing module: Uses bone conduction technology to play real-time synthesized TTS voice prompts, with the speech rate corresponding to the danger level; Visual interaction module: A 0.8-inch OLED micro-screen is set on the inside of the brim to display red / green indicator lights and simple text prompts; Vibration warning module: Six miniature linear vibration motors are distributed on the inside of the brim. The vibration direction is linked to the avoidance direction, and the vibration intensity and frequency are linked to the hazard level.

10. A smart guide hat system based on multimodal perception according to claim 6, characterized in that: The collaboration and optimization module includes: a data communication module, an interaction and sharing module, a data storage module, and a perception optimization module; Data communication module: Receives data output from the multi-terminal interaction and prompting module, and uses the MQTT protocol to realize data transmission between the device and associated terminals; Interaction and Sharing Module: Supports two-way command interaction between devices and associated terminals, allowing users to mark dangerous road sections and share them within the area; Data storage module: Stores trajectory data, multimodal data, and related information about user responses; Perception optimization module: It captures the user's micro-expressions through the camera, combines the tone of voice collected by the microphone array to identify the user's emotional state, and uses the data in the data storage module to adjust the prompting strategy through reinforcement learning algorithm.

Citation Information

Patent Citations

  • Intelligent navigation equipment and navigation method

    CN110368272A

  • Distributed voice blind guiding system and method

    CN116549267A