Active visual explanation system for field travel of blind people
By integrating the positioning system, depth vision module, voice module and other technologies on glasses, the active visual interpretation system for blind people on the field tourism is realized, solving the problem that blind people cannot understand the landscape while traveling, and improving travel experience and safety.
Patent Information
- Application Number
- CN202510388851.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-27
AI Technical Summary
The existing blind assisting technologies focus on the safe travel of blind people and fail to meet their psychological needs to understand the human landscape during their travel.
Design an active visual interpretation system for field tourism for blind people, including positioning system, depth vision module, voice module, storage module and processing module set on glasses. Through the collaborative work of these modules, environmental information can be collected and analyzed in real time, and real-time explanations can be performed through voice modules.
This system not only improves the travel safety of blind people, but also provides real-time explanations of attractions along the way, improving the travel experience of blind people, so that they can have a more comprehensive understanding of the tourist attractions they are participating in.
Smart Images

Figure CN120204016A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of human-computer interaction, and particularly relates to an active vision interpretation system for blind people's on-site tourism. Background Art
[0002] With the continuous progress of science and technology, blind people can be equipped with invisible "guides" to avoid obstacles when traveling, improving travel safety. However, they cannot have an active vision interpretation to inform the blind person of the front scene in real time and communicate with the blind person. Blind people usually can only passively receive instructions.
[0003] Traditional walking assistance methods mainly include blind canes, guide dogs, and manual guidance. The blind cane is the most commonly used auxiliary tool for blind people. It senses obstacles in front by knocking on the ground, but its usage scenario is limited and it cannot provide comprehensive environmental information. Guide dogs can guide blind people to travel safely after professional training, but the training cost is high, the number is limited, and their usage is limited in some occasions (such as public transportation). Manual guidance requires blind people to seek help from others. Although it is effective, it is not convenient enough and difficult to obtain at any time, which hinders them from exploring more complex and interesting travel experiences.
[0004] The research on blind people's travel technology is developing towards diversification and intelligence, mainly in the following three directions: First, intelligent aids such as AI blind aids and guide glasses, which use technologies such as lidar and computer vision to help blind people accurately identify obstacles and improve travel safety. Second, auxiliary wearable and non-wearable devices using ultrasonic air touch technology, which provide navigation information for blind people through vibration feedback, making the devices more portable and daily. Third, research on improving the computational efficiency and accuracy of stereo matching algorithms, solving problems such as specular reflection, and improving the real-time performance and accuracy of blind people's travel.
[0005] The specific working principle of intelligent blind assistance technology includes the comprehensive application of computer vision and sensor fusion, speech recognition and interaction, positioning and navigation, and intelligent intention inference. Computer vision and sensor fusion refer to using sensors such as cameras and lidar to collect surrounding environment information, and through computer vision technology for image recognition and processing to identify key information such as obstacles and road conditions. Speech recognition and interaction refer to equipping with a speech recognition module to understand the speech commands of blind people, such as navigation requests and obstacle avoidance commands, and providing navigation information and road condition tips through speech feedback. Positioning and navigation refer to combining multi-sensor fusion technologies such as GPS and MEMS gyroscopes to achieve accurate positioning, plan the optimal travel path, and actively remind users when obstacles are detected. Intelligent intention inference refers to inferring the intention of blind people by continuously learning their language and behavior habits, and providing more personalized services, such as predicting the next action and recommending suitable routes.
[0006] Currently, many institutions and companies are actively researching blind travel technologies to improve the quality of life of the blind; developing intelligent guiding devices to provide real-time perception and obstacle avoidance guidance for the blind.
[0007] AI commentary has been successfully extended to the live broadcast of niche sports events and the field of medium and small-scale sports events, serving a more segmented audience group, indicating the wide applicability of AI sports commentary services and the infinite possibilities for future development. Since 2019, IBM's "Henry" AI commentary system has been applied to the Masters Tournament in golf. Through a large language model, it generates commentaries for golf, converts data into narrative text, and converts it into voice through Watson Text-to-Speech service. IBM has also been applied to the US Open. Its AI system analyzes the historical game data of players, establishes a statistical model to predict the game results, and at the same time provides detailed game analysis, enhancing the interactivity of the audience during the game and the depth of information acquisition.
[0008] As can be seen from the above, current blind assistance technologies only focus on the safety travel needs of the blind, while ignoring their psychological needs to understand the cultural landscapes during travel. Therefore, if there is a "commentator" to describe and explain the scenery along the way in real time, it will help improve the travel experience of the blind. Summary of the Invention
[0009] Aiming at the above problems existing in the prior art, the purpose of the present invention is to provide an active visual commentary system for blind people's on-site tourism.
[0010] The present invention provides the following technical solution: An active visual commentary system for blind people's on-site tourism, which includes a positioning system, a depth vision module, a voice module, a storage module, and a processing module arranged on glasses;
[0011] The positioning system is used to collect obstacle position and speed information;
[0012] The depth vision module is used to collect object image information and identify it through a deep neural network model to obtain the type information of the object in front;
[0013] The voice module is used to prompt the position, speed, and type of the object in front; the storage module is used to collect the travel behavior process, travel language process data, and user travel records of the blind to obtain statistical information and facts of the user's highest interest;
[0014] The processing module is used to receive the statistical information and facts of the user's highest interest and perform AI real-time picture commentary and reminder of obstacle avoidance.
[0015] Further, the positioning system includes a set of cameras disposed in front of the glasses frame, and a MEMS gyroscope, a GPS module, and an infrared ranging module disposed inside the glasses frame; the set of cameras, the infrared ranging module, and the MEMS gyroscope all perform mapping operations on the outdoor scene based on the Simultaneous Localization and Mapping method to obtain the geometric structure of the outdoor scene where the user is located.
[0016] Further, after receiving the voice signal for reporting the scenery issued by the user, the depth vision detection module collects the real-time image information in front of the acquisition device, and inputs the real-time image information into the yolov11 depth neural network model to obtain the type information of the object in front.
[0017] Further, the voice module is used to report the position status information of the object in front, the obstacle avoidance prompt information, and the AI real-time picture commentary information; the signal of the voice module is transmitted through the bone conduction structure, and a voice switch is installed on the glasses frame to start or stop the reporting.
[0018] Further, the storage module is used to collect the data of the blind person's travel behavior process and travel voice process, generate corresponding statistical information and facts, generate corresponding averages based on the statistical information and facts, and the generated statistical information and facts are coupled to the processing module when the processing module and the voice module execute.
[0019] Further, after receiving the voice signal for reporting the scenery issued by the user, the storage module transmits the statistical information and facts of the user's highest interest to the processing module. The processing module actively predicts the content that the user is interested in, generates commentary text according to the content that the user is interested in when detecting the target, and transmits it to the voice module to provide real-time scene commentary for the user.
[0020] Further, the processing module corrects the statistical information and facts of the highest user interest based on the obtained user behavior information and voice feedback information, updates them to the statistical information and facts of the corresponding user's highest interest, and stores them in the storage module, and is coupled to the processing module when the processing module and the voice module execute next time.
[0021] Further, a task model that models the mapping relationship between the journey progress operation and the user state is preset in the processing module, and different task types correspond to different task models. After receiving the voice signal of the task to be executed, the processing module determines the type of the task to be executed, and then searches for the shortest state sequence that can complete the task to be executed from the task model corresponding to the task to be executed.
[0022] By adopting the above technologies, compared with the prior art, the beneficial effects of the present invention are as follows:
[0023] Based on the system of the present invention, users can select to broadcast navigation obstacle avoidance or scenic descriptions. The real-time image information collected by the depth vision detection module is recognized through a deep neural network model, and the speed and type information of the objects in front and below are obtained in combination with the positioning system. When the user selects to broadcast navigation obstacle avoidance, after generating the audio instruction sequence corresponding to each sub-state according to the shortest state sequence, the user performs actions under the guidance of the audio instructions, and each action will receive feedback from the system, greatly improving the correctness and efficiency of the operations during the journey of tourism; when the user selects to broadcast scenic descriptions, the AI model generates an audio instruction sequence according to the recognized object information to inform the user; not only enhancing the safety of the user during the on-site tourism, but also helping the user understand the tourist attractions they are participating in, providing full-process visual supervision for the user's on-site tourism. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic diagram of the hardware structure of the system module of the present invention;
[0025] Figure 2 It is a data processing flow chart of the system of the present invention;
[0026] Figure 3 It is a schematic diagram of the three-dimensional structure of the glasses of the present invention, where 3-1, 3-2, 3-3, and 3-4 are four cameras 3 of the same model, and 4-1 and 4-2 are bone conduction structures of the same structure;
[0027] Figure 4 It is a front view structure schematic diagram of the glasses of the present invention;
[0028] Figure 5 It is a top view structure schematic diagram of the glasses of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings of the specification and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0030] On the contrary, the present invention covers any alternatives, modifications, equivalent methods and solutions made within the spirit and scope of the present invention as defined by the claims. Further, in order to enable the public to have a better understanding of the present invention, some specific details are described in detail in the following detailed description of the present invention. Those skilled in the art can fully understand the present invention without the description of these details.
[0031] Please refer to Figure 1, A blind field-trip active vision interpretation system, including a positioning system, a depth vision module, a voice module, a storage module and a processing module set on glasses; a strap is provided at the rear of the glasses for fixation and load-bearing.
[0032] The positioning system is used to collect obstacle position and speed information;
[0033] The depth vision module is used to collect object image information and identify it through a depth neural network model to obtain the type information of the object in front;
[0034] The voice module is used to prompt the position, speed and type of the object in front; the storage module is used to collect the travel behavior process, travel language process data of the blind and the user travel record to obtain the statistical information and facts of the user's highest interest;
[0035] The processing module is used to receive the statistical information and facts of the user's highest interest and perform AI real-time picture interpretation and obstacle avoidance reminder.
[0036] Specifically, the positioning system includes a set of cameras arranged in front of the glasses frame and a MEMS gyroscope, a GPS module and an infrared ranging module arranged inside the glasses frame; the set of cameras, the infrared ranging module and the MEMS gyroscope all perform mapping operations on the outdoor scene based on the simultaneous localization and mapping method to obtain the geometric structure of the outdoor scene where the user is located.
[0037] The cameras adopt monocular or binocular optical cameras, and multiple cameras are located at the upper and lower parts on both sides of the front bracket of the glasses, used to observe objects in visual blind areas such as in front and on the ground; the MEMS gyroscope, GPS module and infrared ranging module are placed inside the front of the glasses.
[0038] Through the set cameras and infrared ranging module, when the producer wears the optical camera and the infrared ranging module into the outdoor scene, based on the simultaneous localization and mapping technology, the image obtained by the optical camera and the point cloud data obtained by the infrared ranging module are matched and fused to perform mapping operations on the outdoor scene to obtain the geometric structure of the outdoor scene where they are located; when the user wears the camera into the outdoor scene, by matching with the visual feature points obtained during the mapping operation, the current position of the infrared ranging module is calculated, so as to obtain the relative position between the user and the surrounding objects and generate an audio instruction sequence.
[0039] With the set camera and MEMS gyroscope, when the producer wears glasses equipped with an optical camera and MEMS gyroscope and enters an outdoor scene, based on the simultaneous localization and mapping technology, the images obtained by the optical camera and the sensor pose and acceleration data obtained by the MEMS gyroscope are fused to perform mapping operations on the outdoor scene, obtaining the geometric structure of the outdoor scene where the user is located; when the user wears the camera and enters the outdoor scene, by matching with the visual feature points obtained during the mapping operation, the current pose of the MEMS gyroscope is calculated, thereby obtaining the relative position between the user and surrounding objects and generating an audio instruction sequence.
[0040] Specifically, the depth vision detection module is placed inside the glasses directly in front of the user. After receiving the voice signal of the user reporting the scenery, it collects the real-time image information in front of the acquisition device and inputs the real-time image information into the yolov11 deep neural network model to obtain the type information of the objects in front.
[0041] Specifically, the voice module is used to report the position status information of the objects in front, obstacle avoidance prompt information, and AI real-time picture commentary information; the signal of the voice module is transmitted through the bone conduction structures on the left and right sides of the tail of the glasses frame. A voice switch is installed at the tail on the right side of the glasses frame to start or stop the reporting.
[0042] Specifically, the storage module is used to collect the data of the means of transportation taken during the trips, the total walking length data, and the destination scene data of multiple trips of blind people with more outdoor experience from various regions. Based on the means of transportation data, the total walking length data, and the destination scene data, statistical information and facts are generated for the travel behavior process of blind people, and based on the means of transportation data, the total walking length data, and the destination scene data, averages are generated for the travel behavior process of blind people; the statistical information and facts with the highest user affirmation are selected from the statistical information and facts generated for the travel behavior process of blind people, and the statistical information and facts are coupled to the processing module when the processing module and the voice module execute.
[0043] The storage module is used to collect the languages, preferred voice packs, preferred time intervals for receiving voice information, preferred key content for broadcast, and sensitive word data that blind people from various regions prefer to use when listening to text. The preferred key content for broadcast includes aspects such as scenic spot introductions, historical and cultural explanations, and natural landscape features. Based on the preferred language, preferred voice pack, preferred time interval for receiving voice information, preferred key content for broadcast, and sensitive word data, statistical information and facts are generated for the blind's travel language process, and an average value is generated for the blind's travel voice process based on the preferred language, preferred voice pack, preferred time interval for receiving voice information, and sensitive words; the statistical information and facts with the highest user affirmation are selected from the statistical information and facts generated for the blind's travel language process, and the statistical information and facts are coupled to the processing module when the processing module and the voice module are executed.
[0044] Specifically, after the processing module receives the voice signal of the user's key broadcast of the scenery, the storage module transmits the statistical information and facts of the user's highest interest to the processing module. The statistical information and facts of the user's highest interest include the blind's travel behavior process and the blind's travel language process. The processing module actively predicts the content that the user may be interested in, generates an explanatory text according to the content that the user is interested in when detecting a key target, and transmits it to the voice module to provide real-time scene explanation for the user; the real-time scene explanation is generated using an AI model, and the AI model is configured to select statistical information and facts for the explanation from the statistical information and facts generated for the blind's travel behavior process and the blind's travel language process that are of interest to potential blind travel enthusiasts; the processing module stops generating when issuing a stop explanation instruction, and the processing module can answer questions and interact according to the user's questions during the explanation process to provide personalized explanations.
[0045] The processing module will provide a reference travel plan for the blind's travel based on the stored data of the means of transportation taken during the trip and the destination scene data.
[0046] The processing module compares the stored average value data of the total walking length with the detected data of the user's already walked length. If the already walked length exceeds 50% of the total walking length data, the public seat or rest area will be regarded as a key target and a prompt will be given when the key target is detected.
[0047] Each time the processing module travels, it obtains the user's behavior information and voice feedback information, and then corrects the statistical information and facts of the highest user interest and updates them to the statistical information and facts of the highest user interest for this user. The statistical information and facts of the highest user interest for the user will be stored in the storage module and coupled to the processing module when the processing module and the voice module are executed next time.
[0048] After receiving the voice signal of the key broadcast navigation forward issued by the user, the processing module plans the travel route according to the voice signal; transmits the speed, position and type information of the object in front to the voice module to broadcast the position status information of the object in front, and reminds of obstacle avoidance when the user approaches the object.
[0049] After receiving the voice signal of the key broadcast navigation forward issued by the user, the processing module plans the travel route according to the voice signal. For each sub-state in the shortest state sequence, the positioning system obtains the position and speed information of the object in front in real time, and then the processing module generates an audio instruction sequence for guiding the user's forward action according to the position and speed information. Among them, the audio instruction sequence is a regional guidance instruction. Each time the user performs an action under the guidance of the regional guidance instruction, the positioning system obtains the position and speed information of the object in front once; the processing module judges whether the position information is the same as the expected position after the execution of the current regional guidance instruction. If it is, the next regional guidance instruction is executed. If not, the processing module regenerates the audio instruction sequence according to the position until the user moves to the final operation area to realize the current sub-state.
[0050] Specifically, the processing module is preset with a task model that models the mapping relationship between the travel operations and the user status, and different task types correspond to different task models. After receiving the voice signal of the task to be executed, the processing module judges the type of the task to be executed, and then searches for the shortest state sequence that can complete the task to be executed from the task model corresponding to the task to be executed.
[0051] Specifically, when the journey is a city stroll, the task types include tasks such as waiting for traffic lights, going straight, turning, avoiding obstacles in front, and going up and down stairs. Among them, when the task to be executed is waiting for traffic lights, going straight, and turning, the user's current status is the traveling speed. If the current traveling speed of the user obtained by the positioning system is different from the expected value, it means that the status switching instruction has not been correctly executed. If the current traveling speed of the user obtained by the positioning system is the same as the expected value, it means that the status switching instruction has been correctly executed.
[0052] Embodiment:
[0053] An active visual commentary for blind people's on-site tourism, which is set on glasses, includes a positioning system, a depth vision detection module, a voice module, a processing module and a storage module.
[0054] Among them, the positioning system is used to collect the position and speed information of the front objects and ground objects; the depth vision detection module is used to process the video frames collected by the camera and obtain the depth vision target results through the yolov11 algorithm training; the voice module is used to describe the front scenery, prompt the obstacle road conditions and obstacle information to the blind according to the position state information of the front obstacles; the processing module is used to fuse the type information of the front objects and the position and speed information of the front objects to obtain the position state information of the front objects and coordinate the input and output signals of other modules; the storage module is used to store the user's behavior and language habits, use big data and intelligent algorithms to judge the content that the user is interested in, and update it after each trip of the user. The updated statistical information and facts of the user's highest interest will be stored in the storage module and coupled to the processing module when the processing module and the voice module are executed next time.
[0055] The laser ranging module is paired with an optical camera. Combining image processing technology, deep learning technology and related algorithms, the system can accurately determine the type, distance and relative speed of obstacles. Compared with traditional blind sticks, guide dogs and other guiding methods, it provides more accurate obstacle information. Compared with guide glasses that can only identify obstacles but cannot provide types, the practicality of the system is greatly improved. The types of obstacles include but are not limited to pedestrians, trees, trash cans, cars, electric vehicles, steps, traffic lights and zebra crossings.
[0056] Compared with the traditional blind stick that relies on feeling to determine the road conditions and obstacle information, the explanatory glasses rely on image analysis to prompt the distance of the obstacles. The voice broadcast prompt adopted by the explanatory glasses (the information of two directions, distance and obstacle type in the (front, below) direction) can provide more clear obstacle road conditions and obstacle information for the blind, reducing the burden on the blind to learn and use the invention.
[0057] Specifically, as Figure 2 shown, the present invention uses Jetson nano as the main control chip, and is peripherally equipped with a depth camera, a positioning system, a storage module and a voice input / output module.
[0058] The MEMS gyroscope sensor communicates with Jetson nano through the IIC bus to collect angle information.
[0059] The TOF laser ranging module communicates with Jetson nano through the serial port to collect distance information.
[0060] Through the fusion result of the laser ranging module and the depth vision, the position information of the obstacle is output. First, the type, position and distance of the obstacle are output to the bone conduction voice input / output module through the Baidu voice synthesis technology; secondly, this information is transmitted to the position of the bone conduction earphone, as Figure 3Give a prompt as shown.
[0061] The following describes a navigation method for the active visual interpretation of blind people's on-site tourism, including the following steps:
[0062] Step 1: Use a GPS module, an optical camera, a MEMS gyroscope, a monocular or binocular camera 3, and a TOF laser ranging module to collect obstacle position and speed information.
[0063] The present invention uses a TOF laser rangefinder to collect obstacle position and speed information. When the distance to the obstacle is far, an optical camera and a laser rangefinder above the glasses are used to fuse depth vision to perceive the environmental information, and the position, speed, and type information of the distant obstacle are obtained. When approaching the obstacle, the cameras and laser rangefinders below the glasses are used to fuse depth vision to perceive the environmental information, and the position, speed, and type information of the obstacle under the feet are obtained, effectively ensuring the travel safety of blind people.
[0064] When the blind person is not close to the steps, it can be recognized through image recognition that there are steps ahead. At this time, the blind person will be prompted that there is an upward step 5.1 meters ahead. When the blind person is about to step onto the steps, the distance detected by the laser rangefinder will decrease sharply, reminding the blind person that they are about to step onto the steps ahead.
[0065] Step 2: Use the camera in the depth vision detection module to collect the video information in front of the glasses device, and input each frame of real-time image information in the video into the yolov11 deep neural network model for recognition to obtain the type information of the obstacle.
[0066] Specifically, label the common obstacles of blind people (including cars, electric vehicles, bicycles, trees, road steps, stairs, people, etc.), collect pictures of common obstacles of blind people to make a dataset, train the dataset and generate a yolov11 deep neural network model, input the image information collected by the camera into the yolov11 deep neural network model for recognition, and mark the position coordinates and type of the recognized object to complete the depth vision target detection.
[0067] Describe the situation when the recognized object is a traffic light. By selecting the recognized traffic light, convert the picture from RGB format to HSV format. Then use the Opencv library function cv2.inRange() to set the threshold to remove the background part, and then perform median filtering. Finally, calculate the number of non-zero pixels, and take the color with the most non-zero pixels in the picture after median filtering as the traffic light color result, and output the traffic light color result as the depth vision target result.
[0068] Step 3: Use the processing module to fuse the obstacle position and speed information with the obstacle type information through a decision-level information fusion algorithm to obtain the position state of the obstacle ahead.
[0069] Step 4: Transmit the position status of the obstacle ahead to the voice module for prompting, and announce the position status information of the obstacle ahead.
[0070] The device detected a motorcycle ahead and calculated that the distance to the obstacle was 3.99 m, and the position was 16 degrees to the left of the blind person. At this time, the device will announce the type, distance, and azimuth information of the obstacle to the blind person through the voice module.
[0071] Step 5: Realize navigation by collecting latitude and longitude information through the GPS module and collecting the posture information of the blind person through the gyroscope module.
[0072] Step 6: Send the latitude and longitude information, real-time image information, and posture information to the processing module through the storage module, record the route taken by the blind person through the database, count the user's travel behavior and language behavior, and couple them to the execution module during the next travel.
[0073] The blind guidance system provided by the present invention consists of four major parts: information collection, information processing, information transmission, and information feedback. The device, in cooperation with a laser rangefinder based on the yolov11 algorithm model for image recognition, can realize voice announcement of the obstacle name and detection distance. A database is built on the server side to record the position information of the route taken by the blind person.
[0074] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An active visual interpretation system for blind people's on-site tourism, characterized in that: It includes a positioning system, a depth vision module, a voice module, a storage module and a processing module arranged on the glasses; The positioning system is used to collect obstacle position and speed information; The deep vision module is used to collect object image information and identify it through a deep neural network model to obtain type information of the object in front; The voice module is used to indicate the position, speed and type of the object ahead; the storage module is used to collect the blind person's travel behavior process, travel language process data and user travel records to obtain the statistical information and facts of the user's greatest interest; The processing module is used to receive statistical information and facts of greatest interest to the user, and to perform AI real-time screen interpretation and obstacle avoidance reminders.
2. The active visual interpretation system for on-site tourism for the blind according to claim 1, characterized in that: The positioning system comprises a group of cameras (3) arranged in front of the glasses support, and a MEMS gyroscope, a GPS module and an infrared distance measurement module arranged inside the glasses support; the group of cameras (3), the infrared distance measurement module and the MEMS gyroscope all perform a mapping operation on an outdoor scene based on a synchronous positioning and map construction method to obtain a geometric structure of the outdoor scene where the user is located.
3. The active visual interpretation system for on-site tourism for the blind according to claim 2, characterized in that: After receiving the voice signal of the user reporting the scenery, the deep vision detection module collects real-time image information in front of the device, and inputs the real-time image information into the yolov11 deep neural network model to obtain the type information of the object in front.
4. The active visual interpretation system for on-site tourism for the blind according to claim 3 is characterized in that: The voice module is used to broadcast the position status information of the object in front, obstacle avoidance prompt information and AI real-time screen explanation information; the signal of the voice module is transmitted through the bone conduction structure (4), and a voice switch (1) is installed on the glasses bracket to start or stop the broadcast.
5. The active visual interpretation system for on-site tourism for the blind according to claim 4, characterized in that: The storage module is used to collect data on the blind person's travel behavior process and travel voice process, generate corresponding statistical information and facts, and generate corresponding average values based on the statistical information and facts. The generated statistical information and facts are coupled to the processing module when the processing module and the voice module are executed.
6. The active visual interpretation system for on-site tourism for the blind according to claim 5, characterized in that: After the processing module receives the voice signal of the user reporting the scenery, the storage module transmits the statistical information and facts of the user's greatest interest to the processing module. The processing module actively predicts the content that the user is interested in, generates an explanatory text according to the content that the user is interested in when a target is detected, and transmits it to the voice module to provide real-time scene explanation for the user.
7. The active visual interpretation system for on-site tourism for the blind according to claim 6, characterized in that: The processing module corrects the statistical information and facts of the highest user interest based on the acquired user behavior information and voice feedback information, updates them to the statistical information and facts of the highest interest of the corresponding user, and stores them in the storage module, which is coupled to the processing module when the processing module and the voice module are executed next time.
8. The active visual interpretation system for on-site tourism for the blind according to claim 7, characterized in that: The processing module is preset with a task model that models the mapping relationship between travel operations and user status, and different task types correspond to different task models. After the processing module receives the voice signal of the task to be executed, it determines the type of the task to be executed, and then searches for the shortest state sequence that can complete the task to be executed from the task model corresponding to the task to be executed.
Citation Information
Cited By
A Smart Voice-Guided Interaction Method and System for Accessible Tourism
CN122573475A