Blind person full-scene perception interaction method and system based on multi-mode AI

By using multimodal AI technology to analyze multi-directional video information and vehicle-to-everything (V2X) connectivity, the safety and dignity of blind people traveling have been addressed, providing panoramic perception and risk warnings, thus improving the safety and dignity of blind people traveling.

CN121330643APending Publication Date: 2026-01-13SHAANXI FASHION ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511460994.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing smart glasses and white canes can only calculate the distance between blind people and surrounding objects when assisting them in traveling, and cannot fully solve the safety and dignity issues of blind people in the process of traveling.

Method used

A multimodal AI-based full-scene perception and interaction method for the blind is adopted. By acquiring and analyzing video information from multiple directions, combined with traffic signals, environmental and interpersonal communication information, audio prompts are generated, and vehicle-to-everything (V2X) is used for risk warning to ensure the safety and dignity of the blind.

Benefits of technology

It enables blind people to quickly obtain traffic signals, surrounding environmental risks, and the true meaning of interpersonal communication during travel, reducing the risk of collisions, ensuring safety, and maintaining dignity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330643A_ABST
    Figure CN121330643A_ABST
Patent Text Reader

Abstract

The multi-mode AI-based blind person full-scene perception interaction method comprises the following steps: in the moving process of a user, acquiring a first video in a first direction, and acquiring a second video in a second direction based on the selection of the user; using a pre-trained traffic signal analysis model to locally analyze the first video and the second video to obtain traffic signal information, and converting the traffic signal information into traffic signal audio; uploading the first video and the second video to a cloud end, analyzing the first video and the second video by using a pre-trained complex environment analysis model to obtain complex environment information, and converting the complex environment information into complex environment audio; and acquiring position data of the user in real time, uploading the position data to the Internet of Vehicles to generate risk early warning information, and converting the risk early warning information into risk early warning audio. The safety of a user in the moving process can be fully guaranteed, and dignity of the user is effectively prevented from being damaged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of walking assistance tools for the blind, specifically a method and system for full-scene perception and interaction for the blind based on multimodal AI. Background Technology

[0002] With the continuous development and progress of social technology, society is becoming increasingly intelligent. While enjoying the convenience and speed brought by intelligent technology, people also have a responsibility to care for vulnerable groups in society. Blind people (visually impaired or blind individuals), as a vulnerable group, face significant inconvenience in walking and daily life due to their visual impairment; perceiving the outside world is crucial for them. Currently, most blind people rely on canes to sense tactile paving and obstacles on roads, while some rely on walking dogs for assistance. However, training walking dogs requires a long time, and their limited number and high cost mean that most blind people cannot obtain assistance from them. Using canes alone is often unsafe.

[0003] In existing technologies, some smart glasses and white canes can solve the problem of blind people's travel difficulties to a certain extent. However, most of these technologies only focus on calculating the distance between blind people and surrounding objects and prompting blind people to avoid surrounding risks. In reality, the difficulties that blind people face when traveling are not limited to the risk of collision. Therefore, the role that existing smart glasses and white canes can play in assisting blind people to travel is still limited. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a multimodal AI-based full-scene perception and interaction method and system for the blind. This system can quickly inform users of traffic signals ahead, accurately inform users of environmental risks, and convey the true meaning of others during interactions, thereby fully ensuring the user's safety during movement and effectively protecting the user's dignity.

[0005] To achieve the above objectives, the specific solution adopted by the present invention is as follows:

[0006] A multimodal AI-based method for full-scene perception and interaction for the blind includes the following steps:

[0007] During the user's movement, a first video in a first direction is acquired, and a second video in a second direction is acquired based on the user's selection, wherein the first direction and the second direction may be the same or different;

[0008] The traffic signal information is obtained by parsing the first video and the second video locally using a pre-trained traffic signal parsing model, and the traffic signal information is converted into traffic signal audio.

[0009] The first video and the second video are uploaded to the cloud and a pre-trained complex environment parsing model is used to parse the first video and the second video to obtain complex environment information, and the complex environment information is converted into complex environment audio.

[0010] The system collects user location data in real time and uploads it to the vehicle network to generate risk warning information, and converts the risk warning information into risk warning audio.

[0011] Play the traffic signal audio, the complex environment audio, and the risk warning audio to the user.

[0012] Preferably, during the user's movement, the first video is acquired using a first device fixed to the user's head, and the second video is acquired using a second device held by the user, with the second direction being changed when the user rotates the second device.

[0013] Preferably, after the first device acquires the first video, it performs frame parsing on the first video to obtain multiple first images. After the second device acquires the second video, it performs frame parsing on the second video to obtain multiple second images. The first device and the second device exchange the first images and the second images, and the second device uploads the first images and the second images to the cloud. The first device uses a pre-trained traffic signal parsing model to parse the first images and the second images locally to obtain traffic signal information.

[0014] Preferably, the traffic signal analysis model is trained based on the YOLOv8-tiny model, and the complex environment analysis model includes an environment analysis part trained based on the YOLOv8-tiny model and a communication analysis part trained based on the ResNet model. The environment analysis part is used to analyze the first image and the second image to generate surrounding environment information, and the communication analysis part is used to analyze the first image and the second image to generate interpersonal communication information.

[0015] Preferably, when the first device performs frame parsing on the first video, it extracts the first audio, and when the second device performs frame parsing on the second video, it extracts the second audio. After the first device sends the first audio to the second device, the second device uploads the first audio and the second audio to the cloud. The complex environment parsing model includes an audio parsing part trained based on an LSTM model. The audio parsing part is used to parse the first audio and the second audio to generate communication semantic information.

[0016] Preferably, the method for real-time collection of user location data and uploading it to the vehicle network to generate risk warning information includes:

[0017] The location data is uploaded to the vehicle network;

[0018] In the Internet of Vehicles (IoV), identify multiple vehicles whose distance from the user is less than a preset first threshold and mark them as target vehicles;

[0019] The risk warning information is generated when the distance between at least one of the target vehicles and the user is less than a preset second threshold.

[0020] Preferably, after the target vehicle is determined, the user's location data is displayed in real time on the vehicle's infotainment system.

[0021] Preferably, the risk warning information includes the risk level, risk direction, and avoidance prompts.

[0022] A multimodal AI-based full-scene perception and interaction system for the blind, used to implement the aforementioned multimodal AI-based full-scene perception and interaction method for the blind, the system comprising:

[0023] A first device is used to acquire the first video while the user is moving;

[0024] The second device is for use by the user to hold and to acquire the second video while the user is moving.

[0025] A cloud server is used to parse the first video and the second video using a pre-trained complex environment parsing model to obtain complex environment information, convert the complex environment information into complex environment audio, and send the complex environment audio to the second device.

[0026] The vehicle-to-everything (V2X) subsystem is used to build the V2X network.

[0027] Preferably, the first device is eyeglasses and the second device is a white cane.

[0028] This invention can quickly inform users of traffic signals ahead, accurately inform users of surrounding environmental risks, and convey the true meaning of others during communication, thereby fully ensuring the user's safety during movement and effectively protecting the user's dignity. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart of the method of the present invention;

[0031] Figure 2 This is a structural block diagram of the system of the present invention;

[0032] Figure 3 This is a structural block diagram of the glasses and guide cane in a specific implementation. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] This invention is applicable to guides, that is, to assist blind people in their daily lives. Hereinafter, blind people will be referred to as users.

[0035] like Figure 1 As shown, the present invention first provides a method for full-scene perception and interaction for the blind based on multimodal AI, including S1 to S5.

[0036] S1. During the user's movement, acquire a first video in a first direction, and acquire a second video in a second direction based on the user's selection. The first and second directions may be the same or different. Specifically, the first direction is in front of the user, and the second direction can be selected according to the user's needs, such as in front of the user, or to the user's left, right, or back, etc.

[0037] To fix the first direction and facilitate the user's selection of the second direction, a first video is acquired using a first device fixed to the user's head during movement, and a second video is acquired using a second device held by the user. The second direction is changed when the user rotates the second device. In one embodiment of the invention, the first device is a pair of glasses, and the second device is a white cane. The first device is fixedly worn on the user's head, and the second device is held by the user, thus facilitating flexible rotation of the second device during movement to select the second direction.

[0038] Furthermore, after acquiring the first video, the first device performs frame parsing to obtain multiple first images. Similarly, after acquiring the second video, the second device performs frame parsing to obtain multiple second images. The first and second devices exchange the first and second images, and the second device uploads both images to the cloud. By performing frame parsing on both the first and second devices, the pressure on the cloud is reduced, preventing slow data processing due to excessive cloud complexity and thus avoiding long waiting times for users.

[0039] On the other hand, when the first device performs frame parsing on the first video, it extracts the first audio, and when the second device performs frame parsing on the second video, it extracts the second audio. After the first device sends the first audio to the second device, the second device uploads the first audio and the second audio to the cloud.

[0040] S2. Using a pre-trained traffic signal parsing model, the first and second videos are parsed locally to obtain traffic signal information, and the traffic signal information is converted into traffic signal audio. More specifically, the first device uses the pre-trained traffic signal parsing model to locally parse the first and second images to obtain traffic signal information. The traffic signal parsing model is trained based on the YOLOv8-tiny model. The YOLOv8-tiny model and its training method are conventional techniques in this field and will not be elaborated upon here.

[0041] S3. Upload the first and second videos to the cloud and use a pre-trained complex environment parsing model to parse the first and second videos to obtain complex environment information, and then convert the complex environment information into complex environment audio. The complex environment parsing model includes an environment parsing part trained on a YOLOv8-tiny model and a communication parsing part trained on a ResNet model. The environment parsing part is used to parse the first and second images to generate surrounding environment information, and the communication parsing part is used to parse the first and second images to generate interpersonal communication information. Additionally, the complex environment parsing model includes an audio parsing part trained on an LSTM model, which is used to parse the first and second audio audio to generate communication semantic information. The ResNet and LSTM models, as well as the specific training methods, are mature existing technologies in this field and will not be elaborated further. The traffic signal parsing model and the complex environment parsing model employ different models for their environment parsing, communication parsing, and audio parsing parts, allowing for targeted training for different types of data and needs, ensuring more efficient processing of various data types and higher accuracy of the parsing results.

[0042] Because the first direction, i.e., the area in front of the user, directly affects the user's movement, the processing time of the first video is crucial. This invention processes the first video locally on a first device to generate traffic signal information, helping the user quickly determine whether they can continue moving forward without waiting for cloud processing results, thus eliminating the time spent on data upload and download, resulting in better timeliness. Furthermore, because the first and second devices exchange the first and second images, the second device simultaneously possesses all the first and second images. Uploading the first and second images to the cloud using the second device allows for more comprehensive processing of the first image with the help of the cloud's powerful computing capabilities, generating complex environmental information to help the user assess whether there are risks in their surroundings. The first and second images held by the first device can serve as backups, used both to review the user's movement process and for subsequent optimization training of the traffic signal analysis model and the complex environment analysis model.

[0043] On the other hand, users also need to communicate with others in their daily lives. To help users more accurately grasp the true meaning of the communication partner, this invention simultaneously extracts the first and second audio recordings. Then, it uses the audio analysis component of a complex environment analysis model to analyze the first and second audio recordings, generating communication semantic information. Simultaneously, the communication analysis component uses the first and second images to analyze the user's facial expressions, generating interpersonal communication information. This interpersonal communication information and communication semantic information are then matched. If a match is successful, it indicates that the communication partner's words truthfully reflect their own thoughts. If a match fails, it indicates that the communication partner's words do not truthfully reflect their own thoughts. For example, if the communication partner's words express a pleasant meaning, but their facial expression contradicts this, it suggests that the communication partner is lying. By analyzing the true state of the communication partner, the user's dignity can be better protected, and the user can successfully engage in interpersonal communication.

[0044] Based on the above analysis process of the first and second videos, the present invention can quickly inform the user of the traffic signal situation ahead, accurately inform the user of the surrounding environmental risks, and also inform the user of the other party's true meaning during communication, thereby fully ensuring the user's safety during movement and effectively protecting the user's dignity from infringement.

[0045] To further ensure user safety during movement, in addition to informing the user of the risks in the surrounding environment as described above, this invention also uses vehicle-to-everything (V2X) technology to alert nearby vehicles, enabling them to proactively avoid users with mobility impairments. The specific method includes step S4.

[0046] S4. Real-time collection of user location data and uploading it to the vehicle network to generate risk warning information, and conversion of the risk warning information into risk warning audio. Specifically, the methods for real-time collection of user location data and uploading it to the vehicle network to generate risk warning information include S41 to S43.

[0047] S41. Upload location data to the vehicle-to-everything (V2X) network. The V2X network is a mature, existing technology; its specific structure and working principles will not be elaborated upon here.

[0048] S42. In the vehicle-to-everything (V2X) network, identify multiple vehicles whose distance to the user is less than a preset first threshold and mark them as target vehicles. Specifically, location sensing modules can be installed in the first and second devices, using existing navigation technologies such as GPS and BeiDou navigation systems to determine the user's location data. After the location data is uploaded to the V2X network, the distance between the vehicle and the user is calculated based on the vehicle's real-time location and the user's location data. When this distance is less than the first threshold, the vehicle is marked as a target vehicle. The target vehicle may come into contact with the user during subsequent travel, posing a danger. It should be noted that the distance value here is the straight-line distance between the vehicle and the user. In real-world scenarios, due to road limitations, the traffic distance between the vehicle and the user may be greater than this distance value; that is, the actual distance the vehicle travels to the user's location along the road will be greater than this distance value. However, the straight-line distance calculation is faster. When both the vehicle and the user are in motion, it can quickly determine whether the vehicle may come into contact with the user. Furthermore, because the traffic distance will be greater than or equal to this distance value, this method can inform the vehicle in advance of the presence of a blind user nearby, thus more fully ensuring the user's safety.

[0049] Furthermore, once the target vehicle is identified, the user's location data is displayed in real time on the vehicle's infotainment system. Specifically, the user's location can be displayed in real time on the navigation software within the vehicle's infotainment system, alerting the driver to the presence of a visually impaired user nearby, allowing the driver to proactively avoid the user and ensure their safety.

[0050] S43. When the distance between at least one target vehicle and the user is less than a preset second threshold, a risk warning message is generated. The risk warning message includes the risk level, risk direction, and avoidance prompts. If the target vehicle is close enough to the user, i.e., the distance between the target vehicle and the user is less than the second threshold, it indicates that the user's risk is higher, and the user needs to be prompted to actively avoid the vehicle. In the risk warning message, the risk level can be set to high risk, medium risk, and low risk; the risk direction can be forward, left, right, or backward; and the avoidance prompts can be forward avoidance, backward avoidance, or turning avoidance, etc. Through the risk warning message, the user can be prompted to actively avoid the vehicle, further ensuring the user's safety.

[0051] Based on S4, the present invention can first prompt the vehicle to actively avoid the user when the distance between the vehicle and the user is relatively close, and then prompt the user to actively avoid the user when the distance is too close. Through the active avoidance by both the vehicle and the user, the risk level of the user can be significantly reduced and the user can be prevented from being in danger.

[0052] S5. Play traffic signal audio, complex environment audio, and risk warning audio to the user. To facilitate the playback of traffic signal audio, complex environment audio, and risk warning audio to the user, a voice broadcast module can be added to the first device, i.e., the glasses. The voice broadcast module can use a speaker or bone conduction generator to play audio to the user. The specific settings are all mature existing technologies and will not be described in detail here.

[0053] like Figure 2 As shown, the present invention further provides a multi-modal AI-based full-scene perception and interaction system for the blind, used to implement the above-mentioned multi-modal AI-based full-scene perception and interaction method for the blind. The system includes a first device, a second device, a cloud server, and a vehicle networking subsystem.

[0054] A first device is used to acquire a first video while the user is moving. In one embodiment of the invention, the first device is configured as glasses.

[0055] A second device is provided for the user to hold and for acquiring a second video while the user is moving. In one embodiment of the invention, the second device is a white cane.

[0056] The cloud server is used to parse the first and second videos using a pre-trained complex environment parsing model to obtain complex environment information, convert the complex environment information into complex environment audio, and send the complex environment audio to the second device. The cloud server can be a conventional general-purpose server built on an x86 processor or an ARM processor, which is a mature existing technology and will not be elaborated further here.

[0057] The vehicle-to-everything (V2X) subsystem is used to build the V2X network. The V2X subsystem is also a mature existing technology, and will not be elaborated upon here.

[0058] Furthermore, such as Figure 3As shown, both the glasses and the white cane include a data processing module, which is electrically connected to an image acquisition module, a wireless communication module, a positioning module, a sound playback module, a sound acquisition module, and a data storage module. The image acquisition module is used to capture a first or second video. The wireless communication module is used for exchanging the first and second images between the glasses and the white cane, and also for the white cane to interact with a cloud server. The positioning module is used to determine the user's location data; the location data determined by the two positioning modules can be cross-checked. The sound playback module is used to play the aforementioned traffic signal audio, complex environment audio, and risk warning audio to the user. The sound acquisition module is used to acquire audio during user interactions or to acquire sounds such as horns from surrounding vehicles. The data processing module, image acquisition module, wireless communication module, positioning module, sound playback module, sound acquisition module, and data storage module are all mature existing technologies and will not be elaborated further here. For example, the mechanical and circuit structures of existing smart glasses and smart white canes can be used.

[0059] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0060] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multimodal AI-based full-scene perception and interaction method for the blind, characterized in that, Includes the following steps: During the user's movement, a first video in a first direction is acquired, and a second video in a second direction is acquired based on the user's selection, wherein the first direction and the second direction may be the same or different; The traffic signal information is obtained by parsing the first video and the second video locally using a pre-trained traffic signal parsing model, and the traffic signal information is converted into traffic signal audio. The first video and the second video are uploaded to the cloud and a pre-trained complex environment parsing model is used to parse the first video and the second video to obtain complex environment information, and the complex environment information is converted into complex environment audio. The system collects user location data in real time and uploads it to the vehicle network to generate risk warning information, and converts the risk warning information into risk warning audio. Play the traffic signal audio, the complex environment audio, and the risk warning audio to the user.

2. The method for full-scene perception and interaction for the blind based on multimodal AI as described in claim 1, characterized in that, During the user's movement, the first video is acquired using a first device fixed to the user's head, and the second video is acquired using a second device held by the user. The second direction is changed when the user rotates the second device.

3. The method for full-scene perception and interaction for the blind based on multimodal AI as described in claim 2, characterized in that, After the first device acquires the first video, it performs frame parsing on the first video to obtain multiple first images. After the second device acquires the second video, it performs frame parsing on the second video to obtain multiple second images. The first device and the second device exchange the first images and the second images, and the second device uploads the first images and the second images to the cloud. The first device uses a pre-trained traffic signal parsing model to parse the first images and the second images locally to obtain traffic signal information.

4. The method for full-scene perception and interaction for the blind based on multimodal AI as described in claim 3, characterized in that, The traffic signal analysis model is trained based on the YOLOv8-tiny model. The complex environment analysis model includes an environment analysis part trained based on the YOLOv8-tiny model and a communication analysis part trained based on the ResNet model. The environment analysis part is used to analyze the first image and the second image to generate surrounding environment information, and the communication analysis part is used to analyze the first image and the second image to generate interpersonal communication information.

5. The method for full-scene perception and interaction for the blind based on multimodal AI as described in claim 4, characterized in that, When the first device performs frame parsing on the first video, it extracts the first audio. When the second device performs frame parsing on the second video, it extracts the second audio. After the first device sends the first audio to the second device, the second device uploads the first audio and the second audio to the cloud. The complex environment parsing model includes an audio parsing part trained based on an LSTM model. The audio parsing part is used to parse the first audio and the second audio to generate communication semantic information.

6. The method for full-scene perception and interaction for the blind based on multimodal AI as described in claim 1, characterized in that, The method for collecting user location data in real time and uploading it to the vehicle network to generate risk warning information includes: The location data is uploaded to the vehicle network; In the Internet of Vehicles (IoV), identify multiple vehicles whose distance from the user is less than a preset first threshold and mark them as target vehicles; The risk warning information is generated when the distance between at least one of the target vehicles and the user is less than a preset second threshold.

7. The method for full-scene perception and interaction for the blind based on multimodal AI as described in claim 6, characterized in that, Once the target vehicle is identified, the user's location data is displayed in real time on the vehicle's infotainment system.

8. The method for full-scene perception and interaction for the blind based on multimodal AI as described in claim 6, characterized in that, The risk warning information includes the level of risk, the direction of risk, and avoidance tips.

9. A multimodal AI-based full-scene perception and interaction system for the blind, characterized in that: The system is used to implement the multimodal AI-based full-scene perception and interaction method for the blind as described in any one of claims 1-8, the system comprising: A first device is used to acquire the first video while the user is moving; The second device is for use by the user to hold and to acquire the second video while the user is moving. A cloud server is used to parse the first video and the second video using a pre-trained complex environment parsing model to obtain complex environment information, convert the complex environment information into complex environment audio, and send the complex environment audio to the second device. The vehicle-to-everything (V2X) subsystem is used to build the V2X network.

10. The multimodal AI-based full-scene perception and interaction system for the blind as described in claim 9, characterized in that, The first device is eyeglasses, and the second device is a white cane.