Real-time remote bone conduction blind-assisting auxiliary method for family guarding

By integrating low-latency video acquisition and wireless communication modules into smart blind-assisting glasses, combined with bone conduction speakers and multimodal large models, real-time interaction and remote navigation guidance between visually impaired people and their families are achieved, solving the real-time and interactivity issues of traditional blind-assisting devices and providing safe and efficient remote protection.

CN120753874APending Publication Date: 2025-10-10ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510924314.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Traditional blind-assistance devices lack real-time and interactivity, and are difficult to cope with complex or dynamic environments. Voice prompts affect the perception of environmental sounds, and they lack the function of remote monitoring by family members.

Method used

Low-latency video acquisition and wireless communication modules are integrated into smart blind-assisting glasses to transmit the wearer's field of view video to family members in real time. Combined with bone conduction speaker broadcast instructions and multimodal large models to generate auxiliary information, remote interaction and navigation guidance for family members are possible.

Benefits of technology

It improves the real-time performance and safety of the blind assistance system, meets the blind's needs for environmental sounds, provides precise obstacle avoidance and navigation assistance, and realizes remote protection functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120753874A_ABST
    Figure CN120753874A_ABST
Patent Text Reader

Abstract

The invention discloses a family guarding real-time remote bone conduction blind-assisting assisting method, which comprises the following steps: integrating a low-delay video acquisition and wireless communication module in intelligent blind-assisting glasses, and acquiring and transmitting a real-time video stream of the visual field of a wearer to family end equipment; the family end checks scenes seen by the blind person in real time through an interface, and generates an instruction by using a voice communication or text labeling function; the instruction is wirelessly returned and broadcasted to a wearer through a bone conduction loudspeaker on the glasses, and meanwhile, the system combines a multi-mode large model guiding module to generate auxiliary information; and the wearer performs obstacle avoidance, navigation or surrounding information recognition according to the instruction and the auxiliary information. Real-time performance and interactivity are guaranteed, the requirement of the blind person for environment sound is met through the bone conduction technology, and privacy protection is considered.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent assistive technology and telecommunication, in particular to a real-time remote bone conduction blind assistance method for family guardianship. BACKGROUND

[0002] With the development of intelligent technology and wireless communication technology, blind assistance devices play an increasingly important role in improving the quality of life of visually impaired people. Traditional blind assistance devices mainly rely on ultrasonic sensors, infrared ranging or simple voice prompts to help visually impaired people perceive the surrounding environment. However, these devices have the following problems: first, they lack real-time and interactivity, making it difficult to cope with complex or dynamic environments; second, they mainly play voice prompts through earphones or loudspeakers, which can easily cover environmental sounds and affect the perception of visually impaired people to the surrounding sounds, thereby reducing safety; third, they lack remote guardianship functions involving family members, and visually impaired people may face unexpected situations when going out and cannot obtain timely help. With the advancement of multi-modal large model technology, it is possible to provide more intelligent and more personalized assistance to visually impaired people using video streams, voice interaction and intelligent analysis technology. Therefore, it is necessary to develop a blind assistance system that integrates real-time remote guardianship and bone conduction technology. SUMMARY

[0003] The present application overcomes the shortcomings of the prior art, based on the advantages of intelligent assistive technology and telecommunication technology, and provides a real-time remote bone conduction blind assistance method for family guardianship.

[0004] The real-time remote bone conduction blind assistance method for family guardianship of the present application comprises the following steps:

[0005] S110, integrating a low-latency video acquisition and wireless communication module in the intelligent blind assistance glasses, collecting and transmitting real-time video streams of the wearer's field of view to the family end device;

[0006] S120, the family end viewing the scene seen by the blind person in real time through the interface, and generating instructions using voice calls or text annotation functions;

[0007] S130, the instructions are transmitted back wirelessly and broadcast to the wearer by the bone conduction speaker on the glasses, while the system generates auxiliary information in combination with the multi-modal large model guide module;

[0008] S140, the wearer avoids obstacles, navigates or identifies surrounding information according to the instructions and auxiliary information.

[0009] Further, the low-latency video acquisition and wireless communication module integrated in the intelligent blind assistance glasses of step S110 specifically comprises:

[0010] S1101, select high-resolution, low-latency camera, real-time collection of video data within the wearer's field of view;

[0011] S1102, integrated wireless communication module (such as 5G or Wi-Fi), to ensure high-speed and stable transmission of video stream to family end device.

[0012] Further, the family end described in step S120 real-time view the scene seen by the blind through the interface and generate instructions, specifically including:

[0013] S1201, display real-time video stream on the family end device, provide intuitive interface of blind person's field of view;

[0014] S1202, family members send voice instructions directly through voice calls, or mark key information (such as obstacle position or navigation direction) on the video interface through text annotation.

[0015] Further, the instructions described in step S130 return and broadcast, and the auxiliary information generated by the multi-modal large model, specifically including:

[0016] S1301, return the instructions generated by the family end to the intelligent blind aid glasses through the wireless communication module;

[0017] S1302, use bone conduction speaker to broadcast instructions to the wearer, to ensure that the perception of environmental sound is not hindered;

[0018] S1303, multi-modal large model analyzes video stream and environmental data, generates auxiliary information for obstacle avoidance, navigation or object recognition.

[0019] Further, the wearer performs operations according to the instructions and auxiliary information described in step S140, specifically including:

[0020] S1401, the wearer avoids obstacles or adjusts walking direction according to the instructions broadcast by the bone conduction speaker;

[0021] S1402, in combination with the auxiliary information of the multi-modal large model, identify surrounding objects or complete navigation tasks;

[0022] S1403, in emergency situations, the wearer can contact family members through one-key call function to obtain real-time guidance.

[0023] This invention integrates low-latency video capture and wireless communication modules into smart glasses, enabling real-time streaming of the wearer's field of view to family devices. Family members can view the scene the blind person sees through the interface and generate instructions, which are then wirelessly transmitted back and broadcast by bone conduction speakers. Simultaneously, a multimodal large model is combined to generate obstacle avoidance, navigation, or identification information. The system ensures real-time and interactivity while meeting the blind person's need for ambient sound through bone conduction technology, while also protecting privacy. This method is widely applicable to areas such as daily travel, shopping, emergency rescue, and remote companionship, providing safe and efficient assistance and support for the visually impaired.

[0024] The beneficial effects of the present invention are that the present invention adopts low-latency video acquisition and wireless communication technology to realize real-time interaction between family members and visually impaired people, significantly improving the real-time performance and safety of the blind assistance system; the application of bone conduction speakers not only ensures the clear transmission of voice commands, but also does not hinder the perception of ambient sounds, fully meeting the needs of visually impaired people for ambient sounds; the introduction of multimodal large models enables the system to intelligently analyze environmental information and provide more accurate obstacle avoidance and navigation assistance; in addition, the system's remote guarding function allows family members to provide companionship and support to visually impaired people in other places, which is suitable for daily travel, shopping, tourism and other scenarios, and has high practicality and humanized value. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 Flowchart of the method of the present invention. DETAILED DESCRIPTION

[0026] The technical solution of the present invention is described below with reference to the accompanying drawings.

[0027] This embodiment provides a real-time, remote bone conduction assistance method for the visually impaired, with family members providing assistance. By combining smart glasses, remote communication technology, and a multimodal large model, this method provides real-time, safe assistance services for the visually impaired. This technology can be widely used in areas such as daily travel, shopping, emergency rescue, and remote companionship.

[0028] Specifically, the method includes:

[0029] S110 integrates low-latency video capture and wireless communication modules into smart glasses for blindness-aiding vision, capturing and transmitting real-time video streams of the wearer's field of view to family devices;

[0030] Specifically, a high-resolution, low-latency camera is used to collect video data within the wearer's field of view in real time; an integrated wireless communication module (such as 5G or Wi-Fi) ensures high-speed and stable transmission of video streams to family devices. Specifically, it includes:

[0031] 1) Video data representation: A real-time video stream can be viewed as a series of image frames arranged in time order. Mathematically, each frame image Ft at time point t can be represented as a three-dimensional matrix:

[0032] Ft=[Pi,j,c]∈RW×H×C

[0033] Where W and H represent the width and height (resolution) of the image, respectively, C represents the number of color channels (for example, RGB has 3 channels), and Pi,j,c is the intensity value of the pixel at position (i,j) in color channel c.

[0034] 2) Data compression: To achieve high-speed and stable transmission, the original video data needs to be compressed. This process uses video coding standards (such as H.265 / HEVC), the core of which is to use mathematical transformations to reduce data redundancy. For example, the discrete cosine transform (DCT) is used to convert the image from the spatial domain to the frequency domain, thereby discarding high-frequency information that the human eye is not sensitive to, in order to achieve the purpose of compression. The compression process can be expressed as a function C: Vcompressed = C(Vraw)

[0035] Among them, Vraw is the original video stream, and Vcompressed is the compressed data stream, whose data volume is much smaller than Vraw.

[0036] 3) Transmission Delay Model: The total delay Ttotal is the key to ensuring real-time performance. It can be modeled as the sum of the delays in each processing stage: Ttotal = Tcapture + Tencode + Ttransmit + Tdecode

[0037] Where Tcapture is the acquisition delay, Tencode is the encoding delay, Ttransmit is the transmission delay, and Tdecode is the decoding delay at the home end. Transmission delay itself depends on the size of the compressed data and the network bandwidth: Ttransmit = BandwidthSize(Vcompressed). The system's goal is to minimize Ttotal to ensure real-time interaction.

[0038] S120, the family member views the scene seen by the blind person in real time through the interface and generates instructions using the voice call or text annotation function;

[0039] Specifically, the real-time video stream is displayed on the family's device, providing an intuitive visual interface for the blind; family members can directly send voice commands through voice calls, or mark key information (such as obstacle locations or navigation directions) on the video interface through text annotations. Specifically, it includes:

[0040] 1) Speech command digitization: The family's voice command first needs to be digitized through analog-to-digital conversion (ADC). This process includes sampling and quantization. A continuous sound wave signal s(t) is converted into a discrete digital signal sequence s[n]: s[n] = Quantize(s(t)) | t = nTs

[0041] where Ts is the sampling period, and Quantize is the quantization function that maps the amplitude of the sampling points to a finite set of numbers.

[0042] 2) Text annotation coordinate mapping: When the family performs text annotation on the video interface, the system needs to record the location of the annotation. This process is a two-dimensional coordinate mapping. If the family performs annotation at screen coordinates (xs, ys), the system converts it to actual coordinates (xv, yv) within the video frame through an affine transformation T: xv yv1 = Tx s y s1, which takes into account the scaling and offset of the video picture on the screen.

[0043] S130, the command is returned through wireless communication and broadcast to the wearer by the bone conduction speaker on the glasses, while the system generates auxiliary information in combination with the multi-modal large model guidance module;

[0044] Specifically, the instructions generated by the family end are returned to the intelligent blind aid glasses through the wireless communication module; the bone conduction speaker is used to broadcast the instructions to the wearer, ensuring that the perception of environmental sound is not hindered; the multi-modal large model analyzes the video stream and environmental data to generate auxiliary information for obstacle avoidance, navigation, or object recognition. Specifically, it includes:

[0045] 1) Multi-modal large model analysis: This is the core intelligence of the system. The model can be represented as a highly complex nonlinear function Mθ, where θ represents the vast number of parameters learned by the model after training.

[0046] (O, N, I) = Mθ(V, Denv)

[0047] where V is the input video stream data, and Denv is other possible environmental data (such as motion data from IMU). The output of the model is structured auxiliary information, including:

[0048] 2) Obstacle recognition O: The model analyzes the video frame and outputs a set of detected obstacles O = {o1, o2,..., ok}. Each obstacle oi is defined by its category ci (such as "step", "door") and location (usually a bounding box coordinate bi = (x, y, w, h)). Its category judgment is usually calculated by a Softmax function to calculate the probability distribution: P(cj | Ft) = ∑k=1Kezkezj

[0049] where zj is the output score of the model for class j.

[0050] 3) Navigation Instructions N: By semantically segmenting the scene (classifying each pixel in the image as "traversable area," "obstacle," "sidewalk," etc.), the model constructs a traversable path graph. It then uses a graph search algorithm such as A* (A-star) or Dijkstra to calculate an optimal path from the current location to the destination within the graph. The A* algorithm finds a path by minimizing a cost function f(n) = g(n) + h(n), where g(n) is the cost of the current path and h(n) is the estimated cost to reach the destination (a heuristic function).

[0051] 4) Audio signal synthesis: Whether it is the voice command of the family or the auxiliary voice generated by the model, it needs to be converted from the digital signal back to the analog electrical signal through the digital-to-analog conversion (DAC) to drive the bone conduction speaker.

[0052] S140: The wearer avoids obstacles, navigates, or identifies surrounding information according to the instructions and auxiliary information.

[0053] Specifically, the wearer can avoid obstacles or adjust walking direction according to the instructions broadcast by the bone conduction speaker; combined with the auxiliary information of the multimodal large model, the wearer can identify surrounding objects or complete navigation tasks; in an emergency, the wearer can contact family members through the one-click call function to obtain real-time guidance. Specifically, it includes:

[0054] 1) Closed-Loop Control System: The entire system can be viewed as a closed-loop feedback control system centered around the wearer. The wearer's state (position, orientation) at time t is St. The system generates guidance instructions It based on the scene V(St) input from the camera. The wearer performs actions At (such as adjusting walking direction) based on It, thereby changing their state to St+1. This new state, in turn, generates new video input V(St+1), forming a continuous loop.

[0055] 2) Risk Minimization: The system's goal is to guide the wearer along a path P that minimizes the risk function R(P). The risk function can be defined as the sum or integral of the reciprocal distances from each point on the path to the nearest obstacle. The navigation instructions generated by the system aim to solve the following optimization problem:

[0056] PminR(P)subject toP∈WalkableArea

[0057] This ensures that the resulting path is not only the shortest or fastest, but also the safest.

[0058] 3) Emergency call: The one-touch call function can be mathematically regarded as an interrupt signal, which will give the communication channel between family members and the wearer the highest priority to ensure instant connection.

[0059] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The scope of protection of the present invention should not be regarded as limited to the specific forms described in the embodiments. The scope of protection of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A real-time remote bone conduction blind assistance method for family members, characterized in that: The following steps are involved: S110 integrates low-latency video capture and wireless communication modules into smart glasses for blindness-aiding vision, capturing and transmitting real-time video streams of the wearer's field of view to family devices; S120, the family member views the scene seen by the blind person in real time through the interface and generates instructions using the voice call or text annotation function; S130: The command is wirelessly transmitted back and broadcast to the wearer through the bone conduction speaker on the glasses. At the same time, the system generates auxiliary information in conjunction with the multimodal large model guidance module. S140: The wearer avoids obstacles, navigates, or identifies surrounding information according to the instructions and auxiliary information.

2. The method according to claim 1, characterized in that Integrating a low-latency video acquisition and wireless communication module into the smart blind-aiding glasses in step S110 specifically includes: S1101 uses a high-resolution, low-latency camera to collect real-time video data within the wearer's field of view; S1102 integrates a wireless communication module (such as 5G or Wi-Fi) to ensure high-speed and stable transmission of video streams to family devices.

3. The method according to claim 1, wherein: The family member in step S120 views the scene seen by the blind person in real time through the interface and generates instructions, which specifically includes: S1201, displays real-time video streams on family devices, providing an intuitive interface for the blind; S1202, the family member directly sends a voice command through a voice call, or marks key information on the video interface through text annotation.

4. The method according to claim 1, wherein: The command transmission and broadcasting described in step S130, as well as the generation of auxiliary information for the multimodal large model, specifically include: S1301, transmitting the command generated by the family member to the smart blind-aiding glasses via the wireless communication module; S1302, using a bone conduction speaker to broadcast the instructions to the wearer, ensuring that the perception of ambient sound is not obstructed; S1303: The multimodal large model analyzes the video stream and environmental data to generate auxiliary information for obstacle avoidance, navigation, or object recognition.

5. The method according to claim 1, wherein: The wearer in step S140 performs operations according to the instructions and auxiliary information, specifically including: S1401, the wearer avoids obstacles or adjusts walking direction according to the instructions broadcast by the bone conduction speaker; S1402, combining auxiliary information from the multimodal large model to identify surrounding objects or complete navigation tasks; S1403, in an emergency, the wearer can contact family members through the one-touch calling function and get real-time guidance.