AI intelligent glasses with visual impairment assisting function

Through the integration of AI smart glasses, a variety of sensors and algorithms are solved, the problem of limited functions of visual impairment assistive tools is achieved, safe navigation and information acquisition for visually impaired users, and the quality of life is improved.

CN120360769AInactive Publication Date: 2025-07-25SHENZHEN LINGXI ZHIXIN TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510430036.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing visual impairment assistance tools such as guide sticks and guide dogs have limited functions or limitations in use, and cannot provide rich environmental information. Moreover, the lack of visual impairment assistance devices for smart devices, resulting in mobility difficulties and difficulty in obtaining information for visually impaired people.

Method used

Design an AI smart glasses, integrating a wide-angle camera, depth sensor, inertial measurement unit, geomagnetic sensor, GPS positioning module, microphone array, speaker and storage module, combining environment perception and object recognition algorithm, navigation algorithm, text recognition and reading algorithm, and voice interaction algorithm to realize real-time environmental perception, navigation and information acquisition.

Benefits of technology

Through accurate environmental perception and obstacle detection, visually impaired users can avoid dangers, quickly identify text information, provide autonomous navigation and natural voice interaction, and improve visually impaired users' freedom of movement and information acquisition ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120360769A_ABST
    Figure CN120360769A_ABST
Patent Text Reader

Abstract

The invention discloses AI intelligent glasses with a visual impairment assisting function, and belongs to the technical field of artificial intelligence. Comprising a glasses body, and a hardware system and a software system which are arranged on the glasses body, the hardware system comprises a wide-angle camera, a depth sensor, an inertial measurement unit, a geomagnetic sensor, an AI processing chip, a microphone array, a loudspeaker, a power supply module and a storage module; the wide-angle camera is used for collecting image data of the surrounding environment of the glasses body, the depth sensor is used for measuring the distance between an object and the glasses body, the inertial measurement unit is used for obtaining posture information of the glasses body, the geomagnetic sensor is used for positioning the direction of the glasses body, and the AI processing chip is used for processing data in real time. The microphone array is used for receiving a voice instruction of a user, the loudspeaker is used for outputting voice feedback information, and the power supply module is used for supplying power; the software system comprises an environment perception and object recognition algorithm, a navigation algorithm, a text recognition and reading algorithm and a voice interaction algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a pair of AI smart glasses with visually impaired assistance. Background Art

[0002] Visual impairment brings many inconveniences to people's lives, limiting their freedom of movement and ability to obtain information. At present, the main aids for the visually impaired are guide sticks and guide dogs. Guide sticks can only provide limited obstacle detection functions and cannot obtain richer environmental information; guide dogs have high training costs, have certain limitations in use, and require professional maintenance.

[0003] With the development of science and technology, smart devices are becoming more and more popular, but intelligent auxiliary devices for the visually impaired are still relatively scarce. Although some mobile phone applications provide some auxiliary functions for the visually impaired, the operation of mobile phones is relatively complicated and inconvenient to use during the action. Therefore, the development of AI smart glasses with auxiliary functions for the visually impaired can provide more convenient, efficient and personalized auxiliary services for the visually impaired, which has important social significance and market value. Summary of the invention

[0004] The purpose of the present invention is to provide an AI smart glasses with visual impairment assistance to solve the problems raised in the above background technology.

[0005] In view of the above problems, the technical solution proposed by the present invention is:

[0006] A pair of AI smart glasses with visual impairment assistance, including a glasses body, a hardware system and a software system arranged on the glasses body, the hardware system including a wide-angle camera, a depth sensor, an inertial measurement unit, a geomagnetic sensor, a GPS positioning module, an AI processing chip, a microphone array, a speaker, a power module, and a storage module; the software system including an environmental perception and object recognition algorithm, a navigation algorithm, a text recognition and reading algorithm, and a voice interaction algorithm.

[0007] As a preferred technical solution of the present invention, the wide-angle camera is used to collect image data of the surrounding environment of the glasses body. The wide-angle camera is installed at the front end of the frame of the glasses body. The depth sensor is used to measure the distance between an object and the glasses body. The depth sensor is installed at the front end of the frame of the glasses body and close to the wide-angle camera. The inertial measurement unit is used to obtain the attitude information of the glasses body. The geomagnetic sensor is used to locate the direction of the glasses body. The GPS positioning module is used to locate the position of the glasses body. The AI processing chip is used to process data in real time. The inertial measurement unit, the geomagnetic sensor, the GPS positioning module, and the AI processing chip are installed inside the frame of the glasses body. The microphone array is used to receive voice commands from the user. The speaker is used to output voice feedback information. The microphone array and the speaker are installed on the temple of the glasses body. The power module is used for power supply. The storage module is used to store offline map data and software systems. The power module and the storage module are installed inside the temple of the glasses body.

[0008] As a preferred technical solution of the present invention, the depth sensor is one of a lidar or a binocular stereo vision sensor. The ranging range of the depth sensor is 0.1 - 10 meters, and the accuracy is ±5 centimeters. The microphone array is a beamforming microphone.

[0009] As a preferred technical solution of the present invention, the environmental perception and object recognition algorithm is based on convolutional neural network and semantic segmentation technology to realize real-time detection of the category, position, and scene semantic information of an object. The specific implementation steps are as follows:

[0010] S1. Image preprocessing: Perform preprocessing operations of scaling and normalization on the image collected by the wide-angle camera. The scaling formula is:

[0011] I resized (x,y) = I(sx,sy)

[0012] where (x,y) are the pixel coordinates of the scaled image, and (sx,sy) are the corresponding pixel coordinates of the original image.

[0013] The normalization operation is to map the image pixel values to the interval [0,1] or [-1,1]. The formula is:

[0014]

[0015] where μ is the mean of the image pixel values, and σ is the standard deviation;

[0016] S2. Object Detection: Input the preprocessed image into an object detection model based on a convolutional neural network to output the category, bounding box, and confidence level of the objects in the image. The bounding box is the position information of the objects;

[0017] S3. Semantic Segmentation: Input the image into a semantic segmentation model to output the category label of each pixel and obtain the scene semantic information.

[0018] As a preferred technical solution of the present invention, the navigation algorithm combines the A* algorithm, the extended Kalman filter algorithm, and the map construction technology to achieve path planning and positioning. The specific implementation steps are as follows:

[0019] S4. Pose Prediction: Use the extended Kalman filter algorithm and combine the monitoring data of the inertial measurement unit, geomagnetic sensor, and GPS positioning module, and cooperate with the offline map in the storage module to determine the position and orientation of the user. The prediction step formula of the extended Kalman filter algorithm is:

[0020]

[0021] where, is the predicted pose, F k is the state transition matrix, is the pose estimated at the previous moment, B k is the control input matrix, u k is the control input, is the predicted covariance matrix, Q k is the process noise covariance matrix, and k is the time.

[0022] The update step formula is:

[0023]

[0024] where, K k is the Kalman gain, H k is the observation matrix, z k is the observation value, R k is the observation noise covariance matrix, and L is the identity matrix;

[0025] S5. Real-time Map Construction: Use the GPS positioning module and the map construction technology, and combine the data of the inertial measurement unit, geomagnetic sensor, and depth sensor to construct the map around the user in real time. Assume the map is M, the environmental information obtained by the sensor is z, and the pose of the user is b. Then the map construction process can be expressed as:

[0026] M = f(b, z)

[0027] where, f is the map construction function;

[0028] S6. On the map constructed in step S5, use the A* algorithm to find the optimal path from the starting point to the ending point through a cost function. The cost function formula is:

[0029] f(n) = g(n) + h(n)

[0030] Each time the A algorithm selects the node with the smallest f(n) value for expansion until the ending point is found or it is determined that there is no path.

[0031] As a preferred technical solution of the present invention, the text recognition and reading algorithm converts the text and navigation information in the image into speech through a text recognition model and a natural language generation algorithm. The specific implementation steps are as follows:

[0032] S7. Text recognition: After preprocessing the image, input it into the text recognition model. The text recognition model outputs the characters in the image. Among them, the text recognition model consists of a convolutional layer, a recurrent layer, and a fully connected layer. The convolutional layer of the text recognition model extracts the image features, the recurrent layer processes the feature sequence, and finally the character prediction result is output through the fully connected layer and the Softmax function. Let the original image be I, the feature map after being processed by the convolutional layer be E, the feature sequence output by the recurrent layer be Y, and the predicted character sequence be C. Then:

[0033] F = CNN(I)

[0034] Y = LSTM(E)

[0035] C = Softmax(FC(Y))

[0036] Among them, CNN represents the convolutional neural network operation, LSTM represents the long short-term memory network operation, and FC represents the fully connected layer operation;

[0037] S8. Path parsing: The natural language generation algorithm parses the optimal path into the natural speech text of the optimal path;

[0038] S9. Speech synthesis: Input the characters in the image and the natural speech text of the optimal path into the natural language generation algorithm. The natural language generation algorithm outputs the Mel spectrogram of the speech, and then converts the Mel spectrogram into a speech waveform through a vocoder, and finally plays it through a speaker. Let the characters in the image and the natural speech text of the optimal path be R, the Mel spectrogram be O, and the speech waveform be W. Then:

[0039] O = Tacotron(R)

[0040] W = Griffin-Lim(O)

[0041] When characters in an image and the optimal path natural speech text need to be processed simultaneously, the characters in the image or the optimal path natural speech text are preferentially processed according to the set priority.

[0042] As a preferred technical solution of the present invention, the speech interaction algorithm is based on a hidden Markov model, a deep neural network speech recognition model, and a pre-trained language model based on the Transformer architecture to implement speech recognition and natural language processing. The specific implementation steps are as follows:

[0043] S10. Speech recognition: The speech signal collected by the microphone array is preprocessed and then input into a speech recognition model based on a hidden Markov model and a deep neural network. The speech signal is finally output in the form of text. The hidden Markov model is used to model the temporal structure of speech, and the deep neural network is used to extract speech features and perform classification. Let the speech signal be V, the preprocessed feature vector sequence be X, and the recognized text be J. Then:

[0044] X = Preprocess(V)

[0045] J = HMM-DNN(X)

[0046] Among them, Preprocess represents the speech preprocessing operation, and HMM-DNN represents the speech recognition operation combining a hidden Markov model and a deep neural network;

[0047] S11. Natural language processing: The converted text is input into a pre-trained language model based on the Transformer architecture. After being processed by multiple Transformer blocks, the semantics of the text are obtained, and then semantic understanding is performed through a fully connected layer. Let the input text be R, and the semantic representation after being processed by the pre-trained language model be G. Then:

[0048] G = BERT(R)

[0049] Among them, BERT represents the operation of the pre-trained language model based on the Transformer architecture, and FC represents the operation of the fully connected layer;

[0050] S12. Function module call: The AI processing chip calls the hardware system and the software system according to the semantics of the text to implement the specified path planning.

[0051] Compared with the prior art, the beneficial effects of the present invention are:

[0052] First, through accurate environmental perception and obstacle detection, it helps visually impaired users to timely discover and avoid dangers, reducing the travel risk;

[0053] Second, quickly recognize and read text information so that visually impaired users can obtain various types of written information like normal people, broadening their channels for acquiring knowledge;

[0054] Third, the real-time navigation function and natural voice interaction method enable visually impaired users to carry out daily activities more independently and improve their quality of life. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 FIG. is a schematic three-dimensional structure diagram of an AI smart glasses with visually impaired assistance disclosed in an embodiment of the present invention;

[0056] Figure 2 FIG. is a system block diagram of an AI smart glasses with visually impaired assistance disclosed in an embodiment of the present invention.

[0057] In the figure: 100, glasses body; 200, wide-angle camera; 300, depth sensor; 400, microphone array; 500, speaker. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] Please refer to Figure 1 - Figure 2 , the present invention provides a technical solution: an AI smart glasses with visually impaired assistance, including a glasses body 100, a hardware system and a software system provided on the glasses body 100. The hardware system includes a wide-angle camera 200, a depth sensor 300, an inertial measurement unit, a geomagnetic sensor, a GPS positioning module, an AI processing chip, a microphone array 400, a speaker 500, a power module, and a storage module. The software system includes an environment perception and object recognition algorithm, a navigation algorithm, a text recognition and reading algorithm, and a voice interaction algorithm;

[0060] As an embodiment of the present invention, further, the wide-angle camera 200 is used to collect image data of the surrounding environment of the glasses body 100. The wide-angle camera 200 is installed at the front end of the frame of the glasses body 100. The depth sensor 300 is used to measure the distance between an object and the glasses body 100. The depth sensor 300 is installed at the front end of the frame of the glasses body 100 and is close to the wide-angle camera 200. The inertial measurement unit is used to obtain the attitude information of the glasses body 100. The geomagnetic sensor is used to locate the direction of the glasses body 100. The GPS positioning module is used to locate the position of the glasses body 100. The AI processing chip is used to process data in real time. The inertial measurement unit, the geomagnetic sensor, the GPS positioning module, and the AI processing chip are installed inside the frame of the glasses body 100. The microphone array is used to receive voice commands from the user. The speaker is used to output voice feedback information. The microphone array 400 and the speaker 500 are installed on the temple of the glasses body 100. The power module is used for power supply, and the power module has the function of wireless charging. The storage module is used to store offline map data and software systems. The power module and the storage module are installed inside the temple of the glasses body 100.

[0061] As an embodiment of the present invention, further, the depth sensor 300 is one of a lidar or a binocular stereo vision sensor. The ranging range of the depth sensor 300 is 0.1 - 10 meters, and the accuracy is ±5 centimeters. The microphone array 400 is a beamforming microphone.

[0062] As an embodiment of the present invention, further, the environmental perception and object recognition algorithm is based on convolutional neural network and semantic segmentation technology to realize real-time detection of the category, position, and scene semantic information of an object. The specific implementation steps are as follows:

[0063] S1. Image preprocessing: Perform preprocessing operations of scaling and normalization on the image collected by the wide-angle camera 200. The scaling formula is:

[0064] I resized (x,y) = I(sx,sy)

[0065] where (x,y) is the pixel coordinate of the scaled image, and (sx,sy) is the pixel coordinate corresponding to the original image.

[0066] The normalization operation is to map the image pixel values to the interval [0,1] or [-1,1]. The formula is:

[0067]

[0068] Among them, μ is the mean of the image pixel values, and σ is the standard deviation;

[0069] S2. Object detection: Input the preprocessed image into an object detection model based on a convolutional neural network, and output the category, bounding box, and confidence of the objects in the image. The bounding box is the position information of the objects;

[0070] S3. Semantic segmentation: Input the image into a semantic segmentation model, and output the category label of each pixel to obtain scene semantic information.

[0071] As an embodiment of the present invention, further, the navigation algorithm combines the A* algorithm, the extended Kalman filter algorithm, and map construction technology to achieve path planning and positioning. The specific implementation steps are as follows:

[0072] S4. Pose prediction: Use the extended Kalman filter algorithm and combine the monitoring data of the inertial measurement unit, geomagnetic sensor, and GPS positioning module, and cooperate with the offline map in the storage module to determine the position and orientation of the user. The prediction step formula of the extended Kalman filter algorithm is:

[0073]

[0074] Among them, is the predicted pose, F k is the state transition matrix, is the pose estimated at the previous moment, B k is the control input matrix, u k is the control input, is the predicted covariance matrix, Q k is the process noise covariance matrix, and k is the time.

[0075] The update step formula is:

[0076]

[0077] Among them, K k is the Kalman gain, H k is the observation matrix, z k is the observation value, R k is the observation noise covariance matrix, and L is the identity matrix;

[0078] S5. Real-time map construction: Use the GPS positioning module and map construction technology, and combine the data of the inertial measurement unit, geomagnetic sensor, and depth sensor to construct a map around the user in real time. Assuming the map is M, the environmental information obtained by the sensor is z, and the pose of the user is b, then the map construction process can be expressed as:

[0079] M = f(b, z)

[0080] Among them, f is the map construction function;

[0081] S6. On the map constructed in step S5, use the A* algorithm to find the optimal path from the starting point to the ending point through the cost function. The cost function formula is:

[0082] f(n) = g(n) + h(n)

[0083] The A algorithm expands by selecting the node with the smallest f(n) value each time until the ending point is found or it is determined that there is no path.

[0084] As an embodiment of the present invention, further, the text recognition and reading algorithm converts the text and navigation information in the image into speech through a text recognition model and a natural language generation algorithm. The specific implementation steps are as follows:

[0085] S7. Text recognition: After preprocessing the image, input it into the text recognition model. The text recognition model outputs the characters in the image. Among them, the text recognition model is composed of a convolutional layer, a recurrent layer, and a fully connected layer. The convolutional layer of the text recognition model extracts image features, the recurrent layer processes the feature sequence, and finally, the fully connected layer and the Softmax function output the character prediction result. Let the original image be I, the feature map after convolutional layer processing be E, the feature sequence output by the recurrent layer be Y, and the predicted character sequence be C. Then:

[0086] F = CNN(I)

[0087] Y = LSTM(E)

[0088] C = Softmax(FC(Y))

[0089] Among them, CNN represents the convolutional neural network operation, LSTM represents the long short-term memory network operation, and FC represents the fully connected layer operation;

[0090] S8. Path parsing: The natural language generation algorithm parses the optimal path into the natural speech text of the optimal path;

[0091] S9. Speech synthesis: Input the characters in the image and the natural speech text of the optimal path into the natural language generation algorithm. The natural language generation algorithm outputs the mel spectrogram of the speech, and then converts the mel spectrogram into a speech waveform through a vocoder. Finally, it is played through the speaker 500. Let the characters in the image and the natural speech text of the optimal path be R, the mel spectrogram be O, and the speech waveform be W. Then:

[0092] O = Tacotron(R)

[0093] W = Griffin-Lim(O)

[0094] When characters in an image, the optimal path natural speech text need to be processed simultaneously, the characters in the image or the optimal path natural speech text are preferentially processed according to the set priority.

[0095] As an embodiment of the present invention, further, the voice interaction algorithm is based on a hidden Markov model, a deep neural network speech recognition model, and a pre-trained language model based on the Transformer architecture to implement speech recognition and natural language processing. The specific implementation steps are as follows:

[0096] S10. Speech recognition: The speech signal collected by the microphone array 400 is preprocessed and then input into a speech recognition model based on a hidden Markov model and a deep neural network. The speech signal is finally output in the form of text. The hidden Markov model is used to model the temporal structure of speech, and the deep neural network is used to extract speech features and classify. Let the speech signal be V, the preprocessed feature vector sequence be X, and the recognized text be J. Then:

[0097] X = Preprocess(V)

[0098] J = HMM-DNN(X)

[0099] Among them, Preprocess represents the speech preprocessing operation, and HMM-DNN represents the speech recognition operation combining a hidden Markov model and a deep neural network;

[0100] S11. Natural language processing: The converted text is input into a pre-trained language model based on the Transformer architecture. After being processed by multiple Transformer blocks, the semantics of the text are obtained, and then semantic understanding is performed through a fully connected layer. Let the input text be R, and the semantic representation after being processed by the pre-trained language model be G. Then:

[0101] G = BERT(R)

[0102] Among them, BERT represents the operation of the pre-trained language model based on the Transformer architecture, and FC represents the operation of the fully connected layer;

[0103] S12. Function module call: The AI processing chip calls the hardware system and the software system according to the semantics of the text to implement the specified path planning.

[0104] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0105] Please refer to Figure 1 - Figure 2 , the present invention provides a technical solution: an AI smart glasses with visual impairment assistance, including the following steps:

[0106] Initialization and preparation stage: The visually impaired person turns on the AI smart glasses, and the power module supplies power to the entire system. The GPS positioning module, inertial measurement unit, and geomagnetic sensor of the glasses start to work, and through the extended Kalman filter algorithm combined with the offline map data, initialization positioning and pose estimation are carried out to determine the user's initial position and orientation. At the same time, the microphone array and speaker perform self-checks to ensure the normal operation of the voice interaction function, and the storage module loads the offline map data and software system to prepare for subsequent use.

[0107] Environmental perception and object recognition: When the visually impaired person walks on the street, the wide-angle camera continuously collects image data of the surrounding environment, and the depth sensor synchronously measures the distance between the object and the glasses body. The collected image data is first preprocessed by scaling and normalization. The preprocessed images are respectively input into the object detection model and semantic segmentation model based on the convolutional neural network. The object detection model outputs the category, bounding box, and confidence of the object. For example, if it is recognized that there is a "power pole" in front, the bounding box determines its position; the semantic segmentation model outputs the category label of each pixel to obtain the scene semantic information. For example, it is judged that the current scene is a "street" and there are "buildings" by the roadside, etc. These information are analyzed and processed by the AI processing chip. If an obstacle is detected in front, a voice reminder will be issued in time through the speaker, such as "There is a power pole 2 meters ahead, please be careful to avoid it."

[0108] Navigation function application: A visually impaired person intends to go to a nearby library. By using the voice interaction function, they say "I want to go to the library". The microphone array collects the voice command, which is preprocessed and then input into a speech recognition model based on the Hidden Markov Model and deep neural network to be converted into text. Then, it undergoes natural language processing through a pre-trained language model with a Transformer architecture. After the AI processing chip understands the semantics, it calls the navigation algorithm. The Extended Kalman Filter algorithm combines the data from the inertial measurement unit, geomagnetic sensor, and GPS positioning module to continuously update the user's position and orientation. Meanwhile, using the GPS positioning module and map construction technology, and combining the data from the inertial measurement unit, geomagnetic sensor, and depth sensor to construct the surrounding map in real-time. On the constructed map, the A* algorithm is used to find the optimal path from the current position to the library through a cost function. After determining the optimal path, the natural language generation algorithm parses it into natural speech text, such as "Go straight ahead for 50 meters, turn right at the intersection, and then walk 80 meters to reach the library". Then, the speech synthesis module processes these texts and the text information in the possible images according to the set priority, outputs the Mel spectrogram of the speech, and then converts it into a speech waveform through a vocoder and plays it through the speaker to guide the user to reach the library smoothly.

[0109] Text recognition and information acquisition: On the way, the visually impaired person encounters a sign that needs to be viewed. The wide-angle camera captures the image of the sign, which is processed by the text recognition model. The text recognition model extracts image features through the convolutional layer, processes the feature sequence through the recurrent layer, and finally outputs the character prediction result through the fully connected layer and the Softmax function, that is, the text on the sign is recognized. This text information, together with the navigation information, is converted into speech through the natural language generation algorithm and the speech synthesis module and read to the user through the speaker to help the user obtain information.

Claims

1. An AI smart glasses with visual impairment assistance, characterized in that, It includes a glasses body (100), a hardware system and a software system provided on the glasses body (100); The hardware system includes a wide-angle camera (200), a depth sensor (300), an inertial measurement unit, a geomagnetic sensor, a GPS positioning module, an AI processing chip, a microphone array (400), a speaker (500), a power module, and a storage module; the wide-angle camera (200) is used to collect image data of the surrounding environment of the glasses body (100), the depth sensor (300) is used to measure the distance between an object and the glasses body (100), the inertial measurement unit is used to obtain the attitude information of the glasses body (100), the geomagnetic sensor is used to locate the direction of the glasses body (100), the GPS positioning module is used to locate the position of the glasses body (100), the AI processing chip is used to process data in real time, the microphone array is used to receive voice commands from the user, the speaker is used to output voice feedback information, the power module is used to supply power, and the storage module is used to store offline map data and the software system; The software system includes an environment perception and object recognition algorithm, a navigation algorithm, a text recognition and reading algorithm, and a voice interaction algorithm. The environment perception and object recognition algorithm is used to detect the category, position, and scene semantic information of an object in real time. The navigation algorithm is used to achieve path planning and positioning. The text recognition and reading algorithm is used to convert the text and navigation information in the image into voice. The voice interaction algorithm is used to achieve speech recognition and natural language processing.

2. The AI intelligent glasses with visual impairment assistance according to claim 1, characterized in that, The wide-angle camera (200) is installed at the front end of the frame of the glasses body (100), the depth sensor (300) is installed at the front end of the frame of the glasses body (100) and is close to the wide-angle camera (200). The inertial measurement unit, the geomagnetic sensor, the GPS positioning module, and the AI processing chip are installed inside the frame of the glasses body (100). The microphone array (400) and the speaker (500) are installed on the temple of the glasses body (100). The power module and the storage module are installed inside the temple of the glasses body (100).

3. The AI smart glasses with visual impairment assistance according to claim 1, characterized in that, The depth sensor (300) is one of a lidar or a binocular stereo vision sensor. The ranging range of the depth sensor (300) is 0.1 - 10 meters, and the accuracy is ±5 centimeters.

4. The AI intelligent glasses with visual impairment assistance according to claim 1, characterized in that, The microphone array (400) is a beamforming microphone.

5. The AI intelligent glasses with visual impairment assistance according to claim 1, characterized in that, The environment perception and object recognition algorithm is based on a convolutional neural network and semantic segmentation technology to achieve real-time detection of the category, position, and scene semantic information of an object. The specific implementation steps are as follows: S1. Image preprocessing: Perform preprocessing operations of scaling and normalization on the image collected by the wide-angle camera (200); S2. Object detection: Input the preprocessed image into an object detection model based on a convolutional neural network, and output the category, bounding box, and confidence of the object in the image. The bounding box is the position information of the object; S3. Semantic segmentation: Input the image into the semantic segmentation model to output the class label of each pixel and obtain the scene semantic information.

6. The AI intelligent glasses with visual impairment assistance according to claim 5, wherein, The navigation algorithm combines the A* algorithm, the extended Kalman filter algorithm, and the map construction technology to achieve path planning and positioning. The specific implementation steps are as follows: S4. Pose estimation: Use the extended Kalman filter algorithm and combine the monitoring data of the inertial measurement unit, the geomagnetic sensor, and the GPS positioning module, and cooperate with the offline map in the storage module to determine the position and orientation of the user. S5. Real-time map construction: Use the GPS positioning module and the map construction technology, and combine the data of the inertial measurement unit, the geomagnetic sensor, and the depth sensor to construct the map around the user in real time. S6. On the map constructed in step S5, use the A* algorithm to find the optimal path from the starting point to the ending point through the cost function.

7. The AI smart glasses with visual impairment assistance according to claim 6, characterized in that, The text recognition and speech synthesis algorithm converts the text and navigation information in the image into speech through the text recognition model and the natural language generation algorithm. The specific implementation steps are as follows: S7. Text recognition: After preprocessing the image, input it into the text recognition model, and the text recognition model outputs the characters in the image. S8. Path parsing: The natural language generation algorithm parses the optimal path into the natural speech text of the optimal path. S9. Speech synthesis: Input the characters in the image and the natural speech text of the optimal path into the natural language generation algorithm. The natural language generation algorithm outputs the mel spectrogram of the speech, and then converts the mel spectrogram into a speech waveform through the vocoder, and finally plays it through the speaker (500).

8. The AI smart glasses with visually impaired assistance according to claim 1, wherein The speech interaction algorithm is based on the hidden Markov model, the deep neural network speech recognition model, and the pre-trained language model of the Transformer architecture to achieve speech recognition and natural language processing. The specific implementation steps are as follows: S10. Speech recognition: The speech signal collected by the microphone array (400) is preprocessed and then input into the speech recognition model based on the hidden Markov model and the deep neural network. The speech signal is finally output in the form of text. S11. Natural language processing: The converted text is input into the pre-trained language model of the Transformer architecture, and the semantics of the text are obtained through the processing of multiple Transformer blocks, and then semantic understanding is performed through the fully connected layer. S12. Function module call: The AI processing chip calls the hardware system and the software system according to the semantic understanding to achieve the specified path planning.

Citation Information

Cited By

  • Navigation device for people with visual impairments

    RU2859069C1