A Blind Voice Navigation Assistance Method with Built-in Offline AI Intelligent Model

Through the fusion of infrared-visible binocular camera and inertial measurement unit data, combined with deep learning and Bayesian model, a personalized navigation path is generated, which solves the problem of obstacle detection and positioning drift for visual navigation for blind people in rainy and snowy weather, and improves navigation accuracy and safety.

CN119984296BActive Publication Date: 2025-07-25HANGZHOU YIYI INTELLIGENT TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510480020.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-25
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Traditional blind visual navigation has a sharp decline in image quality in rainy and snowy weather, the detection rate of obstacles is high, the path planning does not conform to the steering inertia of blind people, and the visual-inertial data time is asynchronous, and the existing path planning algorithm lacks human kinematic modeling, and the positioning error is large.

Method used

An infrared-visible binocular camera is used to collect environmental image data in real time, combine the six-degree of freedom motion posture data of the inertial measurement unit, generate three-dimensional semantic maps and dynamic obstacle detection through a multi-task neural network, and build a personalized habit route database using deep reinforcement learning, calculate the optimal navigation path with Bayesian probability graph model, and provide navigation instructions through a multimodal interactive device.

Benefits of technology

Effectively detect obstacles in rainy and snowy weather, generate navigation paths that meet the walking habits of blind people, reduce positioning drift, improve navigation accuracy and safety, and enhance navigation comfort and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119984296B_ABST
    Figure CN119984296B_ABST
Patent Text Reader

Abstract

The present invention is a blind voice navigation assistance method with a built-in offline AI intelligent model, which relates to the technical field of computer processing and includes the following steps: S01. Real-time collect environmental image data through an infrared-visible light binocular camera, and simultaneously obtain six-degree-of-freedom motion attitude data of an inertial measurement unit. The infrared-visible light binocular camera includes a thermal imaging sensing channel; S02. Input the environmental image data into a pre-trained multi-task neural network model to generate a three-dimensional semantic map containing the topological structure of the blind path and a dynamic obstacle detection result; S03. Based on a deep reinforcement learning algorithm, perform spatio-temporal clustering analysis on the user's historical path to construct a personalized habit route database that integrates human body turning inertia. This invention significantly improves the perception accuracy, path planning humanization, positioning accuracy, and extreme weather adaptability of the navigation system, and ensures the safe navigation of the blind.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer processing, and particularly to a blind voice navigation assistance method with a built-in offline AI intelligent model. Background Art

[0002] With the emergence of intelligent AI models, big data processing has become faster, promoting the development of current artificial intelligence. Especially in the application of blind visual navigation, the problems of traditional blind visual navigation are as follows:

[0003] In terms of environmental perception, in traditional monocular vision solutions, the image quality drops sharply in rainy and snowy weather, and multi-spectral data cannot be effectively fused, resulting in a missed detection rate of obstacles as high as 40% in interference scenarios such as water reflection and ice crystal attachment. And in existing path planning algorithms (such as a method for determining line weights for the Dijkstra algorithm disclosed in a patent with publication number CN117091599A, which uses the Dijkstra algorithm), kinematic modeling of the human body is lacking, and the generated right-angle turning paths do not conform to the turning inertia characteristics of the blind. In actual measurements, the probability of users deviating from the path increases by 2.3 times.

[0004] In terms of sensor data fusion, the mainstream solutions (such as a positioning method and device for deep fusion of visual-inertial data disclosed in a patent with publication number CN109238277B) do not solve the problem of time asynchrony of visual-inertial data. When the frame rate of an infrared-visible light binocular camera (30fps) does not match the sampling rate of an IMU (100Hz), the motion compensation residual causes a positioning drift of up to 0.5m / min.

[0005] In summary, how to solve the above problems is expected to be well solved! Summary of the Invention

[0006] In view of the above technical problems, the technical solution adopted by the present invention is a blind voice navigation assistance method with a built-in offline AI intelligent model, including the following steps:

[0007] S01. Real-time collect environmental image data through an infrared-visible light binocular camera, and synchronously obtain six-degree-of-freedom motion attitude data of an inertial measurement unit. The infrared-visible light binocular camera includes a thermal imaging sensing channel;

[0008] S02. Input the environmental image data into a pre-trained multi-task neural network model to generate a three-dimensional semantic map including the topological structure of the blind path and the detection result of dynamic obstacles;

[0009] S03. Conduct spatio-temporal clustering analysis on the user's historical path based on a deep reinforcement learning algorithm, and construct a personalized habitual route database integrating the turning inertia of the human body;

[0010] S04. Integrate real-time perception data with an offline map database, and calculate the optimal navigation path that conforms to human kinematic constraints through a Bayesian probability graph model;

[0011] S05. Convert the result of the optimal navigation path into a natural language instruction containing a description of the spatial reference system, and implement multimodal interaction through a bone conduction headset and a tactile feedback device.

[0012] Preferably, in step S01, the motion attitude data in the synchronous acquisition of the motion attitude data of the inertial measurement unit includes six-degree-of-freedom motion attitude data corresponding to the time phase of each frame of the environmental image data, and its steps include:

[0013] S11. Establish a rigid transformation matrix from the IMU coordinate system to the infrared-visible light binocular camera coordinate system:

[0014] ;

[0015] where R is the rotation matrix and t is the translation vector;

[0016] S12. Insert synchronous timestamps into each frame of image and the IMU data stream;

[0017] S13. When it is detected that the IMU data delay exceeds 10 ms, start the Kalman filter to predict the motion state:

[0018] ;

[0019] where is the state estimate value at time k, is the state transition matrix, is the dynamic evolution law from time to time k, is the control input matrix, is the control input vector,

[0020] Preferably, in step S01, the real-time acquired environmental image data is preprocessed, including:

[0021] S14. Based on the multi-scale Retinex algorithm, separate the luminance component V(x, y) in the HSV color space, and extract the image reflection layer R(x, y) and the illumination layer L(x, y) through a Gaussian difference filter bank. The formula is as follows:

[0022] ;

[0023] where is a Gaussian kernel with a radius of {5, 15, 30} pixels, and wi = {0.3, 0.5, 0.2};

[0024] S15. Construct a weather classifier based on Swin Transformer, and activate the path curvature constraint module when hail is recognized, restricting the turning angle ≤ 45°;

[0025] S16. Based on the non-local mean filtering algorithm, construct the similarity weights of pixel blocks within a two-dimensional search window;

[0026] S17. Embed the prior model of rain and snow degradation generated by adversarial training in the U-Net decoder.

[0027] Preferably, the training of the weather classifier in step S15 includes:

[0028] S151. Construct a synthetic training dataset, use Unreal Engine to render 200 precipitation scenarios, and generate samples with different precipitation densities and wind direction angles through domain randomization;

[0029] S152. Through a dual-branch feature extraction structure, the dual-branch feature includes:

[0030] The main branch processes RGB images;

[0031] The auxiliary branch analyzes the thermal imaging temperature distribution pattern;

[0032] S153. Use light rain / snow samples in the initial stage of training, and finally increase the weight coefficients of extreme weather such as heavy rain / heavy snow, and import them into step S16 to obtain the similarity weights of pixel blocks within the two-dimensional search window to form a data set.

[0033] Preferably, the construction of the personalized habit route database in step S03 includes:

[0034] S31. Record trajectory data through the GNSS / IMU fusion positioning module, and the semi-major axis of the positioning error ellipse ≤ 0.8 meters;

[0035] S32. Adopt a spatio-temporal density clustering algorithm, and the spatio-temporal neighborhood determination conditions are:

[0036] ;

[0037] ;

[0038] where r = 1.2 × average step length, Δt = 5 minutes, (x i , y i ) are the two-dimensional coordinates of the i-th trajectory point of the user, (x j , y jThe two-dimensional coordinates of the user's j-th trajectory point, where ti and tj are the timestamps of trajectory points i and j.

[0039] Preferably, when the GNSS signal is lost during the execution of step S03 above, it includes:

[0040] S71. Dead reckoning based on IMU data:

[0041] ;

[0042] ;

[0043] Among them, = 0.2, v(t) is the linear velocity of the user at time t, a(τ) is the IMU acceleration measurement value at time τ, θ(t) is the heading angle of the user at time t, ω(τ) is the IMU gyroscope angular velocity at time τ, and δ is the deviation between the heading angle measured by the magnetometer and the integrated heading angle of the gyroscope;

[0044] S62. Correct the cumulative error through the landmark recognition result every 2 seconds, and the correction amount is calculated as:

[0045] ;

[0046] Among them, k p = 0.8, k i = 0.05, e x is the deviation between the landmark recognition position and the dead reckoning position at the current moment, and Δx is the positioning error correction amount.

[0047] Preferably, in step S04, calculating the optimal navigation path that conforms to human kinematic constraints through the Bayesian probability graph model includes:

[0048] S41. Obtain the offline map database, define the current walking section, and obtain the coordinates of obstacles around the current position;

[0049] S42. The obtained real-time perception data is distributed in the offline map database in the form of dots, and synchronously obtain the first coordinate set of the maximum swing amplitude of the left and right of the blind stick and the second coordinate set of the left and right sway of the walking body;

[0050] S43. Connect the multiple dot-shaped icons located in the offline map database, judge the trajectory of the human body walking, predict the obstacles that will be touched during future walking, extract the coordinates of the obstacles, and calculate and generate a correction instruction based on the Bayesian probability graph model.

[0051] The present invention has at least the following beneficial effects:

[0052] 1. By using an infrared-visible light binocular camera and an image enhancement algorithm, problems such as poor image quality and reflection interference in rainy and snowy weather are solved. Thus, the present invention can still "see" the environment under harsh conditions such as waterlogging, hail, and heavy rain, avoiding collisions caused by blind people missing obstacles during detection;

[0053] 2. By combining the user's historical walking habits and a human kinematics model, a path that conforms to the natural walking pattern of blind people is generated, reducing counter-intuitive instructions such as sharp turns and lowering the probability of deviating from the path. Thus, it is possible to avoid blind people frequently adjusting their directions due to unreasonable path planning, improving the comfort and safety of navigation;

[0054] 3. By using timestamp synchronization and the Kalman filtering algorithm, the positioning drift problem caused by asynchronous data between the camera and the IMU is solved. Thus, the positioning can be corrected in real time, avoiding potential dangers to blind navigation caused by positioning errors, ensuring accurate position calculation during the walking of blind people, and reducing navigation errors caused by positioning drift;

[0055] 4. The weather classifier is trained using extreme weather data synthesized by a virtual engine, the noise reduction parameters are dynamically adjusted, and the maximum turning angle in hail weather is restricted, improving the robustness of the present invention in extreme scenarios such as heavy rain and heavy snow, and avoiding navigation failure caused by sudden weather changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0057] Figure 1 It is a flowchart of a blind voice navigation assistance method with a built-in offline AI intelligent model provided by Embodiment 1 of the present invention;

[0058] Figure 2 It is an image generated by integrating real-time perception data and an offline map database provided by Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention.

[0060] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0061] Embodiment 1

[0062] This embodiment provides a blind voice navigation assistance method with a built-in offline AI intelligent model. The method includes the following steps, as Figure 1 shown:

[0063] S01. Real-time collect environmental image data through an infrared-visible light binocular camera, and synchronously obtain the six-degree-of-freedom motion attitude data of the inertial measurement unit. The infrared-visible light binocular camera includes a thermal imaging sensing channel;

[0064] Specifically, the above-mentioned real-time collection of environmental image data by the infrared-visible light binocular camera means that the infrared-visible light binocular camera is set on both sides of the blind person's eyes, and high-definition images and infrared images can be obtained as needed. The two image data are used in an overlapping manner, and the purpose is to supplement high-definition images at night and in rainy and snowy weather, so as to increase the identification of moving targets in rainy, snowy and night conditions, and thus avoid collisions between other pedestrians and animals and the user.

[0065] Furthermore, in the above step S01, the motion attitude data in the synchronously obtained motion attitude data of the inertial measurement unit includes six-degree-of-freedom motion attitude data corresponding to the time phase of each frame of the environmental image data. The steps include:

[0066] S11. Establish a rigid transformation matrix from the IMU coordinate system to the infrared-visible light binocular camera coordinate system:

[0067] ;

[0068] where R is the rotation matrix and t is the translation vector;

[0069] S12. Insert synchronous timestamps into each frame of image and the IMU data stream;

[0070] S13. When it is detected that the IMU data delay exceeds 10 ms, start the Kalman filter to predict the motion state:

[0071] ;

[0072] wherein, is the state estimation value at time k, is the state transition matrix, is the dynamic evolution law from time to time k, is the control input matrix, is the control input vector, is the noise vector.

[0073] In order to ensure that the data of the camera and the inertial measurement unit can be "matched", a coordinate system transformation matrix is established to unify the data of the two into one coordinate system. Then, a timestamp is added to each frame of image and the IMU data stream. Through the comparison of the two data by the timestamp, if the data with the same timestamp is consistent, it means that the IMU data is correct. If it is inconsistent, it means that the IMU data is delayed. It will use the Kalman filter to predict the motion state and use the value obtained by the Kalman filter as the value of the IMU data stream, thus greatly improving the accuracy and stability of the data.

[0074] Furthermore, in step S01, the real-time collected environmental image data will be preprocessed, including:

[0075] S14. Based on the multi-scale Retinex algorithm, the luminance component V(x, y) is separated in the HSV color space, and the image reflection layer R(x, y) and the illumination layer L(x, y) are extracted through a Gaussian difference filter bank. The formula is as follows:

[0076] ;

[0077] wherein, is a Gaussian kernel with a radius of {5, 15, 30} pixels, and wi = {0.3, 0.5, 0.2};

[0078] S15. Construct a weather classifier based on Swin Transformer. When hail is recognized, activate the path curvature constraint module to limit the turning angle ≤ 45°;

[0079] S16. Based on the non-local mean filtering algorithm, construct the similarity weight of pixel blocks within a two-dimensional search window;

[0080] S17. Embed the rain and snow degradation prior model generated by adversarial training in the U-Net decoder.

[0081] After the environmental image is collected, it will be preprocessed first to make it clearer and more accurate. The multi-scale Retinex algorithm is used to separate the brightness component of the image from the reflection layer and the illumination layer, so that the image is less affected by illumination. Then, a weather classifier is used to identify the weather conditions. If it is raining or snowing, it will perform special processing on the image to reduce the impact of the weather on navigation. Finally, the non-local means filtering algorithm is used to reduce the noise and blur in the image. Thus, the clarity of the image is improved, the noise and blur in the image are reduced, and at the same time, special processing is performed according to the weather conditions using the weather classifier, so as to ensure the navigation requirements in rainy and snowy weather.

[0082] S02. Input the environmental image data into a pre-trained multi-task neural network model to generate a three-dimensional semantic map containing the topological structure of the blind path and the dynamic obstacle detection result;

[0083] The above-mentioned image information captured from the environment, such as pictures of streets, indoor scenes or any other visual scenes taken by a camera. Then the image data is input into a trained neural network model. This model is "multi-task", which means it can perform multiple tasks simultaneously, rather than just focusing on a specific output (such as only identifying objects or only generating maps). Then the trained neural network model generates a three-dimensional semantic map according to the input image data. It also contains the connection information about the actual blind path and the virtual blind path. The so-called virtual blind path refers to a temporary blind path generated when it is detected that the actual blind path is occupied, and the virtual blind path disappears until it guides to the actual blind path. The above-mentioned trained neural network model belongs to the prior art and can refer to the model logic of intelligent driving of automobiles.

[0084] S03. Based on the deep reinforcement learning algorithm, perform spatio-temporal clustering analysis on the user's historical path to construct a personalized habitual route database integrating human body turning inertia;

[0085] Specifically, in step S03, constructing the personalized habitual route database includes:

[0086] S31. Record the trajectory data through the GNSS / IMU fusion positioning module, and the semi-major axis of the positioning error ellipse ≤ 0.8 meters;

[0087] S32. Adopt the spatio-temporal density clustering algorithm, and the spatio-temporal neighborhood determination condition is:

[0088] ;

[0089] ;

[0090] where r = 1.2 × average step length, Δt = 5 minutes, (x i, y i ), the two-dimensional coordinates of the i-th trajectory point of the user, (x j , y j ), the two-dimensional coordinates of the j-th trajectory point of the user, and ti, tj are the timestamps of the trajectory points i and j.

[0091] The GNSS / IMU fusion positioning module is used to accurately record trajectory data, providing a solid foundation for building a personalized route database. And through the spatio-temporal density clustering algorithm, these trajectory data are deeply analyzed to obtain the routes and preferred paths that blind people often take, so as to provide more personalized services for blind people according to these habitual routes.

[0092] Furthermore, it is executed when the GNSS signal is lost in the above step S03, including:

[0093] S71: Dead reckoning based on IMU data:

[0094] ;

[0095] ;

[0096] Among them, = 0.2, v(t) is the linear velocity of the user at time t, a(τ) is the IMU acceleration measurement value at time τ, θ(t) is the heading angle of the user at time t, ω(τ) is the IMU gyroscope angular velocity at time τ, and δ is the deviation between the heading angle measured by the magnetometer and the integrated heading angle of the gyroscope;

[0097] S62: Correct the cumulative error through the landmark recognition result every 2 seconds, and the correction amount is calculated as:

[0098] ;

[0099] Among them, k p = 08, k i = 0.05, e x is the deviation between the landmark recognition position and the dead reckoning position at the current moment, and Δx is the positioning error correction amount.

[0100] In the case of GNSS signal loss, through means such as dead reckoning based on IMU data and correction of landmark recognition results, the continuity and accuracy of navigation are maintained. When the GNSS signal is lost, the approximate position of the blind person is predicted based on the action state of dead reckoning based on IMU data. At the same time, the previous dead reckoning results are corrected by using the landmark recognition function at regular intervals to ensure the accuracy of the position. The coordinated operation of the two execution strategies ensures that the navigation system can also provide continuous and accurate navigation services when the GNSS signal is unstable.

[0101] S04. Integrate real-time perception data with the offline map database, and calculate the optimal navigation path that conforms to human kinematic constraints through a Bayesian probability graph model;

[0102] Specifically, in combination with Figure 2 As shown, the detailed implementation steps of step S04 are as follows:

[0103] S41. Obtain the offline map database, define the current walking section, and obtain the coordinates of obstacles around the current location;

[0104] S42. The obtained real-time perception data is distributed in the offline map database in the form of points, and simultaneously obtain the first coordinate set of the maximum swing amplitudes of the left and right sides of the blind stick and the second coordinate set of the left and right shakes of the walking body;

[0105] S43. Connect the multiple point-shaped icons in the offline map database, judge the walking trajectory of the human body, predict the obstacles that will be touched during future walking, extract the coordinates of the obstacles, and calculate and generate a correction instruction based on the Bayesian probability graph model.

[0106] It should be noted that the first coordinate set in step S42 above is the two endpoint coordinates of the maximum swing amplitudes of the left and right sides of the blind stick when the device is enabled, and the second coordinate set is the straight-line distance between the two sides of the image under the environmental image data obtained by the infrared-visible light binocular camera, and the endpoint coordinates at both ends of the straight line located in the current section are obtained based on this straight-line distance.

[0107] · By using the Bayesian probability graph model, various factors such as human kinematic constraints, real-time perception data, and offline map information are integrated, and the path is adaptively designed according to the behavior habits of each blind person. At the same time, it can also predict in real time the obstacles that may be encountered during future walking and give correction instructions in advance to ensure that blind people can easily cope with various complex environments. It significantly improves the accuracy and safety of navigation, enhances the humanization and comfort of navigation, and provides more considerate services for blind people.

[0108] S05. Convert the result of the optimal navigation path into natural language instructions including spatial reference system descriptions, and realize multimodal interaction through bone conduction headphones and tactile feedback devices.

[0109] Specifically, generate spatial language according to the obtained result of the optimal navigation path, that is, natural language instructions described by the spatial reference system, such as turn left and walk 100 meters, pass by that red building, there is no interaction with the red building currently, and walk in the current walking direction.

[0110] In summary, the method provided in the first embodiment solves problems such as poor image quality and reflection interference in rainy and snowy weather through an infrared-visible light binocular camera and an image enhancement algorithm. Thus, the present invention can still "see" the environment under harsh conditions such as waterlogging, hail, and heavy rain, avoiding collisions caused by blind people missing obstacles. By combining the user's historical walking habits and the human kinematics model, a path that conforms to the natural walking pattern of blind people is generated, reducing counterintuitive instructions such as sharp turns and lowering the probability of deviating from the path. Therefore, it is possible to avoid blind people frequently adjusting their directions due to unreasonable path planning, improving the comfort and safety of navigation. Secondly, through timestamp synchronization and the Kalman filter algorithm, the problem of positioning drift caused by asynchronous data between the camera and the IMU is solved, thereby correcting the positioning in real time and avoiding potential dangers to blind navigation caused by positioning errors, ensuring accurate position calculation for blind people during walking and reducing navigation errors caused by positioning drift. Moreover, the weather classifier is trained using extreme weather data synthesized by the virtual engine, dynamically adjusting the noise reduction parameters and restricting the maximum turning angle in hail weather, improving the robustness of the present invention in extreme scenarios such as heavy rain and heavy snow, and avoiding navigation failure due to sudden weather changes.

[0111] Embodiment 2

[0112] Based on the above first embodiment, this embodiment aims to train the "construction of a weather classifier based on SwinTransformer" in the above step S01, and its steps include:

[0113] S151. Construct a synthetic training data set, use Unreal Engine to render 200 precipitation scenarios, and generate samples with different precipitation densities and wind direction angles through domain randomization;

[0114] S152. Through a dual-branch feature extraction structure, the dual branches include:

[0115] The main branch processes RGB images;

[0116] The auxiliary branch analyzes the thermal imaging temperature distribution pattern;

[0117] S153. In the initial stage of training, use light rain / snow samples, and finally increase the weight coefficients of extreme weather such as heavy rain / heavy snow, and import them into step S16 respectively to calculate the similarity weights of pixel blocks within the two-dimensional search window to form a data set.

[0118] In this embodiment, a rich precipitation scene is rendered using Unreal Engine, and a large synthetic training dataset is constructed. A dual-branch feature extraction structure is adopted to process RGB images and thermal imaging temperature distribution patterns respectively, so as to extract richer image features. During the training process, it is also considered that the weights of samples need to be dynamically adjusted according to the severity of the weather to ensure that the classifier can pay more attention to those extreme weather conditions. As a result, the recognition accuracy and robustness of the meteorological classifier are significantly improved.

[0119] Embodiment 3

[0120] An embodiment of the present invention provides a non-transitory computer-readable storage medium, in which at least one instruction or at least one program segment is stored, and at least one instruction or at least one program segment is loaded and executed by a processor to implement the steps of:

[0121] Real-time acquisition of environmental image data through an infrared-visible light binocular camera, and synchronous acquisition of six-degree-of-freedom motion attitude data of an inertial measurement unit. The infrared-visible light binocular camera includes a thermal imaging sensing channel;

[0122] Input the environmental image data into a pre-trained multi-task neural network model to generate a three-dimensional semantic map containing the topological structure of the blind path and the dynamic obstacle detection result;

[0123] Based on the deep reinforcement learning algorithm, perform spatio-temporal clustering analysis on the user's historical path to construct a personalized habitual route database that integrates human body turning inertia;

[0124] Fuse real-time perception data and the offline map database, and calculate the optimal navigation path that conforms to human kinematic constraints through the Bayesian probability map model;

[0125] Convert the result of the optimal navigation path into a natural language instruction containing a spatial reference system description, and implement multi-modal interaction through a bone conduction headset and a tactile feedback device.

[0126] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0127] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0128] Embodiment 4

[0129] An embodiment of the present invention provides an electronic device, including a processor and a memory. At least one instruction or at least one program segment is stored in the memory. The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the steps:

[0130] Real-time collect environmental image data through an infrared-visible light binocular camera, and synchronously obtain six-degree-of-freedom motion attitude data of an inertial measurement unit. The infrared-visible light binocular camera includes a thermal imaging sensing channel;

[0131] Input the environmental image data into a pre-trained multi-task neural network model to generate a three-dimensional semantic map including the topological structure of the blind path and the dynamic obstacle detection result;

[0132] Based on a deep reinforcement learning algorithm, perform spatio-temporal clustering analysis on the user's historical path to construct a personalized habit route database that integrates human body turning inertia;

[0133] Fuse real-time perception data with an offline map database, and calculate the optimal navigation path that conforms to human kinematic constraints through a Bayesian probability graph model;

[0134] Convert the result of the optimal navigation path into a natural language instruction including a description of the spatial reference system, and implement multimodal interaction through a bone conduction headset and a tactile feedback device.

[0135] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments with equivalent changes within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A blind voice navigation assistance method with a built-in offline AI intelligent model, characterized in that, It includes the following steps: S01. Real-time collect environmental image data through an infrared-visible light binocular camera, and synchronously obtain six-degree-of-freedom motion attitude data of an inertial measurement unit. The infrared-visible light binocular camera includes a thermal imaging sensing channel; S02. Input the environmental image data into a pre-trained multi-task neural network model to generate a three-dimensional semantic map including a blind path topological structure and a dynamic obstacle detection result; S03. Based on a deep reinforcement learning algorithm, perform spatio-temporal clustering analysis on the user's historical path to construct a personalized habitual route database integrating human body turning inertia; S04. Integrate real-time perception data and an offline map database, and calculate an optimal navigation path that conforms to human kinematic constraints through a Bayesian probability graph model; S05. Convert the result of the optimal navigation path into a natural language instruction including a spatial reference system description, and implement multi-modal interaction through a bone conduction headset and a tactile feedback device; In step S01, the real-time collected environmental image data will be pre-processed, including: S14. Based on the multi-scale Retinex algorithm, separate the luminance component V(x, y) in the HSV color space, and extract the image reflection layer R(x, y) and the illumination layer L(x, y) through a Gaussian difference filter bank. The formula is as follows: ; Among them, is a Gaussian kernel with a radius of {5, 15, 30} pixels, and wi = {0.3, 0.5, 0.2}; S15. Construct a weather classifier based on SwinTransformer. When hail is recognized, activate the path curvature constraint module to limit the turning angle ≤ 45°; S16. Based on the non-local mean filtering algorithm, construct the similarity weight of pixel blocks within a two-dimensional search window; S17. Embed a rain and snow degradation prior model generated by adversarial training in the U-Net decoder; In step S04, the calculation of the optimal navigation path that conforms to human kinematic constraints through the Bayesian probability graph model includes: S41. Obtain an offline map database, define the current walking section, and obtain the coordinates of obstacles around the current location; S42. The obtained real-time perception data is located in the offline map database in a dot-like manner, and synchronously obtain the first coordinate set of the maximum swing amplitude of the left and right of the blind stick and the second coordinate set of the left and right swaying of the walking body; S43. Connect multiple dot-like icons in the offline map database, judge the trajectory of the human body walking, predict the obstacles that will be touched during future walking, extract the coordinates of the obstacles, and calculate and generate a correction instruction based on the Bayesian probability graph model.

2. The blind voice navigation assistance method with a built-in offline AI intelligent model according to claim 1, characterized in that, In step S01, the motion attitude data in the synchronous acquisition of the motion attitude data of the inertial measurement unit includes six-degree-of-freedom motion attitude data corresponding to the time phase of each frame of the environmental image data. The steps include: S11. Establish a rigid transformation matrix from the IMU coordinate system to the infrared-visible light binocular camera coordinate system: ; where R is the rotation matrix and t is the translation vector; S12. Insert synchronous timestamps for each frame of image and the IMU data stream; S13. When it is detected that the IMU data delay exceeds 10 ms, start the Kalman filter to predict the motion state; ; wherein, is the state estimate value at time k, is the state transition matrix, which represents the dynamic evolution law from time to time k, is the control input matrix, is the control input vector, is the noise vector.

3. The blind voice navigation assistance method with a built-in offline AI intelligent model according to claim 1, characterized in that, The training of the weather classifier in step S15 includes: S151. Construct a synthetic training dataset, render 200 precipitation scenarios using Unreal Engine, and generate samples with different precipitation densities and wind direction angles through domain randomization; S152. Through a dual-branch feature extraction structure, the dual-branch feature includes: The main branch processes RGB images; The auxiliary branch analyzes the thermal imaging temperature distribution pattern; S153. In the initial stage of training, use light rain / snow samples, and finally increase the weight coefficients of extreme weather such as heavy rain / blizzard, and import them into step S16 to obtain the similarity weights of pixel blocks within the two-dimensional search window respectively to form a data set.

4. The blind voice navigation assistance method with a built-in offline AI intelligent model according to claim 1, characterized in that In step S03, constructing the personalized habit route database includes: S31. Record trajectory data through the GNSS / IMU fusion positioning module, and the semi-major axis of the positioning error ellipse ≤ 0.8 meters; S32. Adopt a spatio-temporal density clustering algorithm, and the spatio-temporal neighborhood determination condition is: ; ; Among them, r = 1.2 × average step length, Δt = 5 minutes, (x i , y i ) is the two-dimensional coordinate of the i-th trajectory point of the user, (x j , y j ) is the two-dimensional coordinate of the j-th trajectory point of the user, t i , t j are the timestamps of trajectory points i and j.

5. The blind voice navigation assistance method with a built-in offline AI intelligent model according to claim 4, characterized in that, When the GNSS signal is lost during the execution of step S03 above, it includes: S71. Dead reckoning based on IMU data: ; ; wherein, = 0.2, v(t) is the linear velocity of the user at time t, a(τ) is the IMU acceleration measurement value at time τ, θ(t) is the heading angle of the user at time t, ω(τ) is the IMU gyroscope angular velocity at time τ, and δ is the deviation between the heading angle measured by the magnetometer and the integrated heading angle of the gyroscope; S72. Correct the cumulative error every 2 seconds through the landmark recognition result, and the correction amount is calculated as: ; where k p = 0.8, k i = 0.05, e x is the deviation between the recognition position and the estimated position of the landmark at the current moment, and Δx is the positioning error correction amount.

6. The blind voice navigation assistance method with a built-in offline AI intelligent model according to claim 1, characterized in that, The first coordinate set in step S42 is the coordinates of the two endpoints of the maximum swing amplitude of the blind stick left and right when the device is enabled, and the second coordinate set is the straight-line distance between the two sides of the image under the environmental image data obtained by the infrared-visible light binocular camera, and the endpoint coordinates at both ends of the straight line located in the current section are obtained based on this straight-line distance.

7. A non-transitory computer-readable storage medium storing at least one instruction or at least one program segment therein, characterized in that, The at least one instruction or the at least one program is loaded and executed by a processor to implement the steps of the blind voice navigation assistance method with the built-in offline AI intelligent model as described in any one of claims 1-6.

8. An electronic device, characterized in that, It includes a processor and a memory, and at least one instruction or at least one program is stored in the memory, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the steps of the blind voice navigation assistance method with the built-in offline AI intelligent model as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Localization method and device based on deep fusion of visual and inertial data

    CN109238277B

  • Line weight determination method for Dijkstra algorithm

    CN117091599A

  • Method and device for planning flight path of unmanned aerial vehicle under wind and rain conditions

    CN117029827A

  • Pedestrian navigation system and path planning method based on environmental perception and human kinematics

    CN117213513A