Blind person voice navigation auxiliary method with built-in off-line AI intelligent model

Through the blind voice navigation assistance method with built-in offline AI smart model, combined with the data of infrared-visible binocular camera and inertial measurement unit, a three-dimensional semantic map and personalized navigation path is generated, which solves the problem of navigation failure in traditional blind navigation in extreme weather, and achieves high accuracy and safe navigation under harsh conditions.

CN119984296AActive Publication Date: 2025-05-13HANGZHOU YIYI INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510480020.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Traditional blind voice navigation has a decline in image quality in rainy and snowy weather, high detection rate of obstacles, unreasonable path planning, resulting in an increase in the probability of deviation from the path, and serious positioning drift problems, especially in extreme weather navigation failure.

Method used

The blind voice navigation assistance method is adopted with a built-in offline AI smart model. Environmental data is collected through infrared-visible binocular cameras, combined with the motion posture data of the inertial measurement unit, and a multi-task neural network is used to generate three-dimensional semantic maps and dynamic obstacle detection results. A personalized habit route database is constructed based on the deep reinforcement learning algorithm, the optimal navigation path is calculated through the Bayesian probability graph model, and multimodal interaction is achieved through bone conduction headphones and tactile feedback device.

Benefits of technology

Under harsh conditions such as water accumulation, hail, and heavy rain, we can still effectively "see clearly" the environment, avoid collisions of blind people due to missed detection obstacles, generate paths that conform to the natural walking mode of blind people, reduce the probability of deviating from the path, ensure the accuracy and safety of blind people's navigation, and maintain the robustness of navigation in extreme weather.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119984296A_ABST
    Figure CN119984296A_ABST
Patent Text Reader

Abstract

The invention provides a blind person voice navigation auxiliary method with a built-in offline AI intelligent model, and relates to the technical field of computer processing, and the blind person voice navigation auxiliary method comprises the following steps: S01, collecting environment image data in real time through an infrared-visible light binocular camera, synchronously obtaining six-degree-of-freedom motion attitude data of an inertial measurement unit, the binocular infrared-visible light binocular camera comprises a thermal imaging sensing channel; s02, inputting the environment image data into a pre-trained multi-task neural network model, and generating a three-dimensional semantic map containing a blind sidewalk topological structure and a dynamic obstacle detection result; and S03, based on a deep reinforcement learning algorithm, carrying out space-time clustering analysis on the historical paths of the user, and constructing a personalized habit route database fused with human body turning inertia. The sensing precision, the path planning humanization, the positioning accuracy and the extreme weather adaptability of the navigation system in a complex environment are remarkably improved, and safe navigation of the blind is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer processing technology, and in particular to a voice navigation assistance method for the blind with a built-in offline AI intelligent model. Background Art

[0002] With the emergence of intelligent AI models, big data processing has become faster, which has promoted the development of artificial intelligence. Especially the application of visual navigation for the blind. The problems of traditional visual navigation for the blind are: In terms of environmental perception, the image quality of traditional monocular vision solutions drops sharply in rainy and snowy weather, and fails to effectively integrate multispectral data, resulting in a 40% missed detection rate of obstacles in interference scenes such as water reflection and ice crystal attachment. In addition, the existing path planning algorithms (such as the patent, publication number: CN117091599A, a method for determining the route weight for the Dijkstra algorithm, which uses the Dijkstra algorithm) lack human kinematic modeling, and the generated right-angle turning path does not conform to the steering inertia characteristics of the blind. In actual measurements, the probability of users deviating from the path increases by 2.3 times.

[0003] In terms of sensor data fusion, mainstream solutions (such as patents, published CN109238277B, a positioning method and device for deep fusion of visual-inertial data) do not solve the problem of temporal asynchrony of visual-inertial data. When the frame rate of the infrared-visible light binocular camera (30fps) does not match the IMU sampling rate (100Hz), the motion compensation residual causes positioning drift of up to 0.5m / min.

[0004] To sum up, how to solve the above problems is expected to be well solved. Summary of the invention

[0005] In view of the above technical problems, the technical solution adopted by the present invention is a voice navigation assistance method for the blind with a built-in offline AI intelligent model, comprising the following steps: S01, collecting environmental image data in real time through an infrared-visible light binocular camera, and synchronously acquiring six-degree-of-freedom motion posture data of an inertial measurement unit, wherein the binocular infrared-visible light binocular camera includes a thermal imaging sensor channel; S02, inputting the environmental image data into a pre-trained multi-task neural network model to generate a three-dimensional semantic map including a topological structure of a blind path and a dynamic obstacle detection result; S03. Based on the deep reinforcement learning algorithm, the user's historical paths are analyzed in time and space to build a personalized habit route database that integrates human steering inertia; S04, integrating real-time perception data with offline map database, and calculating the optimal navigation path that meets human kinematic constraints through the Bayesian probability graph model; S05. Convert the result of the optimal navigation path into a natural language instruction containing a description of the spatial reference system, and realize multimodal interaction through bone conduction headphones and a tactile feedback device.

[0006] Preferably, the motion posture data in the motion posture data of the inertial measurement unit synchronously acquired in step S01 includes six-degree-of-freedom motion posture data corresponding to the time phase of each frame of the environmental image data, and the steps include: S11. Establish the rigid transformation matrix from the IMU coordinate system to the infrared-visible light binocular camera coordinate system: ; Among them, R is the rotation matrix and t is the translation vector; S12, inserting a synchronization timestamp into each frame image and IMU data stream; S13. When it is detected that the IMU data delay exceeds 10ms, the Kalman filter is started to predict the motion state: ; in, is the estimated value of the state at time k, is the state transfer matrix, The dynamic evolution law from time to time k, is the control input matrix, is the control input vector, is the noise vector.

[0007] Preferably, in step S01, the real-time collected environment image data is preprocessed, including: S14. Based on the multi-scale Retinex algorithm, the brightness component V(x, y) is separated in the HSV color space, and the image reflection layer R(x, y) and the illumination layer L(x, y) are extracted through the Gaussian difference filter group. The formula is as follows: ; Among them, F i is a Gaussian kernel with radius {5, 15, 30} pixels, wi = {0.3, 0.5, 0.2}; S15. Build a meteorological classifier based on SwinTransformer. When hail is identified, activate the path curvature constraint module to limit the turning angle to ≤45°. S16. Based on the non-local mean filtering algorithm, a similarity weight of pixel blocks in the two-dimensional search window is constructed, and the formula is as follows: w(p,q)=exp(-||N(p)-N(q)||² / (2σ²))·G σs (||pq||); Among them, σ is dynamically adjusted according to the precipitation type, and the rain and snow weather are respectively {12,18}, G σs is the spatial Gaussian kernel; S17. The rain and snow degradation prior model generated by adversarial training is embedded in the U-Net decoder, and the loss function is: L=λ ce L ce +λ adv L adv ; Among them, λ ce =1.0,λ adv =0.1.

[0008] Preferably, the training of the meteorological classifier in step S15 includes: S151. Build a synthetic training dataset, use UnrealEngine to render 200 precipitation scenes, and generate samples of different precipitation densities and wind directions through domain randomization; S152, extracting a structure through a dual-branch feature, wherein the dual-branch feature includes: The main branch processes RGB images; Auxiliary branch analysis of thermal imaging temperature distribution patterns; S153, light rain / light snow samples are used in the initial stage of training, and finally the weight coefficients of extreme weather such as heavy rain / heavy snow are increased, and are respectively imported into step S16 to obtain the similarity weights of the pixel blocks in the two-dimensional search window to form a data set.

[0009] Preferably, the step S03 of constructing a personalized habitual route database includes: S31. Record trajectory data through GNSS / IMU fusion positioning module, and the major semi-axis of the positioning error ellipse is ≤ 0.8 meters; S32, using the spatiotemporal density clustering algorithm, the spatiotemporal neighborhood determination conditions are: ; ; Where r = 1.2 × average step length, Δt = 5 minutes, (x i ,y i ) The two-dimensional coordinates of the user's i-th trajectory point, (x j ,y j ) The two-dimensional coordinates of the j-th trajectory point of the user, ti,tj is the timestamp of trajectory point i,j.

[0010] Preferably, when the GNSS signal is lost in the above step S03, the following steps are performed: S71. Dead reckoning based on IMU data: ; ; in, =0.2, v(t) is the linear velocity of the user at time t, a(τ) is the IMU acceleration measurement value at time τ, θ(t) is the heading angle of the user at time t, ω (τ) is the angular velocity of the IMU gyroscope at time τ, δ is the deviation between the heading angle measured by the magnetometer and the gyroscope integrated heading angle; S72, every 2 seconds, the cumulative error is corrected by the landmark recognition result, and the correction amount is calculated as: ; Among them, k p =08, k i =0.05, is the deviation between the identified position and the estimated position of the landmark at the current moment, is the positioning error correction.

[0011] Preferably, the step S04 of calculating the optimal navigation path that complies with human kinematic constraints by using the Bayesian probability graph model includes: S41, obtaining an offline map database, defining the current walking section, and obtaining the coordinates of obstacles around the current positioning; S42, the real-time perception data obtained is distributed in the offline map database in a point shape, and a first coordinate set of the maximum left-right swing amplitude of the cane and a second coordinate set of the left-right swing amplitude of the walking body are synchronously obtained; S43, connecting multiple dot icons located in the offline map database, determining the walking trajectory of the human body, predicting obstacles that will be touched during future walking, extracting the coordinates of the obstacles, and calculating and generating correction instructions based on the Bayesian probability graph model.

[0012] The present invention has at least the following beneficial effects: Through the infrared-visible light binocular camera and image enhancement algorithm, the problems of poor image quality and reflection interference in rainy and snowy weather are solved. Therefore, the present invention can still "see clearly" the environment under adverse conditions such as water accumulation, hail, and rainstorms, and avoid collisions caused by missed obstacles by blind people. Combined with the user's historical walking habits and human kinematics model, a path that conforms to the natural walking pattern of the blind is generated, which reduces counterintuitive instructions such as right-angle turns and the probability of deviation from the path, thereby preventing the blind from frequently adjusting the direction due to unreasonable path planning, and improving navigation comfort and safety; Through timestamp synchronization and Kalman filter algorithm, the positioning drift problem caused by the asynchrony between camera and IMU data is solved, so as to correct the positioning in real time and avoid the potential danger of positioning error to the navigation of blind people, so as to ensure the accurate position calculation of blind people while walking and reduce navigation errors caused by positioning drift. The extreme weather data synthesized by the virtual engine is used to train the meteorological classifier, dynamically adjust the noise reduction parameters, and limit the maximum turning angle in hail weather, so as to improve the robustness of the present invention in extreme scenarios such as heavy rain and heavy snow, and avoid navigation failure due to sudden weather changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0014] Figure 1 A flow chart of a voice navigation assistance method for the blind with a built-in offline AI intelligent model provided in the first embodiment of the present invention; Figure 2 This is an image generated by integrating real-time perception data and an offline map database provided in the first embodiment of the present invention. DETAILED DESCRIPTION

[0015] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0016] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0017] Embodiment 1:

[0018] This embodiment provides a voice navigation assistance method for the blind with a built-in offline AI intelligent model, the method comprising the following steps: Figure 1 As shown: S01. Collect environmental image data in real time through an infrared-visible light binocular camera, and synchronously obtain six-degree-of-freedom motion posture data of an inertial measurement unit. The binocular infrared-visible light binocular camera includes a thermal imaging sensor channel; Specifically, the above-mentioned infrared-visible light binocular camera collects environmental image data in real time, which means that infrared-visible light binocular cameras are set on both sides of the blind person's eyes to obtain high-definition images and infrared images as needed. The image data of the two are superimposed and used. The purpose is to supplement the high-definition images at night and in rainy and snowy weather, thereby increasing the recognition of moving targets in rainy and snowy weather at night, thereby avoiding collisions between other pedestrians and animals and the user.

[0019] Furthermore, the motion posture data in the motion posture data of the inertial measurement unit synchronously acquired in the above step S01 includes six-degree-of-freedom motion posture data corresponding to the time phase of each frame of the environmental image data, and the steps include: S11. Establish the rigid transformation matrix from the IMU coordinate system to the infrared-visible light binocular camera coordinate system: ; Among them, R is the rotation matrix and t is the translation vector; S12, inserting a synchronization timestamp into each frame image and IMU data stream; S13. When it is detected that the IMU data delay exceeds 10ms, the Kalman filter is started to predict the motion state: ; in, is the estimated value of the state at time k, is the state transfer matrix, The dynamic evolution law from time to time k, is the control input matrix, is the control input vector, is the noise vector.

[0020] In order to ensure that the data from the camera and the inertial measurement unit can "match", a coordinate system transformation matrix is ​​established to unify the data of the two into one coordinate system. Then, a timestamp is added to each frame of the image and the IMU data stream, and the two data are compared by the timestamp. If the data with the same timestamp are consistent, it means that the IMU data is correct. If they are inconsistent, it means that the IMU data is delayed. It will use the Kalman filter to predict the motion state, and the value obtained by the Kalman filter is used as the value of the IMU data stream, thereby greatly improving the accuracy and stability of the data.

[0021] Furthermore, in step S01, the real-time collected environment image data is preprocessed, including: S14. Based on the multi-scale Retinex algorithm, the brightness component V(x, y) is separated in the HSV color space, and the image reflection layer R(x, y) and the illumination layer L(x, y) are extracted through the Gaussian difference filter group. The formula is as follows: ; Among them, F i is a Gaussian kernel with radius {5, 15, 30} pixels, wi = {0.3, 0.5, 0.2}; S15. Build a meteorological classifier based on SwinTransformer. When hail is identified, activate the path curvature constraint module to limit the turning angle to ≤45°. S16. Based on the non-local mean filtering algorithm, a similarity weight of pixel blocks in the two-dimensional search window is constructed, and the formula is as follows: w(p,q)=exp(-||N(p)-N(q)||² / (2σ²))·G σs (||pq||); Among them, σ is dynamically adjusted according to the precipitation type, and the rain and snow weather are respectively {12,18}, G σs is the spatial Gaussian kernel; S17. The rain and snow degradation prior model generated by adversarial training is embedded in the U-Net decoder, and the loss function is: L=λ ce L ce +λ adv L adv ; Among them, λ ce =1.0,λ adv =0.1.

[0022] After the environmental image is collected, the image will be preprocessed to make it clearer and more accurate. The multi-scale Retinex algorithm is used to separate the brightness component of the image from the reflection layer and the illumination layer, making the image less susceptible to the influence of illumination. Then, the meteorological classifier is used to identify the weather conditions. If it is raining or snowing, it will perform special processing on the image to reduce the impact of weather on navigation. Finally, the non-local mean filtering algorithm is used to reduce noise and blur in the image. This improves the clarity of the image, reduces image noise and blur, and uses the meteorological classifier to perform special processing according to the weather conditions to ensure navigation needs in rainy and snowy weather.

[0023] S02, inputting the environmental image data into the pre-trained multi-task neural network model to generate a three-dimensional semantic map including the topological structure of the blind path and dynamic obstacle detection results; The above-mentioned image information captured from the environment, such as pictures of streets, indoor scenes or any other visual scenes taken by a camera. The image data is then input into a trained neural network model. This model is "multi-tasking", which means that it can perform multiple tasks at the same time, rather than focusing on a specific output (such as only identifying objects or only generating maps). The trained neural network model then generates a three-dimensional semantic map based on the input image data. It also contains connection information about the actual blind path and the virtual blind path. The so-called virtual blind path refers to a temporary blind path generated when it is detected that the actual blind path is occupied, and the virtual blind path will disappear after it is guided to the actual blind path. The above-mentioned trained neural network model belongs to the prior art, and the model logic of intelligent driving of automobiles can be referred to.

[0024] S03. Based on the deep reinforcement learning algorithm, the user's historical paths are analyzed in time and space to build a personalized habit route database that integrates human steering inertia; Specifically, the step S03 of building a personalized habit route database includes: S31. Record trajectory data through GNSS / IMU fusion positioning module, and the major semi-axis of the positioning error ellipse is ≤ 0.8 meters; S32, using the spatiotemporal density clustering algorithm, the spatiotemporal neighborhood determination conditions are: ; ; Where r = 1.2 × average step length, Δt = 5 minutes, (x i ,y i ) The two-dimensional coordinates of the user's i-th trajectory point, (x j ,y j) The two-dimensional coordinates of the j-th trajectory point of the user, ti,tj is the timestamp of trajectory point i,j.

[0025] The GNSS / IMU fusion positioning module is used to accurately record trajectory data, providing a solid foundation for building a personalized route database. The trajectory data is deeply analyzed through the spatiotemporal density clustering algorithm to obtain the routes and paths that the blind often take, so as to provide more personalized services for the blind based on these habitual routes.

[0026] Furthermore, when the GNSS signal is lost in the above step S03, the following steps are performed: S71. Dead reckoning based on IMU data: ; ; in, =0.2, v(t) is the linear velocity of the user at time t, a(τ) is the IMU acceleration measurement value at time τ, θ(t) is the heading angle of the user at time t, ω (τ) is the angular velocity of the IMU gyroscope at time τ, δ is the deviation between the heading angle measured by the magnetometer and the gyroscope integrated heading angle; S72, every 2 seconds, the cumulative error is corrected by the landmark recognition result, and the correction amount is calculated as: ; Among them, k p =08, k i =0.05, is the deviation between the identified position and the estimated position of the landmark at the current moment, is the positioning error correction.

[0027] In the case of GNSS signal loss, the continuity and accuracy of navigation are maintained through dead reckoning based on IMU data and correction of landmark recognition results. When the GNSS signal is lost, the approximate position of the blind person is predicted based on the dead reckoning of IMU data. At the same time, the landmark recognition function is used to correct the previous calculation results at regular intervals to ensure the accuracy of the position. The coordinated operation of the two execution strategies ensures that the navigation system can provide continuous and accurate navigation services even when the GNSS signal is unstable.

[0028] S04, integrating real-time perception data with offline map database, and calculating the optimal navigation path that meets human kinematic constraints through the Bayesian probability graph model; Specific, combined Figure 2 As shown, the detailed implementation steps of step S04 are as follows: S41, obtaining an offline map database, defining the current walking section, and obtaining the coordinates of obstacles around the current positioning; S42, the acquired real-time perception data is distributed in the offline map database in a point shape, and a first coordinate set of the maximum left-right swing amplitude of the cane and a second coordinate set of the left-right swing amplitude of the walking body are synchronously acquired; S43, connecting multiple dot icons located in the offline map database, determining the walking trajectory of the human body, predicting obstacles that will be touched during future walking, extracting the coordinates of the obstacles, and calculating and generating correction instructions based on the Bayesian probability graph model.

[0029] It should be noted that the first coordinate set in the above step S42 is the coordinates of the two endpoints of the maximum left and right swing amplitude of the cane when the device is enabled, and the second coordinate set is the straight-line distance between the two side edges of the image under the environmental image data obtained by the infrared-visible light binocular camera, and the endpoint coordinates of the two ends of the straight line located on the current road section are obtained based on the straight-line distance.

[0030] The above-mentioned Bayesian probability graph model integrates multiple factors such as human kinematic constraints, real-time perception data, and offline map information, and adaptively designs paths according to the behavioral habits of each blind person. At the same time, it can also predict in real time the obstacles that may be encountered during future walking, and give correction instructions in advance to ensure that the blind can easily cope with various complex environments. It significantly improves the accuracy and safety of navigation, enhances the humanity and comfort of navigation, and provides more considerate services for the blind.

[0031] S05. Convert the result of the optimal navigation path into a natural language instruction containing a description of the spatial reference system, and realize multimodal interaction through bone conduction headphones and a tactile feedback device.

[0032] Specifically, a spatial language is generated based on the result of the optimal navigation path obtained, that is, a natural language instruction described by the spatial reference system, such as turning left and walking 100 meters, passing the red building, and currently there will be no interaction with the red building, and walking in the current walking direction.

[0033] In summary, the method provided in the first embodiment solves the problems of poor image quality and reflection interference in rainy and snowy weather through an infrared-visible light binocular camera and an image enhancement algorithm, so that the present invention can still "see clearly" the environment under adverse conditions such as water accumulation, hail, and rainstorms, and avoid collisions caused by blind people due to missed obstacles. In addition, combined with the user's historical walking habits and human kinematics model, a path that conforms to the natural walking pattern of the blind is generated, counter-intuitive instructions such as right-angle turns are reduced, and the probability of deviation from the path is reduced, thereby avoiding the blind from frequently adjusting the direction due to unreasonable path planning, improving navigation comfort and safety. Secondly, through timestamp synchronization and Kalman filtering algorithms, the positioning drift problem caused by the asynchrony between the camera and IMU data is solved, so as to correct the positioning in real time and avoid the potential danger of positioning errors to the navigation of the blind, so as to ensure that the position calculation of the blind is accurate during walking and reduce navigation errors caused by positioning drift. Furthermore, the extreme weather data synthesized by the virtual engine is used to train the meteorological classifier, dynamically adjust the noise reduction parameters, and limit the maximum turning angle in hail weather, so as to improve the robustness of the present invention in extreme scenarios such as heavy rain and heavy snow, and avoid navigation failure due to sudden weather changes.

[0034] Embodiment 2: Based on the above embodiment 1, this embodiment aims to train the "constructing a meteorological classifier based on SwinTransformer" in the above step S01, and the steps include: S151. Build a synthetic training dataset, use UnrealEngine to render 200 precipitation scenes, and generate samples of different precipitation densities and wind directions through domain randomization; S152, extracting a structure through a dual-branch feature, wherein the dual-branch feature includes: The main branch processes RGB images; Auxiliary branch analysis of thermal imaging temperature distribution patterns; S153, light rain / light snow samples are used in the initial stage of training, and finally the weight coefficients of extreme weather such as heavy rain / heavy snow are increased, and are respectively imported into step S16 to obtain the similarity weights of pixel blocks in the two-dimensional search window to form a data set.

[0035] In this embodiment, UnrealEngine is used to render a rich precipitation scene and construct a large synthetic training data set. A dual-branch feature extraction structure is used to process RGB images and thermal imaging temperature distribution patterns respectively, thereby extracting richer image features. During the training process, the need to dynamically adjust the weight of samples according to the severity of the weather is also taken into account to ensure that the classifier can pay more attention to those extreme weather conditions. As a result, the recognition accuracy and robustness of the meteorological classifier are significantly improved.

[0036] Embodiment three: An embodiment of the present invention provides a non-transitory computer-readable storage medium, wherein at least one instruction or at least one program is stored in the non-transitory computer-readable storage medium, and the at least one instruction or at least one program is loaded and executed by a processor to implement the steps: The infrared-visible light binocular camera collects environmental image data in real time and simultaneously obtains the six-degree-of-freedom motion posture data of the inertial measurement unit. The binocular infrared-visible light binocular camera includes a thermal imaging sensor channel. Input the environmental image data into the pre-trained multi-task neural network model to generate a 3D semantic map containing the topological structure of the blind path and dynamic obstacle detection results; Based on the deep reinforcement learning algorithm, the user's historical paths are clustered in time and space to build a personalized habit route database that integrates human steering inertia; Integrate real-time perception data with offline map database and calculate the optimal navigation path that meets human kinematic constraints through Bayesian probabilistic graphical model; The results of the optimal navigation path are converted into natural language instructions containing a description of the spatial reference system, and multimodal interaction is achieved through bone conduction headphones and tactile feedback devices.

[0037] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0038] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0039] Embodiment 4:

[0040] An embodiment of the present invention provides an electronic device, including a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the steps: The infrared-visible light binocular camera collects environmental image data in real time and simultaneously obtains the six-degree-of-freedom motion posture data of the inertial measurement unit. The binocular infrared-visible light binocular camera includes a thermal imaging sensor channel. Input the environmental image data into the pre-trained multi-task neural network model to generate a 3D semantic map containing the topological structure of the blind path and dynamic obstacle detection results; Based on the deep reinforcement learning algorithm, the user's historical paths are clustered in time and space to build a personalized habit route database that integrates human steering inertia; Integrate real-time perception data with offline map database and calculate the optimal navigation path that meets human kinematic constraints through Bayesian probabilistic graphical model; The results of the optimal navigation path are converted into natural language instructions containing a description of the spatial reference system, and multimodal interaction is achieved through bone conduction headphones and tactile feedback devices.

[0041] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technician familiar with this profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A voice navigation assistance method for the blind with a built-in offline AI intelligent model, characterized in that: The following steps are involved: S01, collecting environmental image data in real time through an infrared-visible light binocular camera, and synchronously acquiring six-degree-of-freedom motion posture data of an inertial measurement unit, wherein the binocular infrared-visible light binocular camera includes a thermal imaging sensor channel; S02, inputting the environmental image data into a pre-trained multi-task neural network model to generate a three-dimensional semantic map including a topological structure of a blind path and a dynamic obstacle detection result; S03. Based on the deep reinforcement learning algorithm, the user's historical paths are analyzed in time and space to build a personalized habit route database that integrates human steering inertia; S04, integrating real-time perception data with offline map database, and calculating the optimal navigation path that meets human kinematic constraints through the Bayesian probability graph model; S05. Convert the result of the optimal navigation path into a natural language instruction containing a description of the spatial reference system, and realize multimodal interaction through bone conduction headphones and a tactile feedback device.

2. According to claim 1, a voice navigation assistance method for the blind with a built-in offline AI intelligent model is characterized in that: The step S01 of synchronously acquiring the motion posture data of the inertial measurement unit includes six-degree-of-freedom motion posture data corresponding to the time phase of each frame of the environmental image data, and the steps include: S11. Establish the rigid transformation matrix from the IMU coordinate system to the infrared-visible light binocular camera coordinate system: ; Among them, R is the rotation matrix and t is the translation vector; S12, inserting a synchronization timestamp into each frame image and IMU data stream; S13. When the IMU data delay is detected to exceed 10ms, the Kalman filter is started to predict the motion state: ; in, is the estimated value of the state at time k, is the state transfer matrix, The dynamic evolution law from time to time k, is the control input matrix, is the control input vector, is the noise vector.

3. According to the method of claim 1, wherein the blind person's voice navigation assistance method with a built-in offline AI intelligent model is characterized in that: In step S01, the real-time collected environment image data is preprocessed, including: S14. Based on the multi-scale Retinex algorithm, the brightness component V(x, y) is separated in the HSV color space, and the image reflection layer R(x, y) and the illumination layer L(x, y) are extracted through the Gaussian difference filter group. The formula is as follows: ; Among them, F i is a Gaussian kernel with radius {5, 15, 30} pixels, wi = {0.3, 0.5, 0.2}; S15. Build a meteorological classifier based on SwinTransformer. When hail is identified, activate the path curvature constraint module to limit the turning angle to ≤45°. S16. Based on the non-local mean filtering algorithm, a similarity weight of pixel blocks in the two-dimensional search window is constructed, and the formula is as follows: w(p,q)=exp(-||N(p)-N(q)||² / (2σ²))·G σs (||p-q||); Among them, σ is dynamically adjusted according to the precipitation type, and the rain and snow weather are respectively {12,18}, G σs is the spatial Gaussian kernel; S17. The rain and snow degradation prior model generated by adversarial training is embedded in the U-Net decoder, and the loss function is: L=λ ce L ce +λ adv L adv ; Among them, l ce =1.0,λ adv =0.1。 4. The method for assisting blind people's voice navigation with a built-in offline AI intelligent model according to claim 1, characterized in that: The training of the meteorological classifier in step S15 includes: S151. Build a synthetic training dataset, use UnrealEngine to render 200 precipitation scenes, and generate samples of different precipitation densities and wind directions through domain randomization; S152, extracting a structure through a dual-branch feature, wherein the dual-branch feature includes: The main branch processes RGB images; Auxiliary branch analysis of thermal imaging temperature distribution patterns; S153, light rain / light snow samples are used in the initial stage of training, and finally the weight coefficients of extreme weather such as heavy rain / heavy snow are increased, and are respectively imported into step S16 to obtain the similarity weights of the pixel blocks in the two-dimensional search window to form a data set.

5. The method for assisting blind people's voice navigation with a built-in offline AI intelligent model according to claim 1, characterized in that: The step S03 of constructing a personalized habitual route database includes: S31. Record trajectory data through GNSS / IMU fusion positioning module, and the major semi-axis of the positioning error ellipse is ≤ 0.8 meters; S32, using the spatiotemporal density clustering algorithm, the spatiotemporal neighborhood determination conditions are: ; ; Where r = 1.2 × average step length, Δt = 5 minutes, (x i ,y i ) The two-dimensional coordinates of the user's i-th trajectory point, (x j ,y j ) The two-dimensional coordinates of the j-th trajectory point of the user, ti,tj is the timestamp of trajectory point i,j.

6. The method for assisting blind people's voice navigation with a built-in offline AI intelligent model according to claim 5, characterized in that: When the GNSS signal is lost in the above step S03, the following steps are performed: S71. Dead reckoning based on IMU data: ; ; in, =0.2, v(t) is the linear velocity of the user at time t, a(τ) is the IMU acceleration measurement value at time τ, θ(t) is the heading angle of the user at time t, ω (τ) is the angular velocity of the IMU gyroscope at time τ, δ is the deviation between the heading angle measured by the magnetometer and the gyroscope integrated heading angle; S72, every 2 seconds, the cumulative error is corrected by the landmark recognition result, and the correction amount is calculated as: ; Among them, k p =08, k i =0.05, is the deviation between the landmark identification position and the estimated position at the current moment, is the positioning error correction.

7. The method for assisting blind people's voice navigation with a built-in offline AI intelligent model according to claim 1, characterized in that: The step S04 of calculating the optimal navigation path that complies with human kinematic constraints by using the Bayesian probability graph model includes: S41, obtaining an offline map database, defining the current walking section, and obtaining the coordinates of obstacles around the current positioning; S42, the real-time perception data obtained is distributed in the offline map database in a point shape, and a first coordinate set of the maximum left-right swing amplitude of the cane and a second coordinate set of the left-right swing amplitude of the walking body are synchronously obtained; S43, connecting multiple dot icons located in the offline map database, determining the walking trajectory of the human body, predicting obstacles that will be touched during future walking, extracting the coordinates of the obstacles, and calculating and generating correction instructions based on the Bayesian probability graph model.

8. The method for assisting blind people's voice navigation with a built-in offline AI intelligent model according to claim 7, characterized in that: The first coordinate set in step S42 is the coordinates of the two endpoints of the maximum left and right swing amplitude of the blind stick when the device is enabled, and the second coordinate set is the straight-line distance between the two side edges of the image under the environmental image data obtained by the infrared-visible light binocular camera, and the endpoint coordinates of the two ends of the straight line located on the current road section are obtained based on the straight-line distance.

9. A non-transitory computer-readable storage medium, wherein at least one instruction or at least one program is stored in the non-transitory computer-readable storage medium, characterized in that: The at least one instruction or the at least one program is loaded and executed by the processor to implement the steps of the blind person voice navigation assistance method with a built-in offline AI intelligent model as described in any one of claims 1-8.

10. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the steps of the blind person voice navigation assistance method with a built-in offline AI intelligent model as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Localization method and device based on deep fusion of visual and inertial data

    CN109238277B

  • Line weight determination method for Dijkstra algorithm

    CN117091599A

  • Method and device for planning flight path of unmanned aerial vehicle under wind and rain conditions

    CN117029827A

  • Pedestrian navigation system and path planning method based on environmental perception and human kinematics

    CN117213513A

  • Blind person intelligent navigation system and method based on image semantic segmentation

    CN118298170A

Cited By

  • Port internal and external fleet positioning system and method

    CN120282264A

  • Port internal and external vehicle fleet positioning system and method

    CN120282264B

  • Production park personnel positioning method and system

    CN120321587A

  • Blind person environment cognition auxiliary system

    CN121280967A