Navigation instruction generation and path updating method and device, equipment and medium

By acquiring environmental images and depth information, performing open vocabulary object detection and generating navigation instructions, the problem of inaccurate obstacle recognition in existing navigation systems is solved, more accurate navigation instruction generation and path updates are achieved, and the safety and practicality of the navigation system are improved.

CN120651241APending Publication Date: 2025-09-16PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510946218.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing navigation systems are unable to provide navigation instructions that include obstacle categories, precise directions, and distance information, and lack semantic-level obstacle recognition and executable directional guidance capabilities, resulting in unstable navigation, poor safety, and poor practicality for visually impaired users in complex environments.

Method used

By acquiring environmental image information and depth information, open vocabulary object detection is performed to generate the spatial position and direction information of obstacles. In combination with the environmental image information, navigation instructions containing scene descriptions, obstacle orientations, and distance descriptions are generated to dynamically update the navigation path.

Benefits of technology

It improves the accuracy of obstacle perception and the practicality of navigation guidance, enhances navigation safety in complex environments, and reduces the path risks of visually impaired users in actual travel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120651241A_ABST
    Figure CN120651241A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of intelligent auxiliary navigation, financial science and technology, medical health and the like, and discloses a navigation instruction generation and path updating method, device, equipment and medium. And fusing the obstacle category information and the environment depth information, generating spatial position information and direction information of the obstacle, generating a navigation instruction containing scene description information, obstacle orientation and distance description information and action suggestion information based on the spatial position information, the direction information and the environment image information, and updating a navigation path. According to the method, the obstacle category information and the environment depth information are fused, the obstacle recognition result with the semantic level is generated in combination with the environment image information, the navigation instruction containing the orientation and the distance is output, and the accuracy of obstacle perception and the practicability of navigation guidance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a navigation instruction generation and path updating method, device, equipment and storage medium. Background Art

[0002] In the field of intelligent assisted navigation technology, especially in applications that provide navigation services for visually impaired elderly people using smart wearable devices, accurate acquisition of environmental and obstacle information is a crucial prerequisite for safe travel. Existing technologies suffer from incomplete perception of surrounding data, inaccurate spatial information acquisition, and lack of targeted navigation feedback, which seriously impact the reliability and practicality of navigation systems in complex real-world environments.

[0003] Specifically, traditional guide devices such as electronic canes and ultrasonic detectors rely primarily on close-range single-point detection, which prevents them from forming a comprehensive perception of the overall environment. They also lack the ability to understand the semantics of obstacle types, relative orientations, and distances, making it difficult for visually impaired users to obtain clear spatial reference information based on these devices. Meanwhile, some systems that rely on multimodal large models (MLLMs) have certain image recognition capabilities, but due to limited training data sources, particularly the lack of professional data annotation standards tailored to the actual needs of the visually impaired, these systems suffer from unstable object detection accuracy and road environment recognition in complex open scenarios, making them unable to meet the high-accuracy navigation needs in dynamic environments such as urban blocks and outdoor transportation hubs.

[0004] In the field of financial technology, with the rise of intelligent outlets and customer escort services, some financial institutions have tried to introduce assisted navigation functions to assist visually impaired users in handling business. However, due to the weak ability of existing equipment to obtain environmental information, inaccurate obstacle positioning, and the lack of a real-time adjustment mechanism for dynamic navigation paths, the autonomous movement of visually impaired groups in and around financial service outlets is inefficient and has high security risks, making it impossible to effectively guarantee the accessibility and convenience of financial services.

[0005] In the medical and health business field, some rehabilitation assistive devices integrate basic environmental perception functions and are used in travel assistance scenarios for the elderly. However, existing devices mostly rely on fixed sensor arrays or single data sources and lack the ability to accurately identify the spatial position and relative direction of obstacles. Especially in urban environments where lighting changes and dynamic obstacles frequently appear, navigation paths are not updated in a timely manner, and there are risks of the elderly getting lost, deviating from safe paths, and colliding with obstacles, which restricts the application and promotion of such devices in the medical and health service system. Summary of the Invention

[0006] The main purpose of the present invention is to provide a navigation instruction generation and path updating method, device, equipment and storage medium, aiming to solve the technical problems that the existing technology cannot provide navigation instructions containing obstacle categories, precise direction and distance information, and the existing navigation system lacks semantic-level obstacle recognition and executable directional guidance capabilities.

[0007] To achieve the above objectives, the present invention provides a navigation instruction generation and path updating method, comprising:

[0008] Obtaining environmental image information and environmental depth information;

[0009] Performing open vocabulary object detection on the environmental image information to obtain obstacle category information;

[0010] Fusion of the obstacle category information and the environment depth information to generate spatial position information and direction information of the obstacle;

[0011] Based on the spatial position information, the direction information, and the environmental image information, generating navigation instructions including scene description information, obstacle orientation and distance description information, and action suggestion information;

[0012] The navigation path is updated according to the navigation instruction.

[0013] Furthermore, to achieve the above-mentioned purpose, the present invention provides a navigation instruction generation and path updating device, comprising:

[0014] Environmental perception module, used to obtain environmental image information and environmental depth information;

[0015] an obstacle recognition module, configured to perform open vocabulary object detection on the environmental image information to obtain obstacle category information;

[0016] A spatial modeling module, configured to fuse the obstacle category information and the environment depth information to generate spatial position information and direction information of the obstacle;

[0017] A navigation instruction generation module, configured to generate a navigation instruction including scene description information, obstacle orientation and distance description information, and action suggestion information based on the spatial position information, direction information, and environmental image information;

[0018] A path planning module is used to update the navigation path according to the navigation instruction.

[0019] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a navigation instruction generation and path update program stored in the memory and executable on the processor. When the navigation instruction generation and path update program is executed by the processor, the steps of the navigation instruction generation and path update method as described above are implemented.

[0020] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a navigation instruction generation and path update program is stored. When the navigation instruction generation and path update program is executed by a processor, the steps of the navigation instruction generation and path update method as described above are implemented.

[0021] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as intelligent assisted navigation, financial technology, and medical health. A method, device, equipment, and medium for generating navigation instructions and updating a path are disclosed, including: obtaining environmental image information and environmental depth information, performing open vocabulary object detection, obtaining obstacle category information, fusing the obstacle category information and environmental depth information, generating spatial position information and direction information of the obstacle, and generating navigation instructions containing scene description information, obstacle orientation and distance description information, and action suggestion information based on the spatial position information, direction information, and environmental image information, and updating the navigation path. The present invention generates obstacle recognition results with a semantic level by fusing obstacle category information and environmental depth information, combining environmental image information, and outputting navigation instructions containing orientation and distance. This can improve the accuracy of obstacle perception and the practicality of navigation guidance, and enhance navigation safety in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0023] Figure 1 A schematic diagram of an application environment of a navigation instruction generation and path updating method according to an embodiment of the present invention;

[0024] Figure 2 This is a flow chart of an embodiment of a method for generating navigation instructions and updating a path according to the present invention;

[0025] Figure 3 A schematic diagram of the functional modules of a preferred embodiment of the navigation instruction generation and path updating device of the present invention;

[0026] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0027] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0029] The navigation instruction generation and path updating method provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can obtain environmental image information and environmental depth information through the user terminal, perform open vocabulary object detection, obtain obstacle category information, fuse the obstacle category information and environmental depth information, generate spatial position information and direction information of the obstacle, and based on the spatial position information, direction information and environmental image information, generate navigation instructions containing scene description information, obstacle orientation and distance description information and action suggestion information, and update the navigation path. The present invention generates obstacle recognition results with semantic level by fusing obstacle category information and environmental depth information, combining environmental image information, and outputting navigation instructions containing orientation and distance, which can improve the accuracy of obstacle perception and the practicality of navigation guidance, and enhance navigation safety in complex environments. Among them, the user terminal can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.

[0030] See also Figure 2 , Figure 2 This is a flowchart of an embodiment of the method for generating navigation instructions and updating a route provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0031] like Figure 2 As shown, the navigation instruction generation and path updating method proposed by the present invention includes the following steps:

[0032] S10, obtaining environmental image information and environmental depth information;

[0033] In this embodiment, obtaining environmental image information and environmental depth information involves collecting multi-source perception data reflecting the environmental structure, spatial distribution, and the presence of obstacles through a device with visual perception capabilities. Environmental image information refers to pixel data containing visible light or near-infrared band images obtained by an image sensor. This image information has a two-dimensional spatial structure and can present visual information in the current scene, specifically including information such as ground structure, object contours, lighting changes, and material texture. The acquisition of image information can be based on a monocular camera, a binocular stereo camera, a panoramic camera, or other sensing device with visual acquisition capabilities. The image format can be RGB, grayscale, infrared, or multi-channel. The resolution, frame rate, and dynamic range of the image can be flexibly adjusted according to environmental requirements to meet application requirements of different lighting, spatial complexity, and scene types.

[0034] Environmental depth information is the distance data between various locations in space and the acquisition device obtained by a sensor device with depth perception capabilities. Depth information is a key data form that reflects the three-dimensional structure of the environment and can assist in understanding the location of objects, spatial layout, and obstacle distribution. The acquisition of depth information can be based on structured light sensors, TOF time-of-flight sensors, lidar, millimeter-wave radar, binocular visual structure reconstruction, or other sensing devices with spatial distance measurement capabilities. When acquiring depth information, it is usually necessary to synchronize image information and depth information to ensure that the two have a unified spatial reference system and timestamp, so as to achieve the joint expression of spatial structure and visual information. Depth information can be expressed in the form of depth maps, point cloud data, dense depth matrices, or sparse depth samples. The specific form is determined by the sensor type and scene adaptation strategy.

[0035] Acquiring environmental image and depth information requires precise calibration of the spatial relationship between the image and depth sensors to ensure consistent mapping between the image and depth coordinate systems. During synchronous acquisition, a unified timestamp or hardware-level synchronization signal trigger mechanism is used to align the image and depth information, avoiding data deviations caused by time delays or asynchrony. To ensure information acquisition quality, exposure parameters, gain parameters, and filtering strategies can be adjusted based on environmental conditions to minimize interference from complex environments such as strong and weak light, rain, and fog.

[0036] Acquiring environmental image information and environmental depth information can be achieved by using a single sensor device. By integrating a color camera and a TOF depth sensor, the RGB image and the corresponding depth matrix can be output synchronously in real time, achieving joint perception of the environmental scene. Alternatively, an external multi-sensor system can be used, with high-resolution image sensors and high-precision lidars deployed separately. The image sensor provides clear visual information, and the lidar provides stable three-dimensional distance data. The system integrates information through a software and hardware synchronization mechanism to achieve multimodal environmental perception. In addition, in specific application scenarios, such as low-light environments, a combination of near-infrared imaging and active structured light depth sensing can be used preferentially to improve information acquisition capabilities at night or under complex lighting conditions, enhancing the system's adaptability to different environments.

[0037] For assisted navigation in healthcare, a low-power, lightweight depth camera combined with a wide-angle image sensor can be used to reduce the burden on elderly users while ensuring effective perception of complex indoor environments such as corridors and furniture. By optimizing sensor parameters, the accuracy of information acquisition in low-light environments can be improved, reducing navigation errors caused by information loss.

[0038] In the assisted navigation scenarios of the financial technology business field, it is suitable for semi-closed environments such as bank branches and self-service equipment areas. It can use high-resolution image sensing equipment combined with high-density point cloud data acquisition devices to achieve comprehensive perception of densely populated areas, dynamic obstacles and spatial layouts, ensuring the safety of users in the process of handling financial business and the accuracy of path guidance.

[0039] Example: In the healthcare business, visually impaired elderly people wear smart glasses with integrated vision and depth perception functions. The system obtains real-time images and depth information of scenes such as corridors, wards, and waiting areas, accurately perceives ground structure, obstacle locations, and dynamic human flow, and combines indoor layout to achieve safe path planning and real-time navigation, effectively reducing the risk of falls or getting lost due to environmental complexity.

[0040] In the field of financial technology business, users wear auxiliary equipment with environmental perception functions in bank business halls, self-service equipment areas and other areas. The system obtains scene images and depth information, and identifies the locations and categories of obstacles such as queues, self-service terminals, and temporarily placed items in real time. It combines spatial structure information to output safe and feasible path navigation information to ensure the spatial safety and path guidance accuracy of users when handling business.

[0041] This embodiment, through the combined acquisition of environmental image information and depth information, can provide a rich and accurate data foundation for subsequent obstacle identification, spatial structure understanding, and path guidance. Image information provides a comprehensive visual representation, assisting in understanding object categories, material properties, and lighting conditions, while depth information complements three-dimensional spatial structure and distance distribution, achieving a multi-dimensional and integrated understanding of the environment, enhancing the system's navigation accuracy and safety capabilities in complex environments.

[0042] S20, performing open vocabulary object detection on the environmental image information to obtain obstacle category information;

[0043] In this embodiment, environmental image information refers to image data collected by visual perception devices that reflects the visual state of the current spatial environment. This information includes multi-dimensional image content, including color information, spatial structure, object outlines, texture characteristics, and light-dark relationships. This information is derived from the visual sensor integrated into the front end of the smart glasses or an external high-precision camera module. It provides real-time and high-resolution capabilities, meeting the navigation system's requirements for environmental understanding and spatial cognition.

[0044] Open vocabulary object detection is a visual recognition process based on a deep learning framework. Unlike traditional closed-category object detection, open vocabulary object detection is not limited to a fixed set of labels. Instead, it incorporates an open semantic system, dynamically adapting the representation of obstacle categories in different scenarios through an extended language model or a large-scale semantic description library. The semantic description library can be built based on a large general language model, a knowledge graph, or a custom vocabulary, covering a variety of natural language categories, attributes, or functional phrases.

[0045] To achieve open-vocabulary object detection, the system first extracts multi-scale visual features from the surrounding image information. Combining deep convolutional neural networks, visual transformers, or other feature encoding structures, it encodes each pixel region, boundary contours, and texture distribution in the image into a high-dimensional feature vector. This feature mapping preserves spatial information, ensuring that object location and category information can be linked.

[0046] Furthermore, the system performs cross-modal matching between visual feature vectors and text embedding vectors from an open semantic description library. Text embedding vectors are derived from large-scale language models or domain-specific training models and possess the ability to understand and express semantic associations. By calculating the similarity score between visual features and text embeddings, the system can determine the degree of match between objects in the image and the category representations in the semantic library.

[0047] Based on similarity assessment, the system selects category labels that are highly relevant to the image content and dynamically generates obstacle category information. This information not only includes basic category names but also incorporates multi-dimensional descriptions based on semantic content such as material properties, dimensional characteristics, and functional usage. This improves obstacle recognition accuracy and information integrity, assisting with subsequent path planning and navigation prompts.

[0048] The generation of obstacle classification information also involves image segmentation techniques such as target localization, bounding box fitting, or mask segmentation to ensure that the output classification information has spatial location attributes and instance-level independence. Although not directly reflected in the original step, image segmentation is an essential technical foundation for the object detection process and meets the overall technical implementation requirements.

[0049] By introducing pre-trained open-vocabulary object detection models, such as adopting a vision-language dual-tower architecture to process image and text inputs separately, cross-modal matching efficiency and category recognition accuracy can be improved. Alternatively, based on a joint embedding space design, images and text can be mapped into the same high-dimensional representation space, enhancing the ability to distinguish different categories of vocabulary in a visual context.

[0050] In practical applications, implementations can dynamically adjust the semantic description library content to suit different environments. For example, in healthcare scenarios, the semantic library can be expanded to include specialized categories such as wheelchairs, hospital beds, mobile medical devices, and nursing carts. In fintech scenarios, the semantic library can be expanded to include typical business area obstacle categories such as self-service terminals, mobile barriers, promotional displays, and cleaning tools.

[0051] In terms of system parameter configuration, similarity thresholds, confidence filtering criteria, and category priority sorting strategies can be set for specific environments to ensure that the output obstacle category information has high accuracy, low false alarm rate, and good semantic adaptability.

[0052] Example: In the field of intelligent assisted navigation, for head-mounted visual assistance devices worn by visually impaired or elderly people, the system obtains environmental image information, captures visual data of the wearer's front and surrounding areas in real time, performs open vocabulary object detection, and accurately identifies various dynamic or static obstacles. In specific applications, the system can distinguish between obstacle category information such as pedestrians, traffic cones, shared bicycles parked on the roadside, low structures, and construction isolation facilities, and provide differentiated navigation strategy support based on the characteristics of different categories. Ensure that obstacle category information has open expression and environmental adaptability, avoid the omission of new category recognition due to fixed label limitations, and ensure that the system continuously outputs accurate obstacle category information in changing environments such as complex urban roads, pedestrian streets, and underground passages. This improves the environmental perception integrity and navigation decision reliability of the assisted navigation system, and further reduces the path risk and probability of getting lost for visually impaired or elderly people during actual travel.

[0053] In the healthcare sector, a rehabilitation-assisted mobility system for visually impaired patients acquires environmental image information and performs open-vocabulary object detection, enabling real-time identification of various obstacles within hospital corridors. These obstacles include wheelchairs, mobile medical equipment, treatment carts, temporary ground signs, and other categories. Based on the detected obstacle categories, the system assists visually impaired patients in safely navigating medical facilities while adapting to environmental changes brought on by frequent adjustments to the hospital's internal spatial layout. This ensures patients maintain stable path perception in diverse scenarios, including diagnosis, rehabilitation, and transportation, minimizing collision risks and ensuring the practicality and safety of healthcare assistance systems.

[0054] In the field of financial technology, intelligent assisted navigation equipment deployed in large bank branches and comprehensive financial service halls can accurately distinguish various obstacle categories by acquiring environmental image information and performing open vocabulary object detection. These include ATMs, customer queue barriers, advertising displays, customer waiting area seating, and other obstacles specific to financial service scenarios. In detecting obstacle categories, the system can adapt to changes in branch layout and traffic density based on real-time business peaks. This ensures path safety and spatial recognition efficiency for visually impaired and elderly customers in financial service venues while handling transactions, finding their way, and waiting in queues. This improves the accessibility and aging-friendly experience of financial technology services and enhances the intelligent assisted navigation capabilities of financial venues.

[0055] This embodiment utilizes an open-vocabulary object detection mechanism, overcoming the technical limitations of traditional object detection, which relies on closed labeling systems, to achieve efficient and accurate identification of various obstacles in complex, dynamic, and ever-changing environments. Combined with the dynamic expansion of the semantic description library, the system can adapt to new category expression requirements in different application scenarios, improving the expressive richness and scenario adaptability of obstacle category information. This in turn enhances the intelligence of navigation path planning and user prompting, significantly reducing the risk of missed recognition and path deviation during navigation.

[0056] S30, fusing the obstacle category information and the environment depth information to generate spatial position information and direction information of the obstacle;

[0057] In this embodiment, the process of fusing obstacle category information with environmental depth information first requires clarifying the corresponding position range of the obstacle in the environmental image information based on the obstacle category information. The obstacle category information comes from open vocabulary object detection and is usually represented by a data structure containing category labels, position annotations, and morphological attributes. The position annotations in this information generally use a two-dimensional pixel coordinate system to reflect the specific position distribution of the obstacle in the image plane. In order to generate spatial position information, it is necessary to map the two-dimensional image position to the actual spatial coordinate system based on the depth data provided by the environmental depth information. The environmental depth information is collected by a depth sensor and usually includes distance values ​​for each pixel point in the environment. Based on this data, the actual position of the obstacle in three-dimensional space can be reconstructed.

[0058] In the implementation, the raw depth data from the environmental depth information is first parsed. This raw depth data is typically stored in a matrix format, where each element corresponds to a single pixel in the image plane, and its value represents the distance from that pixel's spatial location to the device. Combined with the obstacle's pixel coordinates, as indicated in the obstacle category information, the depth value of the corresponding pixel can be extracted to determine the obstacle's spatial distance from the user device.

[0059] To improve the comprehensibility of spatial distance information, the original depth data is usually converted into distance data based on step units. The step unit is set according to the user's actual stride length, commonly 0.7 meters. Through this conversion, the position of obstacles can be described in the form of "distance is X steps", enhancing the user's perception of spatial distance.

[0060] Furthermore, based on the pixel coordinate positions marked in the obstacle category information and combined with the size parameters of the environmental image information, the coordinates of the center point of the image plane can be determined. The center point of the image plane is usually located at the midpoint of the image width and height, which corresponds to the direction directly in front of the user device in space. By calculating the horizontal offset angle of the obstacle pixel coordinate position relative to the image plane center point coordinate, the horizontal position of the obstacle can be determined. The horizontal offset angle is usually referenced by the image coordinate system and converted into real space angle information in combination with the camera parameters. In order to facilitate user understanding, the horizontal offset angle is further converted into a clock azimuth interval value. The clock azimuth interval value is divided into several standard azimuths with reference to the traditional clock face position, such as "1 o'clock direction", "3 o'clock direction", etc., to help users quickly understand the obstacle direction information.

[0061] Based on obstacle category information, distance data, and clock azimuth interval values, the obstacle's spatial position and direction information are generated. Spatial position information primarily reflects the obstacle's specific location in actual three-dimensional space, while direction information, combined with the azimuth description, clarifies the obstacle's orientation relative to the user device. This process integrates category attributes, spatial position, and direction information to achieve a comprehensive and intuitive description of the obstacle.

[0062] The method for acquiring environmental depth information can be adjusted in different scenarios. Depth information can be derived from different technical paths, such as structured light sensors, time-of-flight lidar, and binocular vision reconstruction. The step unit setting can be dynamically adjusted based on individual user differences, improving the personalized adaptation of spatial distance information. The method for determining the center point of the image plane can be optimized based on the device installation angle and user height parameters to ensure the accuracy of directional information. Camera intrinsic parameter correction can be applied during the calculation of the horizontal offset angle to eliminate the effects of lens distortion and ensure the reliability of orientation information. The granularity of the clock orientation interval value division can be flexibly adjusted according to application requirements, for example, divided into 12 intervals in complex scenarios and 8 intervals in simplified mode to adapt to different users' perception abilities and usage preferences.

[0063] Example: In the healthcare field, for hospital indoor navigation assistance systems, the process of integrating obstacle category information and environmental depth information can accurately calibrate the position and direction information of medical equipment, waiting area seats, and temporarily placed items in hospital corridors. This helps visually impaired patients or elderly people with limited mobility to safely avoid obstacles, ensure path safety in different scenarios such as diagnosis, rehabilitation, and visiting, and reduce the risk of accidental collisions in the hospital environment.

[0064] In the field of financial technology, intelligent navigation systems for bank branches and financial service halls can generate real-time spatial location and direction information of obstacles by integrating obstacle category information with environmental depth information, accurately identifying the location and direction of queue isolation zones, advertising stands, self-service equipment, customer seats, etc., assisting visually impaired customers or elderly users to complete business route planning efficiently and safely in crowded environments during peak hours, and improving the level of barrier-free navigation in financial service venues.

[0065] In the field of intelligent assisted navigation, wearable navigation devices deployed in urban road scenarios integrate obstacle category information and environmental depth information, and can perceive the spatial position and direction information of vehicles, bicycles, temporary construction facilities, and pedestrian gathering areas on the road in real time. Combined with the standardized description form of step units and clock azimuth intervals, it helps visually impaired users or the elderly to clearly grasp the specific position and direction of surrounding obstacles, improve travel safety and route planning capabilities in complex traffic environments, and avoid the risk of directional errors or accidental collisions due to incomplete perception of complex environmental information.

[0066] Through the above-mentioned process of generating spatial position information and direction information, this embodiment combines abstract obstacle category information with specific spatial position and direction relationships, opens up the conversion link between two-dimensional image information and three-dimensional spatial perception, improves the integrity and intuitiveness of obstacle information, helps users more accurately understand the specific location and direction of obstacles in the surrounding environment, enhances the environmental perception capability of the auxiliary navigation system, and reduces the risk of collision during users' travel. It is particularly suitable for the safe navigation needs of visually impaired people or elderly users with decreased cognitive abilities in complex environments.

[0067] S40, generating a navigation instruction including scene description information, obstacle orientation and distance description information, and action suggestion information based on the spatial position information, direction information, and the environmental image information;

[0068] In this embodiment, the process of generating navigation instructions based on spatial position information, direction information, and environmental image information first requires clarifying the source and structure of the spatial position information and direction information. Spatial position information includes the three-dimensional coordinate data of the obstacle relative to the user's position, while direction information uses clock azimuth interval values ​​to specify the horizontal direction of the obstacle relative to the user's direction. Environmental image information provides comprehensive visual data of the current environment, including scene elements, background information, and dynamic change characteristics.

[0069] The first step in generating navigation instructions is to identify scene semantic labels in environmental image information. Scene semantic labels are extracted from the visual features of environmental image information based on a deep learning model. By comparing them with a preset semantic classification library, they can determine the typical scene category to which the current environment belongs, such as streets, supermarkets, hospital lobbies, financial service areas, etc. Scene semantic labels provide environmental background information for navigation instructions, enhancing the situational awareness of the instructions.

[0070] On this basis, the distance data in the spatial location information and the clock bearing interval values ​​in the direction information are integrated to generate obstacle position and distance description information. The distance data quantifies the spatial distance from the obstacle to the user using step units, while the clock bearing interval values ​​specify the horizontal direction of the obstacle. Combining the two generates a standardized position and distance expression, such as "Obstacle 2 steps away at 3 o'clock." This description information is intuitive, understandable, and provides operational guidance.

[0071] Furthermore, action recommendations are generated based on the distribution of obstacle position and distance descriptions. This distribution analyzes the relative positional relationships and dynamic changes of multiple obstacles in space, comprehensively assessing their impact on user path safety. The action recommendations are then tailored to the specific distribution and provide corresponding action suggestions, including detours, waiting, stopping, and adjusting direction. This ensures that users have clear and actionable navigation guidance in complex environments.

[0072] Ultimately, the scene description, obstacle location and distance descriptions, and action recommendations are combined to form complete navigation instructions. These instructions are output in a structured format and typically communicated to the user through speech synthesis, tactile feedback, or visual display. This helps users choose paths and avoid situations based on clear scene understanding, accurate obstacle location perception, and clear action recommendations.

[0073] In different application scenarios, the classification granularity of scene semantic labels can be adjusted according to needs. In urban block navigation scenarios, it can be refined to street types and traffic facility categories, and in indoor environments, it can be refined to functional area divisions. The generation method of obstacle orientation and distance description information can adjust the step unit based on user habits, support custom stride parameters, and adapt to the spatial perception needs of different groups of people. The generation logic of action suggestion information can be adjusted in real time based on the dynamic state of the environment. When dynamic obstacles gather, the output probability of waiting suggestions is enhanced, and detour path guidance is generated first when fixed obstacles exist. The output form of navigation instructions can be flexibly adapted through various modes such as voice playback, vibration prompts, and image enhancement display to improve information communication effects and user experience.

[0074] Example: In the healthcare field, the hospital navigation system generates navigation instructions that include scene description information, obstacle location and distance description information, and action suggestions. It can meet patients' navigation needs in different functional areas such as inpatient areas, outpatient areas, and testing areas. It can output voice prompts such as "You are in the lobby on the first floor of the inpatient building. There is a wheelchair 3 steps ahead at 2 o'clock. It is recommended to pass to the left." This helps visually impaired patients or elderly users complete hospital navigation efficiently and safely, reducing travel risks caused by equipment placement or crowds.

[0075] In the field of financial technology, the intelligent assistance system of bank branches is based on a navigation command generation mechanism. It can identify functional places such as customer waiting areas, self-service equipment areas, and financial service areas in real time during peak hours of the branch. Combined with crowd density and obstacle information, it outputs voice prompts such as "You are in the self-service area. There is a queue isolation belt 2 steps ahead at 3 o'clock. It is recommended to turn right and detour." This helps visually impaired users or elderly customers accurately identify service areas, avoid path selection errors caused by spatial cognitive impairment, and ensure the smoothness and safety of the business process.

[0076] In the field of intelligent assisted navigation, wearable navigation devices for urban roads can perceive complex environmental information such as streets, intersections, and bus stops in real time through navigation instructions generated based on spatial position information, direction information, and environmental image information. They can output prompt information such as "You are on the main street, there is a construction fence 2 steps ahead at 1 o'clock, it is recommended to wait or choose to detour on the left", assisting visually impaired people or the elderly to complete route planning and obstacle avoidance safely and efficiently, and improve their ability to travel independently in urban environments.

[0077] This embodiment, through the joint analysis and information fusion of spatial position information, direction information and environmental image information, can generate structured, contextualized and actionable navigation instructions in real time, helping users to clearly grasp scene categories, obstacle distribution and specific action suggestions in changing and complex environments, improving the environmental perception ability and instruction output accuracy of the navigation system, reducing path deviation and collision risks during travel, and enhancing the practicality and safety of the system.

[0078] S50: Update the navigation path according to the navigation instruction.

[0079] In this embodiment, the navigation path update is based on the navigation instructions, which include scene description information, obstacle orientation and distance description information, and action suggestion information. The process of updating the navigation path requires parsing various types of information in the navigation instructions, and dynamically adjusting the existing path planning results in combination with the user's current location and environmental status to ensure that the navigation path is highly matched with the actual environmental conditions, obstacle changes, and user needs.

[0080] First, the navigation instructions are parsed, including extracting the action suggestions contained therein. These action suggestions are typically presented in a structured format, clarifying the recommended route adjustment strategy, such as a detour, stop, or wait. This information serves as the basis for route adjustment decisions. Furthermore, the obstacle's location and distance information contained in the navigation instructions are combined to further determine the spatial distribution and relative orientation of the obstacles, assisting in determining whether the current route is obstructed or poses a safety hazard.

[0081] Subsequently, the system obtains the current user location information and the dynamic state of the environment. This information is collected in real time using high-precision positioning equipment, inertial navigation sensors, or visual positioning algorithms. Dynamic changes in the environment are monitored using depth cameras, lidar, or multimodal sensors, enabling the location and movement trajectory of dynamic obstacles and changes in environmental structure. These two types of information provide a real-time, accurate data foundation for dynamic adjustments to the navigation path.

[0082] Path update parameters are calculated based on action suggestion information, user location information, and the dynamic state of the environment. These parameters include key data such as new path node coordinates, steering angles, and path smoothing curvature. They guide navigation path replanning. Path update parameters are dynamically generated based on the real-time location of obstacles, the user's current position, the target location, and action suggestions, ensuring that the path adjustment process is highly consistent with the actual environment.

[0083] According to the path update parameters, the navigation path trajectory is adjusted and the path trajectory is recalculated through the path planning algorithm to ensure that the adjusted navigation path avoids obstacles, dynamic obstacles and unsafe areas in space, while meeting the path smoothness, accessibility and safety requirements, and generating a new navigation path.

[0084] After the navigation path adjustment is completed, verify whether the updated navigation path complies with the preset safety constraint strategy. The safety constraint strategy includes three parameters: minimum curvature radius, minimum obstacle spacing, and maximum slope angle. These parameters are used to ensure that the path turning radius meets the human gait and mobile device safety requirements, maintain a sufficient safety distance between the path and obstacles, and that the overall slope of the path adapts to the user's mobility range.

[0085] If the updated navigation path meets all safety constraint policies, the updated navigation path is output. This path can be fed back to the user through various methods such as a visual interface, voice broadcast, or vibration prompt, helping the user to perceive path change information in real time and ensure continuous safe navigation in a dynamically changing environment.

[0086] Path update parameter calculations can be based on traditional path planning algorithms (A), dynamic path update algorithms (D), or deep reinforcement learning-based path optimization methods. User location information can be acquired through GNSS satellite positioning, indoor Bluetooth positioning, or spatial positioning technologies based on visual SLAM. Environmental dynamics monitoring can be combined with visual sensors, millimeter-wave radar, or ultrasonic equipment to adapt to different environmental scenarios and user needs. The specific parameters of the safety constraint strategy can be flexibly set based on the usage scenario. In urban outdoor environments, the minimum spacing and maximum slope angle restrictions can be appropriately relaxed. In complex indoor environments, the safety constraint level can be increased to meet higher safety requirements.

[0087] Example: In the healthcare field, the hospital's navigation system monitors the patient's location and the dynamic changes of the surrounding environment in real time, and updates the navigation path based on navigation instructions. It can dynamically adjust the path in wheelchairs, medical carts, or crowded areas, and output optimized travel routes to assist patients in smoothly reaching the examination department or ward area, avoiding the risk of path blockage caused by equipment stacking or dense crowds, and ensuring the efficient and safe movement of patients in the hospital.

[0088] In the field of financial technology, in bank branches or financial service halls, the path update mechanism based on navigation instructions can perceive customer distribution, waiting queues and temporary obstacle locations in real time, dynamically optimize customer travel paths, and output prompt information such as "Please detour from the right" to avoid path obstructions caused by factors such as customer gathering and temporary isolation facility layout, thereby ensuring the smoothness and safety of customer navigation in the financial service environment.

[0089] In the field of intelligent assisted navigation, wearable navigation devices dynamically update routes in combination with navigation instructions. They can sense road construction, temporary traffic controls, or dynamic obstacle changes in urban road environments, adjust route guidance information in real time, and output voice prompts such as "The road ahead is closed, it is recommended to turn right and detour." This ensures that visually impaired people or the elderly can travel independently and safely in complex and changing urban environments.

[0090] This embodiment dynamically analyzes navigation instructions, combines real-time user location and environmental status, and dynamically updates the navigation path. This can effectively respond to environmental changes, obstacle movement, and path blockages, ensure the continuity, feasibility, and safety of the navigation path, enhance the intelligent response capability and environmental adaptability of the navigation system, reduce the risk of users getting lost, colliding, or deviating from the path in complex environments, and improve travel safety and navigation accuracy.

[0091] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as intelligent assisted navigation, financial technology, and healthcare. A method, apparatus, device, and medium for generating navigation instructions and updating a path are disclosed, including: obtaining environmental image information and environmental depth information, performing open vocabulary object detection to obtain obstacle category information, fusing the obstacle category information and environmental depth information, generating spatial position information and direction information of the obstacle, and generating navigation instructions containing scene description information, obstacle orientation and distance description information, and action suggestion information based on the spatial position information, direction information, and environmental image information, and updating the navigation path. By fusing obstacle category information with environmental depth information and combining it with environmental image information to generate semantically-level obstacle recognition results, and outputting navigation instructions containing orientation and distance, the present invention can improve the accuracy of obstacle perception and the practicality of navigation guidance, thereby enhancing navigation safety in complex environments.

[0092] In one embodiment, the above step S10 includes:

[0093] S101, capturing initial environment image information and initial environment depth information through a visual depth sensor;

[0094] S102, performing spatiotemporal synchronization processing on the initial environment image information and the initial environment depth information to generate synchronized environment image data and synchronized environment depth data;

[0095] S103, detecting a dynamic change state in the environment based on the synchronized environment image data and the synchronized environment depth data;

[0096] S104, marking the motion trajectory of the moving object according to the dynamic change state;

[0097] S105, extracting scene illumination features of the synchronized environment image data;

[0098] S106, predicting a spatial position offset based on the motion trajectory of the moving object, and correcting corresponding coordinate values ​​of the synchronized environment depth data based on the spatial position offset to generate dynamically updated depth data;

[0099] S107, fusing the dynamically updated depth data with scene illumination features to generate illumination-compensated depth data, and using the illumination-compensated depth data as environmental depth information;

[0100] S108 : Marking the predicted location area of ​​the dynamic obstacle in the synchronized environment image data according to the motion trajectory of the moving object to generate environment image information.

[0101] In this embodiment, the process of obtaining environmental image information and environmental depth information involves multi-level image acquisition, data synchronization, dynamic monitoring and information fusion. The purpose is to accurately reflect the environmental spatial structure and real-time changing status in a complex dynamic environment, and ensure that the data basis for subsequent obstacle detection, navigation command generation and other links is timely and accurate.

[0102] First, the initial environment image information and initial environment depth information are captured by a visual depth sensor. The visual depth sensor includes but is not limited to a stereo vision camera, a structured light sensor, a time-of-flight camera or a multimodal imaging device, which can simultaneously collect two-dimensional image information of the environment and the corresponding three-dimensional depth data. The initial environment image information reflects the optical reflection characteristics and spatial texture of the environment surface, and the initial environment depth information describes the spatial distance between each position point in the environment and the sensor.

[0103] After data collection is completed, spatiotemporal synchronization processing needs to be performed on the initial environmental image information and the initial environmental depth information. The spatiotemporal synchronization processing maintains strict consistency between the image data and depth data from different sensing channels in the time and space dimensions through timestamp comparison, data frame reordering and spatial coordinate alignment, and generates synchronized environmental image data and synchronized environmental depth data to ensure that the environmental status reflected by the two types of data is the result at the same time and the same spatial perspective, avoiding spatial position deviation caused by acquisition delays or equipment errors.

[0104] Based on synchronized environmental image data and synchronized environmental depth data, dynamic change state detection in the environment is performed. Dynamic change state detection uses background modeling, motion vector analysis or multi-frame difference technology to identify spatial areas in the environment with position movement, shape change, or new or disappeared areas. Dynamic change states include but are not limited to the appearance of moving obstacles, changes in environmental lighting conditions, and the temporary addition of structural obstructions, ensuring that the system can perceive environmental changes in real time.

[0105] According to the detected dynamic change state, the motion trajectory of the moving object is marked. In the process of motion trajectory marking, time series information is combined to analyze the spatial position change trend of the moving object in a continuous time period. Multi-target tracking algorithm, Kalman filtering or visual-inertial fusion technology are used to continuously track the path, speed and direction of the moving object to form structured motion trajectory information. The motion trajectory reflects the dynamic behavior pattern of the moving object in the environment.

[0106] Based on synchronized environmental image data, scene lighting features are extracted. Scene lighting feature extraction quantifies the overall lighting intensity, distribution characteristics, and local light and dark relationships of the environment through image brightness statistical analysis, shadow area segmentation, and light source direction estimation. The lighting feature information is used for subsequent data compensation and perception result optimization to reduce perception errors caused by uneven lighting or strong reflections.

[0107] Furthermore, the spatial position offset is predicted based on the motion trajectory of the moving object. The spatial position offset prediction is combined with historical trajectory data, current movement trends and environmental constraints to infer the possible position changes of the moving object in a short period of time in the future. The corresponding coordinate values ​​of the synchronized environmental depth data are corrected based on the prediction results to ensure that the environmental depth information maintains the accuracy of the spatial position under the interference of the moving object, and dynamically updated depth data is generated. The dynamically updated depth data reflects the real-time spatial structure of the environment under dynamic changes.

[0108] The dynamically updated depth data is fused with the scene illumination features. The fusion process improves the stability and accuracy of the depth data under different lighting conditions through multi-channel data alignment, weighted compensation and error correction, and generates illumination-compensated depth data. The illumination-compensated depth data serves as the final environmental depth information for subsequent obstacle positioning and navigation path generation.

[0109] Combined with the motion trajectory of the moving object, the predicted position area of ​​the dynamic obstacle is marked in the synchronized environmental image data. The predicted position area is determined based on the current motion state of the moving object and the prediction result of the spatial position offset. The position, range and dynamic change trend of the dynamic obstacle in the image information are clarified by using regional division, spatial marking and visual highlighting, and the environmental image information is generated. The environmental image information contains the dynamic obstacle position mark and the overall visual characteristics of the scene, providing accurate data support for subsequent visual detection and navigation command generation.

[0110] Through multi-level information collection, synchronization and dynamic change analysis, this embodiment enables the system to accurately reflect the spatial structure and dynamic change state of the environment in real time, overcome perception errors caused by sensor asynchrony, changes in ambient lighting or interference from moving obstacles, improve the spatial accuracy and temporal consistency of environmental image information and environmental depth information, enhance the ability to perceive and understand dynamic environments, ensure the real-time, stability and reliability of the navigation system in complex environments, effectively support subsequent obstacle detection, path planning and navigation instruction generation processes, and improve the intelligence level and environmental adaptability of the overall navigation system.

[0111] In one embodiment, the above step S20 includes:

[0112] S201, extracting a visual feature vector of the environmental image information;

[0113] S202, matching the visual feature vector with a text embedding vector in an open vocabulary semantic description library, and determining a similarity score between the visual feature vector and the text embedding vector;

[0114] S203, determining a category label of an obstacle that is not a predefined category according to the similarity score;

[0115] S204, analyzing the texture features and reflection features of the visual feature vector to determine the material properties of the obstacle surface;

[0116] S205 , fusing the obstacle category label and material attributes to generate a composite obstacle description, and using the composite obstacle description as obstacle category information.

[0117] In this embodiment, the process of open vocabulary object detection on environmental image information uses multi-dimensional visual feature extraction and semantic matching technology to identify various types of obstacles in the environment, and combines material properties to generate high-precision obstacle category information, thereby improving the recognition ability of complex environments.

[0118] First, the visual feature vectors of the environmental image information are extracted. This extraction is based on the multi-layered information structure of the image data, primarily including but not limited to color, texture, edge contour features, and spatial structure. Common methods include convolutional neural network (CNN) feature extraction, SIFT, HOG, or multi-scale image pyramid analysis. Through feature extraction, the environmental image information is converted into a structured, high-dimensional feature vector that reflects the spatial morphology, visual texture, and appearance attributes of each object region in the image.

[0119] After obtaining the visual feature vector, it is matched against text embeddings in the Open Lexicon Semantic Description Library. Text embeddings are generated based on large-scale natural language models and reflect the semantic spatial distribution of different words, phrases, or category labels in the language description. The matching process uses multimodal alignment algorithms, such as the CLIP framework or cross-modal similarity calculations, to measure the proximity of the visual feature vector to each text embedding in a unified semantic space, ensuring a highly accurate association between visual data and linguistic semantics.

[0120] Based on the similarity score between the visual feature vector and the text embedding vector, the system determines whether there are obstacle category labels in the environmental image that fall outside the traditional fixed category list. These include newly emerged transportation facilities, temporary obstacles, or rare objects in a specific scene. The system automatically generates corresponding obstacle category labels through similarity threshold judgment and a dynamic label expansion mechanism, enhancing its adaptability to open environments.

[0121] Furthermore, the texture features and reflection features of the visual feature vector are analyzed. The texture features include surface pattern, structural repeatability and roughness information. The reflection features reflect the object surface's ability to reflect light and the reflection pattern. Accurate material characterization information is obtained through gradient direction statistics, light intensity distribution analysis and multispectral imaging technology.

[0122] The obstacle category label is fused with the material attributes. The fusion process adopts a multi-dimensional information integration strategy, combining the semantic information of the category label with the physical characteristics of the material attributes to generate a composite obstacle description. The composite obstacle description not only contains classification information at the category level, but also includes material, optical and structural information at the physical attribute level, improving the completeness and accuracy of the obstacle category information. Finally, the composite obstacle description is output as the obstacle category information, providing multi-dimensional information support for subsequent spatial positioning, direction recognition and navigation decision-making.

[0123] This embodiment, through matching visual feature vectors with an open vocabulary semantic description library, enables the system to break through the traditional object detection's reliance on fixed categories, dynamically identify various obstacles in open environments, and enhance the adaptability and flexibility of object detection. Combined with material properties, the generated obstacle category information not only has semantic differentiation capabilities but also includes physical properties such as material and reflection, enhancing the ability to understand environmental complexity and improving the reliability and intelligence of subsequent path planning and navigation command generation. Through this process, the system's overall recognition accuracy and environmental adaptability are significantly improved, effectively coping with the interference of dynamic changes and uncertain factors in open environments, and ensuring the stability and security of the navigation process.

[0124] In one embodiment, the above step S30 includes:

[0125] S301, parsing original depth data in the environmental depth information, and converting the original depth data into distance data based on a step unit;

[0126] S302, identifying pixel coordinate positions of obstacles in the environment image information;

[0127] S303, determining the coordinates of the center point of the image plane according to the size parameters of the environmental image information;

[0128] S304, determining a horizontal offset angle of the pixel coordinate position relative to the coordinates of the center point of the image plane, and converting the horizontal offset angle into a clock azimuth interval value;

[0129] S305: Generate spatial position information and direction information of the obstacle based on the obstacle category information, the distance data, and the clock azimuth interval value.

[0130] In this embodiment, the fusion of obstacle category information and environmental depth information aims to determine the spatial position and orientation attributes of obstacles based on multi-source perception data, providing structured, quantified environmental description information for subsequent navigation decisions. First, the raw depth data in the environmental depth information is parsed. Raw depth data refers to numerical information obtained by the depth sensing device that reflects the distance from each point in the environment to the sensor. It is usually measured in millimeters or centimeters and can be in the form of a two-dimensional depth map, a three-dimensional point cloud, or a multi-channel matrix. The parsing process includes data decoding, coordinate system conversion, and spatial structure reconstruction to ensure that the depth information has spatial positioning accuracy and the ability to express physical distances.

[0131] The parsed raw depth data is converted into distance data based on step units. The introduction of step units is in line with user cognitive habits and is particularly suitable for assisted walking or navigation systems. The determination of the step unit is based on the user's average step length parameter. The step length parameter can be generated through the user's historical walking data, sensor dynamic measurement, or system preset standards. The common step unit is set to 0.7 meters or dynamically adjusted based on individual differences. During the conversion process, the system maps the raw depth value to a step number expression through division or scaling, making it easier for users to understand and execute navigation commands.

[0132] Identifying the pixel coordinates of obstacles in environmental image information involves extracting the two-dimensional position parameters of obstacle targets in the image based on visual detection algorithms. Specific methods can use deep learning object detection networks, bounding box localization, or semantic segmentation techniques to ensure pixel-level accuracy in the positioning results. Pixel coordinates typically include horizontal and vertical coordinates, which are used to describe the relative position of obstacles in the image plane.

[0133] The coordinates of the center point of the image plane are determined based on the size parameters of the environmental image information. The size parameters include the width and height of the image. The center point coordinates are obtained through mathematical calculation. The horizontal center point is half of the width, and the vertical center point is half of the height. The center point serves as a direction reference to ensure that subsequent offset angle calculations are based on a unified reference system, thereby improving the standardization of spatial orientation expression.

[0134] Determine the horizontal offset angle of the pixel coordinate position relative to the coordinates of the center point of the image plane. The horizontal offset angle reflects the positional deviation of the obstacle relative to the observer's front within the horizontal viewing angle. It is usually calculated using inverse trigonometric functions or angle mapping relationships and can be expressed in degrees or radians. Convert the horizontal offset angle to a clock azimuth interval value. The clock azimuth interval value is based on the spatial division of an analog clock face, dividing the 360-degree horizontal viewing angle into 12 azimuth intervals. For example, the front is 12 o'clock, the right front is 3 o'clock, and the left front is 9 o'clock. This simplifies the representation of spatial orientation and improves user understanding and response efficiency.

[0135] Based on obstacle category information, distance data, and clock azimuth interval values, the spatial position and direction information of the obstacle are generated. The spatial position information includes the relative position, distance, and category attributes of the obstacle in the user reference frame. The direction information expresses the directional distribution of the obstacle through the clock azimuth interval values, forming a structured multi-dimensional environmental description data to support the precise execution of subsequent navigation decisions, obstacle avoidance strategies, and path planning.

[0136] This embodiment integrates obstacle category information with environmental depth information to achieve a joint expression of the obstacle's spatial position and directional attributes. This overcomes the limitations of a single information source in environmental cognition and improves spatial perception accuracy. Step unit conversion lowers the threshold for data comprehension, enhancing the feasibility and user acceptance of navigation information. Pixel-level obstacle positioning and image center reference calculation ensure standardization and consistency in azimuth expression. Clock azimuth interval mapping further optimizes the intuitiveness and language expression efficiency of azimuth information, overall improving the real-time, accuracy, and user adaptability of the navigation system in dynamic environments, effectively supporting safe navigation needs in highly complex environments.

[0137] In one embodiment, the above step S40 includes:

[0138] S401, identifying scene semantic tags of the environmental image information and generating scene description information;

[0139] S402, fusing the distance data in the spatial position information and the clock azimuth interval value in the direction information to generate obstacle azimuth and distance description information;

[0140] S403, generating action strategy information based on the distribution state of the obstacle orientation and distance description information;

[0141] S404: Combining the scene description information, obstacle position and distance description information, and action strategy information to form a navigation instruction.

[0142] In this embodiment, navigation instructions including scene description information, obstacle orientation and distance description information, and action suggestion information are generated based on spatial position information, direction information, and environmental image information, which involves multi-level data fusion and semantic information extraction, aiming to form structured, understandable, and guiding navigation output. First, the scene semantic labels of the environmental image information are identified. The environmental image information is the environmental image data obtained by the visual sensor. The scene semantic labels are used to describe the type, spatial function, or usage scenario of the current environment. The recognition process can be achieved through a deep learning image classification network, a visual language joint model, or semantic segmentation technology. Common scene semantic labels include categories such as streets, indoor corridors, public squares, and shopping mall interiors. Scene semantic labels provide spatial environment background information for generating navigation content, and assist in judging the walkable area, obstacle distribution patterns, and potential safety risks.

[0143] The distance data in the spatial location information is fused with the clock direction interval values ​​in the direction information to generate obstacle position and distance description information. The distance data in the spatial location information is derived from the quantitative calculation of the obstacle position and is usually expressed in units of steps or physical distance. The clock direction interval values ​​in the direction information are based on standardized direction divisions, expressing the obstacle direction as interval values ​​analogous to a clock dial, such as 1 o'clock, 3 o'clock, and 9 o'clock. The fusion process integrates the distance and direction information of the obstacle into an integrated description, generating structured, intuitive, and easy-to-understand obstacle direction and distance description information, such as "there is an obstacle 3 steps away at 2 o'clock" or "there is an obstacle 2 meters ahead."

[0144] Based on the distribution of obstacle location and distance information, action strategy information is generated. This distribution includes the number, density, and dynamic change trends of obstacles at different locations and distances. By analyzing this distribution and combining it with established safety policies or path planning rules, the system generates action strategy information that is executable. These action strategy information provides specific navigation guidance, such as suggestions for detours, stopping and waiting, swerving to the left, or slowing down, ensuring safe and efficient path selection in complex environments.

[0145] The scene description information, obstacle orientation and distance description information, and action strategy information are combined to form navigation instructions. Navigation instructions are structured, user-oriented navigation information outputs, which usually include environmental background descriptions, obstacle locations and risk warnings, and specific behavioral suggestions. The output can be in the form of voice broadcasts, graphical interface displays, or multimodal interactions. Navigation instructions improve users' understanding of the environment and the efficiency of navigation behavior execution by integrating multi-source information, thereby enhancing the practicality and reliability of the auxiliary navigation system.

[0146] This embodiment generates navigation instructions containing multi-dimensional environmental information based on spatial position information, direction information, and environmental image information. The system achieves a deep integration of environmental perception, obstacle location, and behavioral decision-making, breaking the limitation of traditional navigation that only provides vague direction information, and improving the structured, semantic, and user-friendly navigation content. Scene semantic tags provide environmental context, obstacle position and distance descriptions quantify risk locations, and action strategy information clarifies navigation suggestions. This overall improves the system's navigation accuracy and user interaction experience in dynamic and complex environments, effectively ensuring walking safety and path optimization.

[0147] In one embodiment, the above step S50 includes:

[0148] S501, parsing the action strategy information in the navigation instruction;

[0149] S502, obtaining the current user location information and the dynamic change status of the environment;

[0150] S503, determining path update parameters based on the action strategy information, current user location information, and dynamic environmental change status;

[0151] S504, adjusting the current navigation path trajectory according to the path update parameters to generate an updated navigation path;

[0152] S505, verifying whether the updated navigation path complies with a preset security constraint policy;

[0153] S506: If the updated navigation path complies with the preset security constraint policy, output the updated navigation path.

[0154] In this embodiment, the process of updating the navigation route based on navigation instructions first requires parsing the action strategy information contained in the navigation instructions. Navigation instructions are multi-dimensional structured data generated by combining environmental information. Action strategy information is the core content used to guide user movement behavior. It is usually presented in the form of text, code, or parameter sets. Action strategy information may include detour direction, forward speed, stop instructions, path deviation angle, and other content. The parsing process accurately extracts action strategy information through data structure analysis, semantic analysis, or protocol decoding, providing decision-making basis for subsequent route updates.

[0155] Acquire the current user's location information and the dynamic state of the environment. User location information includes current location coordinates, heading angle, and movement speed. The dynamic state of the environment indicates real-time changes in factors related to the user's navigation path, commonly including obstacle movement trajectories, ambient lighting conditions, and temporary road closures. This acquisition method relies on a multi-sensor fusion positioning system, a visual analysis module, or an environmental monitoring device to achieve real-time and accurate perception of the user and environmental status.

[0156] Path update parameters are determined based on action strategy information, user location information, and the dynamic state of the environment. Path update parameters are control variables used to adjust the navigation path. Typical examples include path offset, direction adjustment angle, path replanning trigger flag, speed limit, etc. The parameter determination process must comprehensively consider the user's current state, environmental changes, and navigation strategy requirements to ensure that path updates are real-time, accurate, and secure.

[0157] The current navigation path trajectory is adjusted based on the path update parameters to generate an updated navigation path. The navigation path trajectory is the spatial curve data that describes the user's intended movement route. The adjustment process involves path recalculation, path smoothing optimization, or dynamic obstacle avoidance processing. The result is an updated navigation path that reflects the path's coherence, obstacle avoidance, and feasibility.

[0158] Verify that the updated navigation path complies with pre-set safety constraints, including minimum curvature radius, minimum obstacle spacing, and maximum slope angle. These constraints define the basic safety boundaries of the navigation path. The minimum curvature radius prevents sharp turns that could cause imbalance or falls, the minimum obstacle spacing ensures a safe distance between the user and obstacles, and the maximum slope angle limits the path's tilt to prevent slips or falls. Verification is accomplished through calculation of path geometry, spatial distance measurement, and terrain analysis to ensure that the updated navigation path meets walking safety requirements in terms of structure and environmental adaptation.

[0159] If the updated navigation path meets the preset safety constraint strategy, the updated navigation path is output. The output form can be path data transmitted to the navigation execution module, visual interface display or voice broadcast, ensuring that the user or navigation device performs the navigation task based on the safety-verified path.

[0160] This embodiment dynamically generates and optimizes navigation paths by parsing action strategy information contained in navigation instructions, combining it with real-time user location and environmental changes. This system effectively adapts to environmental changes and user mobility, ensuring the timeliness, rationality, and safety of navigation paths. A safety constraint verification mechanism is embedded in the path update process to ensure that path adjustments do not introduce new security risks. This overall approach enables real-time updates and dynamic optimization of navigation paths, enhancing the path adaptability and navigation safety of the auxiliary navigation system in complex environments.

[0161] In one embodiment, after the above step S50, the method further includes:

[0162] S601, parsing obstacle position and distance description information in the navigation instruction, and extracting distance data from the obstacle position and distance description information;

[0163] S602, identifying a clock azimuth description value in the obstacle azimuth and distance description information;

[0164] S603, establishing an inverse proportional correlation relationship between the distance data and the vibration intensity, and adjusting the intensity change rate of the inverse proportional correlation relationship according to the user movement speed parameter;

[0165] S604, calculating a basic vibration intensity level using the inverse proportional relationship, and performing brightness compensation on the basic vibration intensity level based on an ambient light intensity parameter to generate a final tactile feedback intensity level;

[0166] S605: Build a mapping strategy for clock orientation intervals to physical feedback positions, and map the left orientation interval to the left head feedback position according to the mapping strategy, map the right orientation interval to the right head feedback position according to the mapping strategy, and map the front orientation interval to the central forehead feedback position according to the mapping strategy;

[0167] S606, determining a target feedback position based on matching the corresponding strategy with the clock orientation description value;

[0168] S607: Generate a trapezoidal vibration waveform, set an amplitude parameter of the trapezoidal vibration waveform to be proportional to the tactile feedback intensity level, set a duration parameter of the trapezoidal vibration waveform to be inversely correlated with the obstacle distance value, and configure a repetition frequency parameter of the trapezoidal vibration waveform to be positively correlated with the degree of dynamic change of the environment;

[0169] S608, outputting the trapezoidal vibration waveform through a driving unit corresponding to the target feedback position;

[0170] S609, during the output of the trapezoidal vibration waveform, when it is detected that the obstacle distance value is less than the safety distance threshold, the full feedback position synchronization vibration state is activated, the vibration intensity is increased to the highest level, and a continuous pulse waveform is used to replace the trapezoidal waveform to maintain the vibration output until the user position status is updated.

[0171] In this embodiment, after the navigation path is updated, the obstacle position and distance description information in the navigation instruction needs to be further parsed. The obstacle position and distance description information is an important component of the navigation instruction structure, primarily representing the spatial position of the obstacle relative to the user. The distance data is a parameter reflecting the spatial distance from the obstacle to the user, and the unit can be set to meters, steps, or other length measurement standards as needed. The clock position description value uses an angular representation similar to the divisions of a clock face to intuitively represent the horizontal position of the obstacle, such as 1 o'clock or 3 o'clock, to facilitate user understanding and perception of the obstacle's location.

[0172] Extracting distance data and clock orientation values ​​provides the foundational input for subsequent tactile feedback. Subsequently, an inversely proportional relationship must be established between distance data and vibration intensity. This inverse proportional relationship means that the closer the obstacle, the higher the vibration intensity, thereby increasing the user's awareness of nearby obstacles. To adapt to different user movements, this inverse proportional relationship must also be dynamically adjusted based on the user's movement speed parameters. Speed ​​parameters can be obtained in real time through an inertial measurement unit, gait analysis, or other motion sensing modules. This adjusted intensity change rate prevents excessive vibration frequency at high speeds or insufficient vibration at low speeds, improving the adaptability and accuracy of tactile feedback.

[0173] Based on this inverse proportional relationship, a base vibration intensity level is calculated. This level reflects the vibration amplitude level under standard conditions. Considering that user sensitivity varies across different lighting environments, the system performs brightness compensation on the base vibration intensity level based on ambient light intensity parameters. Light intensity parameters can be obtained through light sensors or image analysis. This compensation process ensures that the vibration signal is appropriately enhanced in low-light environments, improving the tactile feedback effect in low-light conditions. Ultimately, the system generates a tactile feedback intensity level that is appropriate for the current environment and user state.

[0174] To achieve accurate orientation feedback, a strategy was developed to map clock orientation intervals to physical feedback locations. Clock orientation intervals are divided based on a circular scale. For example, the 0 to 4 o'clock direction is defined as the left orientation interval, the 4 to 8 o'clock direction is defined as the right orientation interval, and the 8 to 12 o'clock direction is defined as the front orientation interval. These are mapped to the left head feedback position, the right head feedback position, and the center forehead feedback position, respectively. This strategy uses spatial mapping to logically associate obstacle orientation information with the corresponding body parts, making it easier for users to intuitively perceive obstacle locations.

[0175] Based on the clock orientation description value matching corresponding strategy, the final target feedback position is determined. The target feedback position is the specific location point where vibration feedback is executed. Combined with the actual hardware structure, it usually corresponds to the independent vibration unit on the wearable device.

[0176] In order to provide a vibration signal with a gradual perception effect, the system generates a trapezoidal vibration waveform. The trapezoidal waveform has the characteristics of clear undulating boundaries and distinct vibration perception levels, which is suitable for tactile information transmission in the field of assisted navigation. The amplitude parameter of the trapezoidal vibration waveform is set to be proportional to the tactile feedback intensity level, ensuring that the closer the obstacle, the larger the amplitude and the more obvious the prompt effect. The duration parameter of the vibration waveform is set to be inversely correlated with the obstacle distance value. The closer the distance, the shorter the vibration duration, creating a high-frequency and intensive feedback feeling. The repetition frequency parameter of the vibration waveform is positively correlated with the degree of dynamic change of the environment. The repetition frequency is increased in a dynamic environment and moderately reduced in a static environment, thereby improving the real-time nature of information transmission.

[0177] The output of the trapezoidal vibration waveform is executed by the driving unit corresponding to the target feedback position. The driving unit can be implemented using structures such as piezoelectric ceramics, micro motors, and magnetostrictive elements to ensure efficient and accurate transmission of physical vibration signals.

[0178] During the vibration waveform output process, the obstacle distance value is detected in real time. If the detected distance value is less than the safe distance threshold, the system activates the full-feedback position synchronized vibration state, that is, all feedback units work simultaneously, forming high-intensity multi-point vibration to enhance the warning effect. At the same time, the vibration intensity is increased to the highest level, and a continuous pulse waveform is used instead of a trapezoidal waveform. The continuous pulse waveform has the characteristics of high frequency, strong rhythm, and short period, which is suitable for warning scenarios where danger is approaching. This state is maintained until the user's position status is updated. The position status update can be judged in real time by the positioning system, inertial navigation unit, etc., ensuring that the emergency warning automatically ends after the risk is eliminated, improving system safety and user comfort.

[0179] Example description: In the field of intelligent assisted navigation, especially in application scenarios where visually impaired people wear smart glasses to navigate in dynamic environments, real-time navigation support for the entire process can be achieved by integrating environmental perception, obstacle detection, path planning and tactile feedback modules.

[0180] In actual use, users wear multi-sensor smart glasses. The system first utilizes the visual depth sensor mounted on the front of the glasses to synchronously collect environmental image and depth information for the current road scene. The smart glasses continuously capture image and depth information at a high frame rate, ensuring the timeliness and integrity of the data source. The synchronously collected data undergoes spatiotemporal synchronization processing in a local or cloud-based computing unit. This synchronization generates synchronized environmental image and depth data through timestamp alignment and coordinate unification, ensuring accurate and consistent data for subsequent analysis.

[0181] Based on the synchronized environmental data, the system calls the dynamic target detection module to track and identify objects in the environment. It calculates the dynamic change state of the environment through continuous image frames and determines whether there are currently moving obstacles or sudden traffic elements. When a moving object is identified, the system uses a motion trajectory prediction algorithm to calculate the spatial position offset of the moving object in a short period of time, and based on the prediction results, it makes real-time corrections to the coordinates of the synchronized environmental depth data to generate dynamically updated depth data. To adapt to scenes with changing lighting, the system further extracts lighting feature parameters from the current environmental image data and applies them to the depth data compensation model to perform lighting compensation on the dynamically updated depth data to generate lighting-compensated depth data, ensuring the stability and perception accuracy of the depth data.

[0182] On synchronized image data, the system annotates the predicted location of moving objects using motion trajectories, annotating dynamic obstacle information in the current environment image in real time. This dynamic annotation information is used in the subsequent generation of navigation instructions. Using an open vocabulary object detection model, the system extracts features from the environment image data, obtains visual feature vectors, and performs similarity matching between these visual feature vectors and text embedding vectors in a pre-established open vocabulary semantic description library. Using a cross-modal matching mechanism, the system automatically detects and identifies obstacle labels that are not in predefined categories. Simultaneously, through image texture analysis algorithms and optical reflection modeling, the system extracts texture and reflection features from obstacle surfaces, infers obstacle material properties, and fuses the identified obstacle category labels with the material properties to form a composite obstacle description, which serves as input for subsequent spatial position information calculations.

[0183] The system further analyzes the raw depth data in the environmental depth information, calls the step unit conversion model, and converts the raw depth data into distance data based on the user's average step length, realizing the user's personalized distance expression. Subsequently, the system automatically calculates the coordinates of the center point of the image plane based on the pixel coordinate position of the obstacle in the environmental image, combined with the image size parameters, and calculates the horizontal offset angle of the obstacle relative to the image center through geometric relationships. The offset angle is divided into the standard clock azimuth interval, and the obstacle's azimuth information is output. The system associates the obstacle category information, the distance data based on the step unit, and the clock azimuth interval value, and generates the spatial position information and direction information of the obstacle in real time as the basic input for generating navigation instructions.

[0184] Based on spatial location information, direction information, and environmental image information, the system uses a scene recognition module to extract scene semantic tags from the image and generate scene description information to characterize the current road or building type. The system combines obstacle distance data and orientation information to generate structured obstacle orientation and distance description information. By analyzing the current obstacle distribution status and assessing the feasibility of the user's passage path, the system automatically generates action strategy information including avoidance, stopping, or detour. The system combines scene description information, obstacle orientation and distance description information, and action strategy information to form navigation instructions that include semantic expression and quantitative guidance.

[0185] The system parses the action strategy information in the navigation instructions, obtains the user's current location information and the dynamic state of the environment, combines the user's location information with dynamic environmental parameters, determines the path update parameters, adjusts the current navigation path trajectory, and generates an updated navigation path. Based on the preset safety constraint strategy, the system verifies whether the updated navigation path meets the minimum curvature radius, minimum obstacle spacing, and maximum slope angle requirements. Once the path is verified, it outputs a navigation path that meets safety standards.

[0186] After the navigation path is updated, the system parses the obstacle position and distance descriptions in the navigation instructions, extracts the distance data and the clock position description value, and establishes an inversely proportional relationship between obstacle distance and vibration intensity. This relationship adjusts the intensity change rate in real time based on the user's movement speed to ensure an appropriate vibration feedback rhythm for different gaits. The base vibration intensity level is compensated for by the light intensity parameter to generate the final tactile feedback intensity level. Based on a mapping strategy between clock position intervals and physical feedback locations, the system maps the position intervals to the feedback modules on the left, right, and forehead of the smart glasses device. After matching the target feedback location, the system generates a trapezoidal vibration waveform with an amplitude proportional to the tactile feedback intensity level, a duration inversely proportional to the obstacle distance, and a repetition frequency proportional to the degree of environmental dynamics. The trapezoidal waveform is output by the driver unit at the target feedback location. If the obstacle distance is detected below the safe distance threshold, the system immediately activates the fully feedback position-synchronized vibration state, increasing the vibration intensity to the highest level and switching to a continuous pulse waveform. This output continues until the user's position status is updated, ensuring timely notification of impending danger. The entire navigation process realizes a closed-loop information flow from environmental perception, target detection, path planning to tactile feedback, providing a real-time, accurate and personalized assisted navigation experience.

[0187] In the field of medical health, especially for the application scenarios of travel assistance equipment for visually impaired patients, a complete safe navigation system is built through environmental perception, obstacle detection, dynamic path adjustment and multi-dimensional feedback mechanism to effectively alleviate the access barriers of visually impaired people in medical institutions, rehabilitation centers or public medical places.

[0188] In practice, patients wear an intelligent assisted navigation device equipped with multimodal sensors. The system first acquires real-time environmental image and depth information through an integrated visual depth sensor. Image information encompasses the visual content of the surrounding space, while depth information reflects the spatial distance and relative position of obstacles. After data collection, the device invokes a synchronization calculation module to perform spatiotemporal alignment of the image and depth information, ensuring consistency in the timestamps and spatial references of the different data sources. This generates synchronized environmental image and depth data, laying the foundation for dynamic change detection.

[0189] In a medical environment, where personnel flow frequently, the system performs dynamic change state detection based on synchronized data, accurately identifies the trajectory information of moving patients, medical staff, and auxiliary equipment, generates motion trajectories in real time, and annotates them in the synchronized environmental image data. Combined with the illumination feature extraction module of image data, the system perceives the lighting conditions of different areas of hospital corridors, outpatient areas, or inpatient departments in real time to prevent navigation accuracy from being affected by changes in ambient brightness. Based on the trajectory of moving objects, the system predicts the spatial position offset, corrects the coordinates of the synchronized environmental depth data, and further combines the illumination feature to compensate for the depth information, forming illumination-compensated depth data as the environmental depth information output, ensuring the accuracy of patients' distance perception in complex medical scenarios.

[0190] Based on synchronized environmental image data, the system uses an open vocabulary object detection model to extract visual feature vectors, match them to a semantic description library dedicated to medical scenarios, and automatically identify obstacle categories such as wheelchairs, stretchers, infusion stands, and emergency carts. It combines image texture and reflection characteristics to infer obstacle material information, such as metal, plastic, or transparent materials, and generates a composite obstacle description, which serves as obstacle category information for subsequent spatial position calculations.

[0191] The system analyzes the raw depth data from the environmental depth information and converts it into distance data per step based on the user's stride parameters, enhancing the patient's understanding of distance. Combining the obstacle's pixel coordinates and image size, the system automatically determines the coordinates of the image plane center point, calculates the obstacle's relative horizontal offset angle, and maps this to a standard clock azimuth interval. Combining the obstacle category, distance data, and azimuth information, the system outputs the obstacle's spatial position and direction.

[0192] Combining spatial location information, direction information, and environmental imagery, the system identifies scene semantic tags, distinguishes between waiting areas, corridors, wards, emergency rooms, and other medical locations, and generates scene descriptions. The system integrates obstacle distance data and orientation information to generate obstacle orientation and distance descriptions. By analyzing obstacle distribution, the system determines safe paths or generates action strategies such as temporary detours and waiting times. Ultimately, these instructions are combined to form navigation instructions, assisting patients in autonomous travel.

[0193] Based on navigation instructions, the system analyzes action strategy information in real time, adjusts the navigation path based on the patient's current location and the dynamic changes in the medical environment, and generates an updated path that meets actual traffic needs. The system verifies whether the updated path meets the safety constraints of minimum curvature radius, minimum obstacle spacing and maximum slope angle, ensures path stability, and outputs updated path information.

[0194] Based on the path update, the system parses the obstacle orientation and distance description information in the navigation instructions, establishes an inversely proportional relationship between the distance data and the vibration intensity, adjusts the feedback rhythm in combination with the patient's movement speed parameters, and generates the final tactile feedback intensity level through illumination compensation. Based on the clock orientation interval and feedback position mapping rules, the system maps obstacle information in different directions to the left, right or forehead feedback modules of the patient's device, generates a trapezoidal vibration waveform and outputs it to the target feedback position, thereby improving the patient's spatial perception of the obstacle position. If the obstacle distance is detected to be lower than the safety distance threshold, the system automatically triggers synchronized vibration, intensity increase and continuous pulse waveforms in the full feedback position, and continuously outputs warning information until the patient's position status is updated, ensuring travel safety in medical environments, reducing the risk of falls and collisions, and improving the patient's independent travel ability in complex environments such as hospitals and rehabilitation institutions.

[0195] In the field of financial technology, intelligent assisted navigation systems for high-density financial business places such as bank branches, securities business departments, and insurance service halls improve the autonomous navigation capabilities of visually impaired users within financial institutions through environmental perception, obstacle recognition, dynamic path adjustment, and multi-dimensional feedback, thereby ensuring safety and efficiency in the business process.

[0196] After a user enters a financial institution, the intelligent assisted navigation device first uses high-precision visual depth sensors to synchronously acquire environmental image and depth information. The image information reflects the real-time spatial layout of the business hall, service area, and queuing area, while the depth information is used to quantify the spatial distance and positional relationships between obstacles, people, and users. The system performs spatiotemporal synchronization on the collected image and depth data to ensure the consistency of information sources, generating synchronized environmental image and depth data, providing a stable data foundation for dynamic change detection and subsequent navigation decisions.

[0197] Financial institutions experience frequent personnel turnover. The system performs dynamic change state detection based on synchronized data, automatically identifying the movement trajectories of customers, staff, and auxiliary equipment within the lobby. It also annotates the predicted locations of dynamic obstacles in the synchronized environmental image data in real time. The system also extracts scene lighting feature information to identify lighting variations in different areas, minimizing the impact of varying lighting conditions on obstacle identification accuracy. Combined with dynamic trajectory prediction, the system corrects the spatial coordinates of the synchronized environmental depth data, generates dynamically updated depth data, and integrates the lighting feature information to form illumination-compensated depth data, which is output as environmental depth information to ensure the stability and accuracy of navigation information.

[0198] Based on environmental image information, the system uses an open vocabulary object detection model to extract visual feature vectors, match them with a semantic description library dedicated to financial business venues, identify obstacle categories such as counters, seats, advertising boards, self-service terminals, and isolation barriers, and combine image texture and reflection characteristics to determine the material properties of the obstacles. It then generates a composite obstacle description and outputs obstacle category information, providing semantic-level data support for spatial location analysis.

[0199] The system further analyzes the raw depth data from the environmental depth information and, combined with the user's historical stride parameters, converts this depth information into distance data in units of stride length, enhancing the user's understanding of actual distance. Based on the obstacle's pixel coordinates and image size parameters in the image, the system automatically calculates the coordinates of the image plane center point, analyzes the obstacle's horizontal offset angle, and converts it into a standard clock azimuth interval value. Combining the obstacle category information, stride length data, and azimuth information, the system generates the obstacle's spatial position and direction information.

[0200] Combining spatial location information, direction information, and environmental imagery, the system identifies scene semantic tags, distinguishing functional areas within financial institutions such as self-service areas, counter areas, waiting areas, and VIP areas, and generating scene descriptions. The system integrates obstacle distance data and orientation information to generate obstacle orientation and distance descriptions. Based on obstacle distribution analysis, the system dynamically generates action strategies, including suggestions to wait, detour, and pause. Finally, the system combines scene descriptions, obstacle orientation and distance descriptions, and action strategy information to generate navigation instructions, improving users' independent navigation efficiency in complex financial environments.

[0201] The system parses action strategy information based on navigation instructions, combines the user's current location information with the dynamic changes in the financial business environment, adjusts the navigation path in real time, and generates an updated navigation path. The system automatically verifies whether the updated path meets the safety constraints of minimum curvature radius, minimum obstacle spacing, and maximum slope angle, ensuring that the path is passable and stable, and ensuring that users can move safely while handling business or waiting.

[0202] After the path update is complete, the system further analyzes the obstacle's position and distance descriptions in the navigation instructions, extracting distance and position data. It then establishes an inversely proportional relationship between distance data and vibration intensity. It dynamically adjusts the feedback intensity change rate based on the user's movement speed, and performs brightness compensation for the vibration intensity based on ambient light intensity parameters to generate the final tactile feedback intensity level. Based on the position information and feedback position mapping strategy, the system determines the mapping relationship between the obstacle position and the left, right, or forehead feedback position of the user's device. It then generates a trapezoidal vibration waveform, sets the amplitude, duration, and repetition frequency parameters, and outputs it to the target feedback position in real time. If the obstacle is detected to be below the set safety distance threshold, the system automatically activates synchronized vibration at all feedback positions, increasing the vibration intensity and switching to a continuous pulse waveform mode, continuously outputting a warning signal until the user's location status is updated. This effectively prevents safety hazards such as collisions and errant entry caused by visual impairments during financial transactions, and enhances the user's independent travel and transaction experience in financial institutions such as banks, securities companies, and insurance companies.

[0203] This embodiment dynamically generates a vibration intensity that is inversely proportional to the distance by parsing the obstacle position and distance information in the navigation instructions, and adjusts the feedback intensity based on the user's speed and ambient light factors to improve the accuracy and adaptability of the tactile prompts. By selecting the feedback position based on the clock position mapping, efficient transmission of spatial direction information is achieved, making it easier for users to intuitively perceive the position of obstacles. The use of trapezoidal waveforms and multi-parameter control strategies gives the vibration feedback a good sense of hierarchy and controllability. An embedded emergency avoidance mechanism can automatically switch to a high-intensity omnidirectional pulse vibration state when the obstacle is too close, enhancing the effect of the impending danger warning and ensuring the safety of users' travel.

[0204] In one embodiment, a navigation instruction generation and path updating device is provided, which corresponds to the navigation instruction generation and path updating method in the above embodiment. Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the navigation instruction generation and path updating device of the present invention. It includes an environment perception module 10, an obstacle recognition module 20, a space modeling module 30, a navigation instruction generation module 40, and a path planning module 50. Each functional module is described in detail below:

[0205] Environmental perception module 10, used to obtain environmental image information and environmental depth information;

[0206] an obstacle recognition module 20 for performing open vocabulary object detection on the environmental image information to obtain obstacle category information;

[0207] A spatial modeling module 30 is configured to fuse the obstacle category information and the environment depth information to generate spatial position information and direction information of the obstacle;

[0208] A navigation instruction generating module 40 is configured to generate a navigation instruction including scene description information, obstacle orientation and distance description information, and action suggestion information based on the spatial position information, direction information, and the environmental image information;

[0209] The path planning module 50 is configured to update the navigation path according to the navigation instruction.

[0210] In one embodiment, the environment perception module 10 is specifically configured to:

[0211] Capturing initial environment image information and initial environment depth information through a visual depth sensor;

[0212] Performing spatiotemporal synchronization processing on the initial environment image information and the initial environment depth information to generate synchronized environment image data and synchronized environment depth data;

[0213] Detecting a dynamic change state in the environment based on the synchronized environment image data and the synchronized environment depth data;

[0214] Marking the motion trajectory of the moving object according to the dynamic change state;

[0215] extracting scene illumination features of the synchronized environment image data;

[0216] Predicting a spatial position offset based on the motion trajectory of the moving object, and correcting corresponding coordinate values ​​of the synchronized environmental depth data based on the spatial position offset to generate dynamically updated depth data;

[0217] Fusing the dynamically updated depth data with scene illumination features to generate illumination-compensated depth data, and using the illumination-compensated depth data as environmental depth information;

[0218] According to the motion trajectory of the moving object, the predicted position area of ​​the dynamic obstacle is marked in the synchronized environmental image data to generate environmental image information.

[0219] In one embodiment, the obstacle identification module 20 is specifically configured to:

[0220] Extracting a visual feature vector of the environmental image information;

[0221] matching the visual feature vector with a text embedding vector in an open vocabulary semantic description library, and determining a similarity score between the visual feature vector and the text embedding vector;

[0222] determining a category label of an obstacle of an undefined category based on the similarity score;

[0223] Analyzing the texture features and reflection features of the visual feature vector to determine the material properties of the obstacle surface;

[0224] The obstacle category label and material attributes are fused to generate a composite obstacle description, and the composite obstacle description is used as obstacle category information.

[0225] In one embodiment, the space modeling module 30 is specifically configured to:

[0226] Parsing raw depth data in the environmental depth information and converting the raw depth data into distance data based on a step unit;

[0227] Identify pixel coordinate positions of obstacles in the environment image information;

[0228] Determine the coordinates of the center point of the image plane according to the size parameters of the environmental image information;

[0229] Determining a horizontal offset angle of the pixel coordinate position relative to the coordinates of the center point of the image plane, and converting the horizontal offset angle into a clock azimuth interval value;

[0230] Based on the obstacle category information, the distance data, and the clock azimuth interval value, spatial position information and direction information of the obstacle are generated.

[0231] In one embodiment, the navigation instruction generating module 40 is specifically configured to:

[0232] Identifying scene semantic labels of the environmental image information and generating scene description information;

[0233] Fusion of the distance data in the spatial position information and the clock azimuth interval value in the direction information to generate obstacle azimuth and distance description information;

[0234] generating action strategy information based on the distribution state of the obstacle orientation and distance description information;

[0235] The scene description information, obstacle position and distance description information, and action strategy information are combined to form a navigation instruction.

[0236] In one embodiment, the path planning module 50 is specifically configured to:

[0237] parsing the action strategy information in the navigation instruction;

[0238] Get the current user location information and the dynamic change status of the environment;

[0239] Determining path update parameters based on the action strategy information, current user location information, and dynamic environmental changes;

[0240] Adjust the current navigation path trajectory according to the path update parameters to generate an updated navigation path;

[0241] Verifying whether the updated navigation path complies with a preset security constraint policy;

[0242] If the updated navigation path complies with the preset security constraint policy, the updated navigation path is output.

[0243] In one embodiment, the path planning module 50 is specifically configured to:

[0244] Parsing the obstacle position and distance description information in the navigation instruction, and extracting the distance data in the obstacle position and distance description information;

[0245] Identifying a clock azimuth description value in the obstacle azimuth and distance description information;

[0246] Establishing an inversely proportional relationship between distance data and vibration intensity, and adjusting the intensity change rate of the inversely proportional relationship according to a user movement speed parameter;

[0247] Calculating a basic vibration intensity level through the inverse proportional relationship, and performing brightness compensation on the basic vibration intensity level based on an ambient light intensity parameter to generate a final tactile feedback intensity level;

[0248] Constructing a correspondence strategy from clock orientation intervals to physical feedback positions, and mapping the left orientation interval to the left head feedback position according to the correspondence strategy, mapping the right orientation interval to the right head feedback position according to the correspondence strategy, and mapping the front orientation interval to the central forehead feedback position according to the correspondence strategy;

[0249] Determining a target feedback position based on matching the corresponding strategy with the clock orientation description value;

[0250] generating a trapezoidal vibration waveform, and setting an amplitude parameter of the trapezoidal vibration waveform to be proportional to the tactile feedback intensity level, setting a duration parameter of the trapezoidal vibration waveform to be inversely correlated with the obstacle distance value, and configuring a repetition frequency parameter of the trapezoidal vibration waveform to be positively correlated with the degree of dynamic change of the environment;

[0251] Outputting the trapezoidal vibration waveform through a driving unit corresponding to the target feedback position;

[0252] During the output process of the trapezoidal vibration waveform, when the obstacle distance value is detected to be less than the safety distance threshold, the full feedback position synchronous vibration state is activated, the vibration intensity is increased to the highest level, and a continuous pulse waveform is used to replace the trapezoidal waveform to maintain the vibration output until the user position status is updated.

[0253] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a navigation instruction generation and path update method.

[0254] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a user-side method for generating navigation instructions and updating a path.

[0255] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0256] Obtaining environmental image information and environmental depth information;

[0257] Performing open vocabulary object detection on the environmental image information to obtain obstacle category information;

[0258] Fusion of the obstacle category information and the environment depth information to generate spatial position information and direction information of the obstacle;

[0259] Based on the spatial position information, the direction information, and the environmental image information, generating navigation instructions including scene description information, obstacle orientation and distance description information, and action suggestion information;

[0260] The navigation path is updated according to the navigation instruction.

[0261] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0262] Obtaining environmental image information and environmental depth information;

[0263] Performing open vocabulary object detection on the environmental image information to obtain obstacle category information;

[0264] Fusion of the obstacle category information and the environment depth information to generate spatial position information and direction information of the obstacle;

[0265] Based on the spatial position information, the direction information, and the environmental image information, generating navigation instructions including scene description information, obstacle orientation and distance description information, and action suggestion information;

[0266] The navigation path is updated according to the navigation instruction.

[0267] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0268] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0269] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0270] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A navigation instruction generation and path updating method, characterized in that: The following steps are involved: Obtaining environmental image information and environmental depth information; Performing open vocabulary object detection on the environmental image information to obtain obstacle category information; Fusion of the obstacle category information and the environment depth information to generate spatial position information and direction information of the obstacle; Based on the spatial position information, the direction information, and the environmental image information, generating navigation instructions including scene description information, obstacle orientation and distance description information, and action suggestion information; The navigation path is updated according to the navigation instruction.

2. The navigation instruction generation and path updating method according to claim 1, wherein: Obtain environmental image information and environmental depth information, including: Capturing initial environment image information and initial environment depth information through a visual depth sensor; Performing spatiotemporal synchronization processing on the initial environment image information and the initial environment depth information to generate synchronized environment image data and synchronized environment depth data; Detecting a dynamic change state in the environment based on the synchronized environment image data and the synchronized environment depth data; Marking the motion trajectory of the moving object according to the dynamic change state; extracting scene illumination features of the synchronized environment image data; Predicting a spatial position offset based on the motion trajectory of the moving object, and correcting corresponding coordinate values ​​of the synchronized environment depth data based on the spatial position offset to generate dynamically updated depth data; Fusing the dynamically updated depth data with scene illumination features to generate illumination-compensated depth data, and using the illumination-compensated depth data as environmental depth information; According to the motion trajectory of the moving object, the predicted position area of ​​the dynamic obstacle is marked in the synchronized environmental image data to generate environmental image information.

3. The navigation instruction generation and path updating method according to claim 1, wherein: Perform open vocabulary object detection on the environmental image information to obtain obstacle category information, including: Extracting a visual feature vector of the environmental image information; matching the visual feature vector with a text embedding vector in an open vocabulary semantic description library, and determining a similarity score between the visual feature vector and the text embedding vector; determining a category label of an obstacle of an undefined category based on the similarity score; Analyzing the texture features and reflection features of the visual feature vector to determine the material properties of the obstacle surface; The obstacle category label and material attributes are fused to generate a composite obstacle description, and the composite obstacle description is used as obstacle category information.

4. The navigation instruction generation and path updating method according to claim 1, wherein: Fusion of the obstacle category information and the environment depth information to generate spatial position information and direction information of the obstacle includes: Parsing raw depth data in the environmental depth information and converting the raw depth data into distance data based on a step unit; Identify pixel coordinate positions of obstacles in the environment image information; Determine the coordinates of the center point of the image plane according to the size parameters of the environmental image information; Determining a horizontal offset angle of the pixel coordinate position relative to the coordinates of the center point of the image plane, and converting the horizontal offset angle into a clock azimuth interval value; Based on the obstacle category information, the distance data, and the clock azimuth interval value, spatial position information and direction information of the obstacle are generated.

5. The navigation instruction generation and path updating method according to claim 1, wherein: Based on the spatial position information, the direction information, and the environmental image information, generating a navigation instruction including scene description information, obstacle orientation and distance description information, and action suggestion information, including: Identifying scene semantic labels of the environmental image information and generating scene description information; Fusion of the distance data in the spatial position information and the clock azimuth interval value in the direction information to generate obstacle azimuth and distance description information; generating action strategy information based on the distribution state of the obstacle orientation and distance description information; The scene description information, obstacle position and distance description information, and action strategy information are combined to form a navigation instruction.

6. The navigation instruction generation and path updating method according to claim 1, wherein: Updating the navigation path according to the navigation instruction includes: parsing the action strategy information in the navigation instruction; Get the current user location information and the dynamic change status of the environment; Determining path update parameters based on the action strategy information, current user location information, and dynamic environmental changes; Adjust the current navigation path trajectory according to the path update parameters to generate an updated navigation path; Verifying whether the updated navigation path complies with a preset security constraint policy; If the updated navigation path complies with the preset security constraint policy, the updated navigation path is output.

7. The navigation instruction generation and path updating method according to claim 1, wherein: After updating the navigation path according to the navigation instruction, the method further includes: Parsing the obstacle position and distance description information in the navigation instruction, and extracting the distance data in the obstacle position and distance description information; Identifying a clock azimuth description value in the obstacle azimuth and distance description information; Establishing an inversely proportional relationship between distance data and vibration intensity, and adjusting the intensity change rate of the inversely proportional relationship according to a user movement speed parameter; Calculating a basic vibration intensity level through the inverse proportional relationship, and performing brightness compensation on the basic vibration intensity level based on an ambient light intensity parameter to generate a final tactile feedback intensity level; Constructing a correspondence strategy from clock orientation intervals to physical feedback positions, and mapping the left orientation interval to the left head feedback position according to the correspondence strategy, mapping the right orientation interval to the right head feedback position according to the correspondence strategy, and mapping the front orientation interval to the central forehead feedback position according to the correspondence strategy; Determining a target feedback position based on matching the corresponding strategy with the clock orientation description value; generating a trapezoidal vibration waveform, and setting an amplitude parameter of the trapezoidal vibration waveform to be proportional to the tactile feedback intensity level, setting a duration parameter of the trapezoidal vibration waveform to be inversely correlated with the obstacle distance value, and configuring a repetition frequency parameter of the trapezoidal vibration waveform to be positively correlated with the degree of dynamic change of the environment; Outputting the trapezoidal vibration waveform through a driving unit corresponding to the target feedback position; During the output process of the trapezoidal vibration waveform, when the obstacle distance value is detected to be less than the safety distance threshold, the full feedback position synchronous vibration state is activated, the vibration intensity is increased to the highest level, and a continuous pulse waveform is used to replace the trapezoidal waveform to maintain the vibration output until the user position status is updated.

8. A navigation instruction generation and path updating device, characterized in that: The navigation instruction generation and path updating device includes: Environmental perception module, used to obtain environmental image information and environmental depth information; an obstacle recognition module, configured to perform open vocabulary object detection on the environmental image information to obtain obstacle category information; A spatial modeling module, configured to fuse the obstacle category information and the environment depth information to generate spatial position information and direction information of the obstacle; A navigation instruction generation module, configured to generate a navigation instruction including scene description information, obstacle orientation and distance description information, and action suggestion information based on the spatial position information, direction information, and environmental image information; A path planning module is used to update the navigation path according to the navigation instruction.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a navigation instruction generation and path update program stored in the memory and capable of running on the processor. When the navigation instruction generation and path update program is executed by the processor, the steps of the navigation instruction generation and path update method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a navigation instruction generation and path update program, which, when executed by a processor, implements the steps of the navigation instruction generation and path update method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Blind person environment cognition auxiliary system

    CN121280967A

  • Equipment moving method and system with autonomous navigation function

    CN121384027A

  • Environmental adaptive navigation strategy adjustment system based on visual semantic segmentation

    CN121541483A

  • Old-age care robot autonomous navigation method and system based on depth vision

    CN121655541A