Positioning and navigation methods, devices, equipment, storage media, and computer program products

By generating a 3D model through feature recognition and semantic segmentation of multi-view 2D images of the target object, and combining environmental video and motion data, the problem of navigation in complex environments without a preset map is solved, and accurate real-time navigation is achieved.

CN122089986APending Publication Date: 2026-05-26CHINA MOBILE GROUP DESIGN INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE GROUP DESIGN INST
Filing Date
2026-02-27
Publication Date
2026-05-26

Smart Images

  • Figure CN122089986A_ABST
    Figure CN122089986A_ABST
Patent Text Reader

Abstract

This application discloses a positioning and navigation method, apparatus, device, storage medium, and computer program product to solve the problem that existing positioning and navigation schemes cannot provide accurate real-time navigation guidance to users in complex environments when a preset map is lacking. The method includes: performing feature recognition on at least two perspective 2D images corresponding to a target object to determine a target region in the 2D images; performing semantic segmentation on the target region to obtain a semantic segmentation mask corresponding to the target region; generating a 3D model corresponding to the target object based on the semantic segmentation mask and the 2D images; determining the pose information of the device to be navigated within the 3D model based on environmental video and motion data collected by the device to be navigated; generating a navigation path based on the position information of the target region and the pose information; and sending the navigation path to the device to be navigated for navigation guidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a positioning and navigation method, apparatus, device, storage medium, and computer program product. Background Technology

[0002] In fields such as equipment inspection and facility maintenance, the ability to quickly and accurately locate damaged equipment and provide intuitive navigation guidance to on-site personnel is a crucial technical requirement. Existing technologies primarily achieve positioning and navigation through the following two methods: The first type of solution primarily involves constructing a 3D model of the object by acquiring images from multiple angles, thereby analyzing and determining the spatial location of the damaged area. However, this technology mainly serves post-event analysis and recording; the generated 3D model is static and fails to dynamically integrate with the operator's real-time position and perspective, thus unable to provide real-time navigation guidance. The second type of solution mainly uses augmented reality technology for real-time navigation. This type of solution typically relies on a pre-built, accurate environmental map or a stable external positioning signal to function, making it unusable in complex environments without a pre-built map.

[0003] Therefore, how to achieve real-time and accurate navigation guidance from the user's real-time location to a specific target point in the absence of a pre-set map has become an urgent technical problem to be solved. Summary of the Invention

[0004] This application provides a positioning and navigation method to solve the problem that existing positioning and navigation schemes cannot provide accurate real-time navigation guidance to users in complex environments when there is a lack of preset maps.

[0005] This application also provides a positioning and navigation device to solve the problem that existing positioning and navigation schemes cannot provide accurate real-time navigation guidance to users in complex environments when there is a lack of preset maps.

[0006] This application also provides a positioning and navigation device to solve the problem that existing positioning and navigation solutions cannot provide accurate real-time navigation guidance to users in complex environments when there is a lack of preset maps.

[0007] This application also provides a computer-readable storage medium to address the problem that existing positioning and navigation schemes cannot provide accurate real-time navigation guidance to users in complex environments when a preset map is lacking.

[0008] A computer program product designed to address the problem that existing positioning and navigation solutions cannot provide accurate real-time navigation guidance to users in complex environments when a preset map is unavailable.

[0009] The embodiments of this application adopt the following technical solutions: A positioning and navigation method includes: performing feature recognition on at least two perspective two-dimensional images of a target object to determine a target region in the two-dimensional images; performing semantic segmentation on the target region to obtain a semantic segmentation mask corresponding to the target region; generating a three-dimensional model of the target object based on the semantic segmentation mask and the two-dimensional images, wherein the three-dimensional model contains position information of the target region; determining the pose information of the device to be navigated within the three-dimensional model based on environmental video and motion data collected by the device to be navigated; generating a navigation path based on the position information of the target region in the three-dimensional model and the pose information, and sending the navigation path to the device to be navigated for navigation guidance.

[0010] A positioning and navigation device includes: a semantic segmentation unit, configured to perform feature recognition on at least two perspective two-dimensional images corresponding to a target object, determine a target region in the two-dimensional images, perform semantic segmentation on the target region, and obtain a semantic segmentation mask corresponding to the target region; a three-dimensional modeling unit, configured to generate a three-dimensional model corresponding to the target object based on the semantic segmentation mask and the two-dimensional images, wherein the three-dimensional model contains position information of the target region; a pose information acquisition unit, configured to determine the pose information of the device to be navigated within the three-dimensional model based on environmental video and motion data acquired by the device to be navigated; and a navigation and positioning unit, configured to generate a navigation path based on the position information of the target region in the three-dimensional model and the pose information, and send the navigation path to the device to be navigated for navigation guidance.

[0011] A positioning and navigation device, comprising: The processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the following operations: perform feature recognition on at least two perspective images of a target object to determine a target region in the two-dimensional images; perform semantic segmentation on the target region to obtain a semantic segmentation mask corresponding to the target region; generate a three-dimensional model of the target object based on the semantic segmentation mask and the two-dimensional images, wherein the three-dimensional model contains position information of the target region; determine the pose information of the device to be navigated within the three-dimensional model based on environmental video and motion data acquired by the device to be navigated; generate a navigation path based on the position information of the target region in the three-dimensional model and the pose information, and send the navigation path to the device to be navigated for navigation guidance.

[0012] A computer-readable storage medium stores one or more programs that, when executed by an electronic device including multiple applications, cause the electronic device to perform the following operations: perform feature recognition on at least two perspectives of a target object to determine a target region in the two-dimensional images; perform semantic segmentation on the target region to obtain a semantic segmentation mask corresponding to the target region; generate a three-dimensional model of the target object based on the semantic segmentation mask and the two-dimensional images, wherein the three-dimensional model contains position information of the target region; determine the pose information of the device to be navigated within the three-dimensional model based on environmental video and motion data acquired by the device to be navigated; generate a navigation path based on the position information of the target region in the three-dimensional model and the pose information, and send the navigation path to the device to be navigated for navigation guidance.

[0013] A computer program product includes a computer program that, when executed by a processor, performs the following: feature recognition on at least two perspective two-dimensional images of a target object to determine a target region in the two-dimensional images; semantic segmentation of the target region to obtain a semantic segmentation mask corresponding to the target region; generating a three-dimensional model of the target object based on the semantic segmentation mask and the two-dimensional images, wherein the three-dimensional model contains position information of the target region; determining the pose information of the device to be navigated within the three-dimensional model based on environmental video and motion data acquired by the device to be navigated; generating a navigation path based on the position information of the target region in the three-dimensional model and the pose information; and sending the navigation path to the device to be navigated for navigation guidance.

[0014] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: Using the positioning and navigation method provided in this application, during positioning and navigation, the first step is to identify the features of at least two perspectives of the target object to determine the target region in the two-dimensional image. Then, semantic segmentation is performed on the target region to obtain a semantic segmentation mask corresponding to the target region. Subsequently, a three-dimensional model carrying semantic information corresponding to the target object is generated based on the semantic segmentation mask and the two-dimensional image. The three-dimensional model contains the position information of the target region. Next, based on the environmental video and motion data collected by the device to be navigated, the pose information of the device to be navigated in the three-dimensional model is determined. Finally, a navigation path is generated based on the position information and pose information of the target region in the three-dimensional model and sent to the device to be navigated for navigation guidance. The positioning and navigation method provided in this application has several advantages. First, by performing feature recognition on the acquired multi-view two-dimensional images, the target region in the image is determined, i.e., the damaged region to be navigated. Then, semantic segmentation is performed only on this target region to obtain an accurate semantic segmentation mask. This concentrates limited computational resources on the key target region, avoiding the computational overhead of performing semantic segmentation on the entire image, and significantly improving processing efficiency. Simultaneously, the semantic segmentation mask can accurately depict the pixel-level contour of the target region, ensuring that the geometric structure and spatial position of the target region are restored with high fidelity in the subsequent three-dimensional reconstruction process, and generating semantically labeled... The 3D model explicitly contains the location information of the target area, providing a precise spatial reference benchmark for subsequent navigation. On the other hand, the positioning and navigation method provided in this application generates a 3D model with semantic labels in real time from multi-view 2D images collected on-site. This 3D model constitutes a high-precision spatial reference system. Subsequently, by fusing motion data collected by the device to be navigated with environmental video, the real-time pose of the device to be navigated is accurately positioned in this 3D model, thereby eliminating the dependence on a preset map. This enables the positioning and navigation of specific targets even in scenarios without a preset map, greatly improving the applicability of the navigation scheme and the accuracy of positioning and navigation. Attached Figure Description

[0015] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This application provides a schematic flowchart of a positioning and navigation method. Figure 2 This is a schematic diagram of the specific structure of a positioning and navigation device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the specific structure of a positioning and navigation device provided in an embodiment of this application. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] This application provides a positioning and navigation method to address the problem that existing positioning and navigation schemes cannot provide accurate real-time navigation guidance to users in complex environments when there is a lack of preset maps.

[0018] The execution subject of the positioning and navigation method provided in this application embodiment may be, but is not limited to, at least one of a navigation server, a positioning server, and a device maintenance server; in addition, the execution subject of the method may also be the system or application (APP) itself running on these servers.

[0019] For ease of description, the following description uses a positioning and navigation system as the execution subject of this method as an example to introduce its implementation. It should be understood that using a positioning and navigation system as the execution subject is merely an illustrative example and should not be construed as a limitation of the method.

[0020] The schematic diagram of the specific implementation process of the positioning and navigation method provided in this application is shown below. Figure 1 As shown, the main steps include the following: Step 11: Perform feature recognition on the two-dimensional images of at least two viewpoints corresponding to the acquired target object, determine the target region in the two-dimensional image, perform semantic segmentation on the target region, and obtain the semantic segmentation mask corresponding to the two-dimensional image of the target region. In this embodiment, multiple two-dimensional images can be captured from at least two different perspectives around the target object using an image acquisition device, such as the camera of a smartphone, tablet, or AR glasses. Furthermore, it should be noted that to ensure the quality of the three-dimensional model obtained from the 3D reconstruction based on these two-dimensional images, this embodiment typically acquires three or more two-dimensional images with sufficient parallax angles.

[0021] After acquiring multi-view two-dimensional images, feature recognition processing can be performed on each image to determine target areas that require special attention, such as damaged areas of target objects. In this embodiment, target detection algorithms, such as YOLO and Faster R-CNN, can be used for feature recognition to identify the bounding boxes of target areas containing features such as damage, cracks, and gaps from the images.

[0022] After determining the target region, the positioning and navigation system can use a pre-trained semantic segmentation model to perform semantic segmentation processing only on the image content within that target region. In one embodiment, a Mask Region-based Convolutional Neural Network (Mask R-CNN) model can be used as the semantic segmentation model. This Mask R-CNN model can identify specific semantic categories in a two-dimensional image and output pixel-level classification results as a semantic segmentation mask.

[0023] It should be noted that the semantic segmentation mask is a binary or multi-label matrix with the same size as the original image, where the value of each pixel position identifies whether the pixel belongs to the navigation target area. In this embodiment, the navigation target area is the damaged area of ​​the target object, such as a crack area or a missing part area.

[0024] By performing the semantic segmentation process in step 11, the resulting semantic segmentation mask can mark the pixels of the damaged area (i.e. the navigation target area) of the target object as 1, and mark the pixels of the undamaged area (i.e. the non-navigation target area) of the target object as 0. This allows for accurate segmentation and labeling of the target areas that require subsequent processing and navigation from the complex background.

[0025] Step 12: Generate a 3D model of the target object based on the semantic segmentation mask and the 2D image obtained by performing Step 11. In this embodiment, the positioning and navigation system can reconstruct a three-dimensional model with semantic labels using multi-view two-dimensional images and their corresponding semantic segmentation masks, and use this three-dimensional model as a spatial reference benchmark for subsequent positioning and navigation processes. Specifically, the positioning and navigation system can construct the three-dimensional model according to the following sub-steps: Sub-step 1201: Based on the semantic segmentation mask, determine the target region (i.e., the damaged region) and the non-target region (i.e., the non-damaged region) corresponding to the target object. Sub-step 1202: Dense feature extraction is performed on the target region in the two-dimensional image to obtain the first image feature; It should be noted that, in order to achieve a balance between navigation accuracy and modeling efficiency, a hybrid feature extraction strategy can be adopted in this embodiment of the application. Different feature extraction strategies are used to extract image features from different regions of the two-dimensional image.

[0026] Specifically, for the target region indicated by the semantic segmentation mask, since the target region is the target point for subsequent navigation and localization, in order to improve navigation accuracy, a deep convolutional neural network can be used for dense feature extraction. For example, in one implementation, a ResNet network can be used to extract high-resolution dense feature maps.

[0027] Assume, I i(x,y) Let represent the value of the two-dimensional image corresponding to the i-th viewpoint at pixel coordinates (x, y). Then, through convolution, the dense feature vector of the image can be determined according to the following formula [1]: [1] Among them, I i Let W represent the image from the i-th viewpoint, where Conv is the convolution operation and W is the image from the ith viewpoint. dense These are the convolution kernel parameters used for dense feature extraction.

[0028] Sub-step 1203: Sparse feature extraction is performed on the non-target region in the two-dimensional image to obtain the second image features.

[0029] For non-target regions, in order to reduce the amount of computation, only traditional feature point detection and description algorithms, such as the Scale-Invariant Feature Transform (SIFT) algorithm, are used for sparse feature extraction.

[0030] Specifically, assuming This represents the set of sparse feature points extracted from the i-th image. It is a two-dimensional vector representing the location of a feature point in the image. For each sparse feature point... The feature vector can be obtained through the feature SIFT descriptor. .

[0031] Sub-step 1204: Generate a 3D model of the target object based on the obtained first image features and second image features.

[0032] Specifically, the 3D model can be generated following the steps below: Step 1: Determine the positional relationship between the two-dimensional images; In this embodiment of the application, sparse feature points can be matched. First, the extracted sparse feature points are used to perform feature matching between two-dimensional images from different perspectives. For the matched feature point pairs, the relative position and pose between the cameras when these images were captured are solved by geometric constraints based on the Structure from Motion (SfM) algorithm, and the positional relationship between each two-dimensional image is determined, thereby establishing the preliminary geometric relationship of all two-dimensional images in a unified three-dimensional space.

[0033] Step 2: Based on the positional relationships determined through the above steps, determine the correspondence between each first image feature; Specifically, the positioning and navigation system can match dense features belonging to the navigation target area between two-dimensional images from different perspectives based on the relative positions and attitudes between cameras obtained through execution process 1.

[0034] Step 3: Based on the correspondence obtained by executing Step 2, perform viewpoint unification processing on the first image features to obtain the third image features; It should be noted that, in order to ensure the accuracy of dense feature matching of the navigation target area, in this embodiment, the first image features corresponding to the multi-view two-dimensional images can be subjected to view unification processing. In this embodiment, the view unification processing of the first image features is constrained by the semantic segmentation mask, and its goal is to minimize a loss function that integrates the sparse and dense feature matching errors. Specifically, in this embodiment, the loss function formula is as follows [2]: [2] Among them, L sparse L represents the reprojection error based on sparse feature points. dense M represents the matching error based on dense feature points. mask Represents a semantic segmentation mask, used in calculating L dense The process considers only features within the navigation target area, with λ being the weighting coefficient. By optimizing the aforementioned loss function, the dense features of the damaged area under different viewpoints are accurately aligned in three-dimensional space, thus obtaining the second image features.

[0035] Step 4: Based on the third image features and the second image features obtained through the above steps, generate a 3D model corresponding to the target object.

[0036] Optimized, consistent third-party image features, along with sparse feature points, are input into a multi-view stereo (MVS) 3D reconstruction algorithm. The MVS algorithm performs dense matching on multiple 2D images, known camera poses, and optimized feature correspondences. By comparing pixel color, intensity, and other information of the same point in space from different viewpoints, the precise location of that point in 3D space is calculated, ultimately generating a complete and detailed 3D mesh or dense point cloud model of the target object. Since the features used to generate the 3D model are semantically information-carrying features from a semantic segmentation mask, the final generated 3D model is a semantically labeled 3D model. The 3D coordinate set of the damaged region in the model is accurately recorded, forming the local coordinate system of that region.

[0037] It should be noted that since using the MVS algorithm for 3D reconstruction is a common technique in the relevant field, the specific method of generating the 3D model will not be elaborated here. Furthermore, it should be noted that this application does not specifically limit the algorithm used for 3D reconstruction.

[0038] Step 13: Based on the environmental video and motion data collected by the device to be navigated, determine the pose information of the device to be navigated within the three-dimensional model; In this embodiment, taking a mobile terminal device (e.g., a mobile phone) held by a maintenance worker as an example, when the maintenance worker arrives at the location of the target object with the mobile terminal, they can activate the positioning and navigation function. The mobile terminal device collects environmental information, such as recording environmental video using its camera. Simultaneously, the mobile terminal device can upload motion data collected by various motion sensors to the positioning and navigation system. This allows the positioning and navigation system to determine the pose information of the mobile device according to the following sub-steps: Sub-step 1301: Based on the collected environmental video, determine the visual feature points corresponding to the device to be navigated; Sub-step 1302: Based on the acquired motion data, determine the inertial parameters corresponding to the device to be navigated, such as the pre-integral quantity of angular velocity and acceleration.

[0039] Sub-step 1303: Based on visual feature points and inertial parameters, determine the pose information of the device to be navigated in the 3D model using the Sliding-Window Nonlinear Optimization (SWNO) algorithm.

[0040] Specifically, the positioning and navigation system can employ a vision-inertial tightly coupled positioning method. The mobile device's camera continuously captures environmental video, and the positioning and navigation system extracts feature points from the video frames. Simultaneously, high-frequency motion data is provided by the mobile device's Inertial Measurement Unit (IMU). The SWNO algorithm, within a fixed-size sliding window, constructs a unified nonlinear optimization problem by combining the geometric constraints of multi-frame visual observations with the motion dynamic constraints provided by IMU pre-integration. This joint optimization solves for the precise position and attitude of the device at each moment within the window in the world coordinate system defined by the 3D model. It should be noted that since using the SWNO algorithm to determine pose information is a common technique in the relevant field, the specific methods for determining the pose information of the device to be navigated within the 3D model will not be elaborated here.

[0041] Additionally, it should be noted that, in order to address the issue of insufficient visual features caused by low-texture areas, such as smooth white pipe surfaces, in this embodiment, artificial markers, such as QR codes or April Tags, can be introduced in these low-texture areas as auxiliary feature points to enhance the robustness of positioning and navigation.

[0042] By performing the above sub-steps, the video stream collected in real time by maintenance personnel through a handheld mobile terminal can be associated with a pre-built 3D model, thereby determining which position and orientation in the 3D model corresponds to the current viewpoint of the device's camera.

[0043] Step 14: Based on the location and pose information of the target area in the 3D model obtained by performing the above steps, generate a navigation path and send it to the device to be navigated for navigation guidance.

[0044] In this embodiment of the application, the positioning and navigation system can specifically provide navigation guidance to the device to be navigated through the following sub-steps, including: Sub-step 1401: Determine the global navigation direction based on the pose information and the position information of the target area in the 3D model.

[0045] Specifically, the positioning and navigation system can determine the current position information of the device in the three-dimensional model world coordinate system based on the optimized real-time pose of the device, and then combine the position coordinates of the target area marked in the three-dimensional model to calculate the vector from the current position to the target position. The direction of this vector is the global navigation direction. In order to display it on the screen, the three-dimensional spatial vector can also be transformed into the screen coordinate system through the camera projection matrix corresponding to the current pose of the device according to the following formula [3]: [3] Where K is the camera intrinsic parameter matrix, R camera, t camera It represents the current pose of the mobile device and is used to indicate the rotation and translation matrix from the world coordinate system to the camera coordinate system. The projected result can be used to generate an arrow or compass indicator pointing to the approximate direction of the target.

[0046] Sub-step 1402: Generate an environmental map based on the environmental video, and perform obstacle detection based on the environmental map to obtain obstacle information; It should be noted that, in order to achieve accurate obstacle avoidance, the positioning and navigation system can use the visual feature points obtained by performing step 13 to generate a local environment map in real time through synchronous positioning and mapping technology, and then detect obstacles between the device and the target based on this map.

[0047] Sub-step 1403: Generate a navigation path based on obstacle information and the global navigation direction.

[0048] It should be noted that, in order to improve navigation accuracy, the positioning and navigation system in this embodiment can adopt a hierarchical strategy for navigation path planning, which can specifically include the following two levels: 1. Global Layer: Based on the aforementioned global navigation directions, coarse-grained navigation can be performed at the global level. 2. Local Layer: Specifically, guided by global direction, the positioning and navigation system can combine real-time obstacle information and employ A... The search algorithm plans a collision-free detailed path on the grid representation of the environment map, which is the final navigation path.

[0049] Sub-step 1404: Send the navigation path to the device to be navigated for navigation guidance.

[0050] In this embodiment of the application, the navigation guidance provided by the positioning and navigation system can be presented in a multimodal manner, for example, in the following two ways: Method 1, Visual Guidance: Specifically, augmented reality (AR) technology can be used to overlay navigation paths (such as guide lines on the ground), directional arrows, and real-time distances to the target onto the AR interface of the mobile device held by maintenance personnel. Furthermore, the positioning and navigation system can dynamically adjust the transparency of these AR overlay markers according to the following formula [4] based on the real-time ambient light intensity and the degree of motion blur of the device, to ensure clear visibility in any environment: [4] Among them, I light B represents the normalized value of light intensity. motion The coefficients represent the motion fuzzy evaluation coefficients, and α and β represent learnable parameters or empirical coefficients.

[0051] Method 2, voice guidance: In one implementation, the positioning and navigation system can also generate voice broadcasts simultaneously to guide the user's navigation. The generated navigation content can include: direction correction instructions, such as "Please move 0.3 meters to the left"; device perspective adjustment instructions, such as "Please raise the camera by 10 degrees"; and arrival prompts, such as "The target is directly in front of you".

[0052] It's also worth noting that, to improve the accuracy of the final arrival location, when the positioning and navigation system detects that the distance between the device to be navigated and the target area is less than a preset threshold (e.g., 1 meter), it automatically switches to a high-precision positioning mode. In this mode, the ultra-wideband (UWB) positioning module can be activated. By calculating the signal flight time between the UWB tag carried by the device to be navigated and the UWB anchor points pre-placed near the target area, a precise distance at the centimeter level is calculated. This allows for correction of the vision-inertial pose information and corresponding fine-tuning of the final navigation path, ensuring that the user is accurately guided to the damage point.

[0053] The following describes the positioning and navigation method provided in this application embodiment, taking the target object as a communication device and the navigation target area as the damaged area of ​​the communication device's casing as an example: After discovering the damage to the communication device, maintenance personnel A uses a terminal to take multiple photos around the device from different angles and uploads the photos to the positioning and navigation system. Based on the received photos, the positioning and navigation system generates a three-dimensional model of the communication device with damage markings by executing steps 11 to 12. After maintenance personnel B arrives at the scene, they open the terminal APP, point the camera at the communication device, and take real-time environmental videos. The collected environmental videos and the motion data of the maintenance personnel B's handheld terminal are uploaded to the positioning and navigation system. The positioning and navigation system calculates the pose of the maintenance personnel B's handheld terminal in real time by executing step 13 and matches it to the three-dimensional model. Subsequently, the positioning and navigation system generates navigation guidance by executing step 14, and displays arrows superimposed on the video of the real equipment on the handheld terminal of repair personnel B in real time. Repair personnel B moves left or right or adjusts the angle of the phone, while voice prompts such as "Move 20 centimeters to the right" and "Til the camera down" are given. This allows repair personnel B to be quickly and accurately guided to the tiny cracks that are difficult to see with the naked eye without having to look through drawings or find markings, which greatly improves repair efficiency and accuracy.

[0054] Using the positioning and navigation method provided in this application, during positioning and navigation, the first step is to identify the features of at least two perspectives of the target object to determine the target region in the two-dimensional image. Then, semantic segmentation is performed on the target region to obtain a semantic segmentation mask corresponding to the target region. Subsequently, a three-dimensional model carrying semantic information corresponding to the target object is generated based on the semantic segmentation mask and the two-dimensional image. The three-dimensional model contains the position information of the target region. Next, based on the environmental video and motion data collected by the device to be navigated, the pose information of the device to be navigated in the three-dimensional model is determined. Finally, a navigation path is generated based on the position information and pose information of the target region in the three-dimensional model and sent to the device to be navigated for navigation guidance. The positioning and navigation method provided in this application has several advantages. First, by performing feature recognition on the acquired multi-view two-dimensional images, the target region in the image is determined, i.e., the damaged region to be navigated. Then, semantic segmentation is performed only on this target region to obtain an accurate semantic segmentation mask. This concentrates limited computational resources on the key target region, avoiding the computational overhead of performing semantic segmentation on the entire image, and significantly improving processing efficiency. Simultaneously, the semantic segmentation mask can accurately depict the pixel-level contour of the target region, ensuring that the geometric structure and spatial position of the target region are restored with high fidelity in the subsequent three-dimensional reconstruction process, and generating semantically labeled... The 3D model explicitly contains the location information of the target area, providing a precise spatial reference benchmark for subsequent navigation. On the other hand, the positioning and navigation method provided in this application generates a 3D model with semantic labels in real time from multi-view 2D images collected on-site. This 3D model constitutes a high-precision spatial reference system. Subsequently, by fusing motion data collected by the device to be navigated with environmental video, the real-time pose of the device to be navigated is accurately positioned in this 3D model, thereby eliminating the dependence on a preset map. This enables the positioning and navigation of specific targets even in scenarios without a preset map, greatly improving the applicability of the navigation scheme and the accuracy of positioning and navigation.

[0055] In one embodiment, this application also provides a positioning and navigation device to address the problem that existing positioning and navigation solutions cannot provide accurate real-time navigation guidance to users in complex environments when a preset map is lacking. A schematic diagram of the specific structure of this positioning and navigation device is shown below. Figure 2 As shown, it includes: a semantic segmentation unit 21, a 3D modeling unit 22, a pose information acquisition unit 23, and a navigation and positioning unit 24.

[0056] The semantic segmentation unit 21 is used to perform feature recognition on the two-dimensional images of at least two perspectives corresponding to the acquired target object, determine the target region in the two-dimensional image, perform semantic segmentation on the target region, and obtain the semantic segmentation mask corresponding to the target region. The 3D modeling unit 22 is used to generate a 3D model corresponding to the target object based on the semantic segmentation mask and the 2D image, wherein the 3D model contains the location information of the target region; The pose information acquisition unit 23 is used to determine the pose information of the device to be navigated in the three-dimensional model based on the environmental video and motion data acquired by the device to be navigated. The navigation and positioning unit 24 is used to generate a navigation path based on the location information and pose information of the target area in the three-dimensional model, and send the navigation path to the device to be navigated for navigation guidance.

[0057] In one embodiment, the 3D modeling unit 22 is specifically used to: extract dense features from the target region based on the semantic segmentation mask to obtain first image features; extract sparse features from non-target regions in the 2D image to obtain second image features; and generate a 3D model corresponding to the target object based on the first image features and the second image features.

[0058] In one embodiment, the 3D modeling unit 22 is specifically used for: determining the positional relationship between each of the two-dimensional images; determining the correspondence between each of the first image features based on the positional relationship; performing viewpoint unification processing on the first image features based on the correspondence to obtain a third image feature; and generating a 3D model corresponding to the target object based on the third image feature and the second image feature.

[0059] In one embodiment, the pose information acquisition unit 23 is specifically used for: determining visual feature points corresponding to the device to be navigated based on the environmental video; determining inertial parameters corresponding to the device to be navigated based on the motion data; and determining the pose information of the device to be navigated within the three-dimensional model based on the visual feature points and the inertial parameters, using a sliding window nonlinear optimization algorithm.

[0060] In one embodiment, the navigation and positioning unit 24 is specifically configured to: determine a global navigation direction based on the pose information and the location information of the target area; generate an environmental map based on the environmental video, and perform obstacle detection based on the environmental map to obtain obstacle information; and generate a navigation path based on the obstacle information and the global navigation direction.

[0061] In one embodiment, the navigation and positioning unit 24 is specifically used to: determine the current position information of the device to be navigated in the three-dimensional model based on the pose information; and determine the global navigation direction based on the current position information and the position information of the target area.

[0062] Using the positioning and navigation device provided in this application embodiment, when performing positioning and navigation, the device first identifies the features of at least two perspectives of the target object to determine the target region in the two-dimensional image. Then, it performs semantic segmentation on the target region to obtain a semantic segmentation mask corresponding to the target region. Subsequently, based on the semantic segmentation mask and the two-dimensional image, it generates a three-dimensional model corresponding to the target object that carries semantic information. The three-dimensional model contains the position information of the target region. Next, based on the environmental video and motion data collected by the device to be navigated, it determines the pose information of the device to be navigated in the three-dimensional model. Finally, based on the position information and pose information of the target region in the three-dimensional model, it generates a navigation path and sends it to the device to be navigated for navigation guidance. The positioning and navigation method provided in this application has several advantages. First, by performing feature recognition on the acquired multi-view two-dimensional images, the target region in the image is determined, i.e., the damaged region to be navigated. Then, semantic segmentation is performed only on this target region to obtain an accurate semantic segmentation mask. This concentrates limited computational resources on the key target region, avoiding the computational overhead of performing semantic segmentation on the entire image, and significantly improving processing efficiency. Simultaneously, the semantic segmentation mask can accurately depict the pixel-level contour of the target region, ensuring that the geometric structure and spatial position of the target region are restored with high fidelity in the subsequent three-dimensional reconstruction process, and generating semantically labeled... The 3D model explicitly contains the location information of the target area, providing a precise spatial reference benchmark for subsequent navigation. On the other hand, the positioning and navigation method provided in this application generates a 3D model with semantic labels in real time from multi-view 2D images collected on-site. This 3D model constitutes a high-precision spatial reference system. Subsequently, by fusing motion data collected by the device to be navigated with environmental video, the real-time pose of the device to be navigated is accurately positioned in this 3D model, thereby eliminating the dependence on a preset map. This enables the positioning and navigation of specific targets even in scenarios without a preset map, greatly improving the applicability of the navigation scheme and the accuracy of positioning and navigation.

[0063] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Please refer to it. Figure 3At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.

[0064] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0065] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.

[0066] The processor reads the corresponding computer program from non-volatile memory into main memory and then executes it, forming a positioning and navigation device at the logical level. The processor executes the program stored in memory and specifically performs the following operations: Feature recognition is performed on at least two perspective 2D images corresponding to the target object to determine the target region in the 2D images. Semantic segmentation is performed on the target region to obtain a semantic segmentation mask corresponding to the target region. Based on the semantic segmentation mask and the 2D images, a 3D model corresponding to the target object is generated, wherein the 3D model contains the position information of the target region. Based on the environmental video and motion data collected by the device to be navigated, the pose information of the device to be navigated within the 3D model is determined. Based on the position information of the target region in the 3D model and the pose information, a navigation path is generated and the navigation path is sent to the device to be navigated for navigation guidance.

[0067] The above is as stated in this application. Figure 3The positioning and navigation electronic device method disclosed in the illustrated embodiments can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0068] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0069] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by a portable electronic device including multiple applications, enable the portable electronic device to perform... Figure 1 The positioning and navigation method shown in the embodiment is specifically used to perform the following operations: Feature recognition is performed on at least two perspective 2D images corresponding to the target object to determine the target region in the 2D images. Semantic segmentation is performed on the target region to obtain a semantic segmentation mask corresponding to the target region. Based on the semantic segmentation mask and the 2D images, a 3D model corresponding to the target object is generated, wherein the 3D model contains the position information of the target region. Based on the environmental video and motion data collected by the device to be navigated, the pose information of the device to be navigated within the 3D model is determined. Based on the position information of the target region in the 3D model and the pose information, a navigation path is generated and the navigation path is sent to the device to be navigated for navigation guidance.

[0070] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0071] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0072] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0073] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0074] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0075] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0076] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0077] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0078] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0079] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A positioning and navigation method, characterized in that, include: Feature recognition is performed on at least two perspective 2D images of the target object to determine the target region in the 2D images, and semantic segmentation is performed on the target region to obtain the semantic segmentation mask corresponding to the target region. Based on the semantic segmentation mask and the two-dimensional image, a three-dimensional model corresponding to the target object is generated, wherein the three-dimensional model contains the location information of the target region; Based on the environmental video and motion data collected by the device to be navigated, the pose information of the device to be navigated within the three-dimensional model is determined; Based on the location information and pose information of the target area in the 3D model, a navigation path is generated and sent to the device to be navigated for navigation guidance.

2. The method according to claim 1, characterized in that, The step of generating a 3D model corresponding to the target object based on the semantic segmentation mask and the 2D image specifically includes: Based on the semantic segmentation mask, dense feature extraction is performed on the target region to obtain the first image features; Sparse feature extraction is performed on the non-target regions in the two-dimensional image to obtain the second image features; A three-dimensional model corresponding to the target object is generated based on the first image features and the second image features.

3. The method according to claim 2, characterized in that, The step of generating a 3D model of the target object based on the first image features and the second image features specifically includes: Determine the positional relationship between each of the two-dimensional images; Based on the positional relationship, determine the correspondence between each of the first image features; Based on the correspondence, the first image features are subjected to viewpoint unification processing to obtain the third image features; A three-dimensional model corresponding to the target object is generated based on the third image features and the second image features.

4. The method according to claim 1, characterized in that, The step of determining the pose information of the device under navigation within the 3D model based on the acquired environmental video and motion data collected by the device under navigation specifically includes: Based on the environmental video, determine the visual feature points corresponding to the device to be navigated; Based on the motion data, determine the inertial parameters corresponding to the device to be navigated; Based on the visual feature points and the inertial parameters, the pose information of the device to be navigated within the three-dimensional model is determined using a sliding window nonlinear optimization algorithm.

5. The method according to claim 1, characterized in that, The step of generating a navigation path based on the location information and pose information of the target region in the 3D model specifically includes: Based on the pose information and the location information of the target area, the global navigation direction is determined; An environmental map is generated based on the environmental video, and obstacle detection is performed based on the environmental map to obtain obstacle information; A navigation path is generated based on the obstacle information and the global navigation direction.

6. The method according to claim 5, characterized in that, The step of determining the global navigation direction based on the pose information and the position information of the target area specifically includes: Based on the pose information, the current position information of the device to be navigated in the three-dimensional model is determined; The global navigation direction is determined based on the current location information and the location information of the target area.

7. A positioning and navigation device, characterized in that, include: The semantic segmentation unit is used to perform feature recognition on at least two perspectives of the acquired target object in two-dimensional images, determine the target region in the two-dimensional images, perform semantic segmentation on the target region, and obtain the semantic segmentation mask corresponding to the target region. A 3D modeling unit is used to generate a 3D model corresponding to the target object based on the semantic segmentation mask and the 2D image, wherein the 3D model contains the location information of the target region; The pose information acquisition unit is used to determine the pose information of the device to be navigated in the three-dimensional model based on the environmental video and motion data acquired by the device to be navigated. The navigation and positioning unit is used to generate a navigation path based on the location information and pose information of the target area in the three-dimensional model, and to send the navigation path to the device to be navigated for navigation guidance.

8. A positioning and navigation device, comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the following operations: Feature recognition is performed on at least two perspective 2D images of the target object to determine the target region in the 2D images, and semantic segmentation is performed on the target region to obtain the semantic segmentation mask corresponding to the target region. Based on the semantic segmentation mask and the two-dimensional image, a three-dimensional model corresponding to the target object is generated, wherein the three-dimensional model contains the location information of the target region; Based on the environmental video and motion data collected by the device to be navigated, the pose information of the device to be navigated within the three-dimensional model is determined; Based on the location information and pose information of the target area in the 3D model, a navigation path is generated and sent to the device to be navigated for navigation guidance.

9. A computer-readable storage medium storing one or more programs that, when executed by an electronic device including a plurality of applications, cause the electronic device to perform the positioning and navigation method as described in any one of claims 1-6.

10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the positioning and navigation method as described in any one of claims 1-6.