Semantic aerial view visual relocation method and device in non-exposed scene, electronic equipment, storage medium and program product
By using a semantic-geometric fusion mechanism and an edge-cloud collaborative architecture, a dense 3D semantic point cloud is generated and semantic bird's-eye view relocalization is performed. This solves the problems of positioning accuracy and stability in non-exposed scenarios, achieving high-precision and robust visual relocalization, which is suitable for intelligent transportation and underground inspection.
Patent Information
- Application Number
- CN202511595280.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-02-10
AI Technical Summary
In non-exposed scenarios where GNSS signals are limited or obstructed, traditional visual relocation methods suffer from insufficient positioning accuracy, poor semantic consistency, low utilization of geometric constraints, and limited real-time performance, making it difficult to achieve high-precision and stable positioning in complex environments.
By constructing a semantic-geometric fusion mechanism, a dense 3D semantic point cloud is generated using a pre-trained deep semantic detection model and a voxel differentiable rendering model. Combined with the short-term trajectory information of the inertial measurement unit, joint mapping and relocalization of semantic and geometric information are achieved. A semantic mask-guided reciprocal matching strategy is adopted to improve matching stability, and real-time deployment is achieved through an edge-cloud collaborative architecture.
It achieves centimeter-level positioning accuracy in non-exposed environments, suppresses matching drift under conditions of varying illumination and complex structures, possesses sub-second real-time performance and high robustness, outputs interpretable semantic bird's-eye view maps, and supports applications such as intelligent transportation and underground inspection.
Smart Images

Figure CN121505029A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer vision, artificial intelligence and intelligent navigation, in particular to a non-exposed scene semantic bird's eye view visual relocalization method and device based on semantic-geometric fusion and bird's eye mapping mechanism, electronic equipment, storage medium and computer program product, which is suitable for various complex environments such as rail transit, underground space inspection, unmanned system positioning and emergency navigation. BACKGROUND
[0002] This part aims to provide background or context for the embodiments of the present disclosure stated in the claims. The following content is only used to help understand the technical background of the present disclosure, and does not mean that the content belongs to the prior art known.
[0003] In non-exposed scenes such as rail transit, underground pipe gallery, tunnel, underground parking lot, the environment is closed, the light is complex, the structure is repeated, and the external signal is easily disturbed, which brings challenges to the spatial positioning and navigation of the device. The traditional positioning scheme based on Global Navigation Satellite System (GNSS) or GNSS+Inertial Measurement Unit (IMU) combined navigation is difficult to work stably in non-exposed scenes due to signal obstruction or reflection; even through inertia compensation, it is easy to appear cumulative drift and long-term error out of control problems.
[0004] In order to get rid of the dependence on GNSS signal, researchers have proposed various positioning technologies based on visual perception, such as feature point matching, sparse three-dimensional mapping (SLAM) and deep learning end-to-end positioning method. This kind of method estimates the pose by extracting geometric features or learning global representation from images, but still has the following shortcomings in non-exposed scenes:
[0005] (1) Weak texture and repeated structure problem: the surface of the tunnel or underground space is mostly smooth concrete structure, lacking of significant feature points, resulting in poor stability of feature extraction and matching;
[0006] (2) Light changes and noise interference: uneven distribution of light sources and strong reflection in non-exposed environments make the image brightness and contrast change significantly, affecting the robustness of visual algorithms;
[0007] (3) Semantic and geometric information is split: existing visual relocalization mostly focuses on geometric feature matching, and it is difficult to use semantic structure information (such as lighting devices, cable bridges, ventilation ports, etc.) in the scene, and it is impossible to enhance the spatial constraints at the semantic level;
[0008] (4) Limited accuracy of three-dimensional mapping: traditional sparse point cloud mapping method has insufficient precision and high noise, which is difficult to generate high-quality semantic map that can directly support relocalization;
[0009] (5) Real-time and deployment complexity: existing high-precision methods often rely on high-performance computing platforms, making it difficult to achieve real-time processing on vehicles or mobile terminals.
[0010] In recent years, with the development of semantic perception and dense mapping technology, researchers have begun to introduce semantic information into visual relocalization, improving the robustness of feature recognition and matching through semantic segmentation or object detection. However, current solutions generally have the following problems: (1) semantic information is mostly limited to two-dimensional images, making it difficult to fully integrate with three-dimensional structures;
[0011] (2) Point cloud and semantic fusion lack unified geometric constraints, with poor semantic consistency and spatial alignment accuracy; (3) Lack of semantic structure representation and relocalization model suitable for non-exposed environments, making it difficult to achieve semantic and geometric collaborative optimization.
[0012] Therefore, how to construct a high-robustness visual relocalization method that integrates semantic understanding and spatial geometry in GNSS-void, weak-texture, and complex-structure non-exposed scenes has become a key technical problem in the field of intelligent navigation and perception.
[0013] The present disclosure proposes an innovative solution to the above problems - by constructing a semantic-enhanced three-dimensional point cloud and two-dimensional bird's eye view mapping model, a semantic-geometric dual-channel fusion visual relocalization framework is realized, which effectively improves the positioning accuracy and stability in non-exposed environments while maintaining real-time and deployment feasibility. SUMMARY
[0014] Therefore, the present disclosure proposes a non-exposed scene semantic bird's eye view visual relocalization method and device based on semantic-geometric fusion mechanism, electronic equipment, storage medium and computer program product, to overcome the problems of insufficient positioning accuracy, poor semantic consistency, low utilization rate of geometric constraints, and limited real-time performance in tunnels, underground spaces and other non-exposed environments.
[0015] The present disclosure constructs a systematic solution from three aspects of semantic understanding, geometric modeling and end-cloud collaboration, forming a semantic-enhanced high-precision visual relocalization framework for non-exposed environments to solve the above technical problems. Based on the full integration of visual semantics and spatial structure information, the framework realizes the complete technical closed loop from image semantic extraction to three-dimensional reconstruction, from structure recognition to relocalization inference.
[0016] I. Overall Idea
[0017] 1. Semantic-geometric fusion mechanism.
[0018] The disclosure first utilizes a pre-trained deep semantic detection model to perform semantic target recognition and region extraction on an input image, and realizes spatial fusion combined with geometric constraints between images to construct a dense three-dimensional semantic point cloud with both semantic labels and structural features. This mechanism realizes the reverse constraint of semantic information on the geometric reconstruction process, thereby improving the stability and consistency of mapping and repositioning in non-exposed scenes.
[0019] 2. High-fidelity semantic mapping and structural expression.
[0020] To further improve mapping accuracy, the disclosure introduces a Volumetric Gaussian Grid Transformer (VGGT) model, which obtains a high-density, differentiable three-dimensional semantic point cloud expression by jointly optimizing color, depth, and edge consistency loss. Subsequently, the system analyzes the normal vector and spatial distribution characteristics of the point cloud, automatically identifies major structural planes such as the ground and left / right walls, and projects them onto a two-dimensional coordinate system to generate a Bird-Eye-View Map (BEV Map) with semantic annotations, realizing unified spatial expression of the semantic and geometric layers.
[0021] 3. Robust repositioning mechanism guided by semantic prior.
[0022] In the repositioning phase, the disclosure proposes a semantic mask guided mutual nearest neighbor matching strategy, which only performs feature point matching in regions of the same semantic class to reduce false matches caused by weak texture or repetitive structures. The system further fuses short-term trajectory information from an inertial measurement unit to realize joint pose estimation of visual and inertial information through Kalman filtering, thereby maintaining high robustness and high accuracy in complex conditions such as large changes in lighting and obvious symmetrical structures.
[0023] 4. End-to-cloud collaboration and real-time deployment architecture.
[0024] To balance computational efficiency and positioning accuracy, the disclosure designs a collaborative architecture that combines lightweight feature extraction on the edge and semantic mapping / reasoning in the cloud. This architecture realizes hierarchical data processing and load balancing, maintaining centimeter-level positioning accuracy while the overall system response has sub-second real-time performance, meeting the online navigation and monitoring needs of typical non-exposed environments such as rail transit trains, underground inspection vehicles, and unmanned systems.
[0025] II. Method Scheme
[0026] Based on the above innovative design, the semantic bird's eye view visual repositioning method in non-exposed scenes of the disclosure includes the following steps:
[0027] S210: Image acquisition and semantic feature extraction.
[0028] The reference image sequence in the non-exposed scene (such as a tunnel, an underground pipe gallery, etc.) is acquired, and the images are subjected to size standardization, distortion correction, brightness equalization and noise suppression processing to ensure the quality of the input data. Then, based on a pre-trained semantic detection model (such as YOLO, etc.), the key semantic target area in the image is identified, and the feature area with a semantic label is extracted to provide a basis for subsequent semantic fusion;
[0029] S220: new image feature matching.
[0030] During operation, based on the semantic instance area of the current acquisition image and the above-mentioned reference image sequence, sparse key points inside the instance are further extracted, and reciprocal matching under instance guidance is performed to construct local geometric correspondence;
[0031] S230: three-dimensional semantic mapping.
[0032] Combined with the geometric relationship between image frames and the depth estimation result, a dense three-dimensional semantic point cloud is generated; and through voxelization and a differentiable rendering mechanism, high-precision fusion of semantic and geometric information is realized, and camera pose is estimated and optimized;
[0033] S240: structure recognition and bird's eye view generation.
[0034] The normal vector distribution of the semantic point cloud is clustered and fitted to identify main structure planes such as the ground and the wall, and project them onto a two-dimensional coordinate system to generate a semantic bird's eye view, realizing structured space expression; S250: repositioning and pose estimation.
[0035] Semantic structure features are extracted from real-time acquired images and matched with a pre-generated semantic bird's eye view to realize high-robustness visual repositioning with semantic enhancement.
[0036] III. System device implementation
[0037] Based on the same inventive concept as the above method, the disclosure also provides a semantic bird's eye view visual repositioning device in a non-exposed scene, which includes several functional modules, as follows:
[0038] 1. Image acquisition module, for acquiring multi-view image data in the direction of the path in a non-exposed scene (such as a tunnel, an underground pipe gallery, etc.) in real time, and performing distortion correction, brightness normalization and denoising processing on the acquired images to ensure the quality and consistency of the subsequent processing data;
[0039] 2. Semantic recognition and feature extraction module, for calling a pre-trained deep semantic detection model to perform target detection and semantic segmentation on the input image, and extract target areas or structure boundaries with semantic categories to provide basic feature information for semantic-enhanced space modeling;
[0040] 3. The 3D mapping and semantic projection module is used to recover the camera pose based on the geometric relationships between image frames and the depth inference results, and generate a dense 3D point cloud model; and combines a differentiable rendering mechanism to jointly optimize color, depth and edge consistency to form a high-fidelity, structured semantic point cloud.
[0041] 4. Semantic mapping and fusion module, which is used to map semantic labels to the three-dimensional spatial point cloud based on the projection relationship between the image and the point cloud, realize the joint expression of semantic information and spatial geometric information, and generate a dense three-dimensional semantic point cloud model with semantic attributes;
[0042] 5. The structure recognition and bird's-eye view generation module is used to analyze the normal vectors and spatial distribution features in the semantic point cloud, identify the main structural planes such as the ground and left / right walls, and project the key planes into a two-dimensional coordinate system to generate a bird's-eye view with semantic annotations, thereby achieving interpretable fusion of semantics and spatial structure.
[0043] 6. The relocalization and pose estimation module is used to estimate the spatial pose of the current device based on the semantic structural features of the newly acquired image and the pre-constructed semantic bird's-eye view. At the same time, it integrates IMU data or other auxiliary sensor information and improves the stability and robustness of localization through filtering and optimization.
[0044] Through the coordinated operation of the above modules, the device can complete the entire process from visual semantic understanding, 3D mapping to pose estimation in non-exposed scenarios where GNSS signals are limited or completely ineffective, achieving high-precision and robust visual repositioning.
[0045] In addition, this disclosure also provides the following technical implementations based on the same concept:
[0046] 1. An electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the program, it implements the non-exposed scene semantic bird's-eye view visual relocalization process as described in the foregoing method.
[0047] 2. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a computer, cause the computer to implement all the steps of the above-described non-exposed scene semantic bird's-eye view visual relocation method.
[0048] 3. A computer program product comprising computer-executable instructions that, when executed on a computer or processing unit, cause to perform the non-exposed scene semantic bird's-eye view visual relocation method.
[0049] Through the above-mentioned device and system design, this disclosure achieves deep integration in key aspects such as image semantic extraction, 3D point cloud mapping, structure recognition and plane extraction, semantic bird's-eye view construction and visual relocation matching, breaking through the bottleneck problems of low accuracy, sparse features and insufficient robustness in traditional non-exposed scene relocation technology.
[0050] This device can effectively improve the accuracy and stability of visual positioning in enclosed structured environments such as tunnels and underground spaces, providing reliable spatial perception and positioning support for intelligent transportation, underground inspection, unmanned systems and digital twin applications.
[0051] IV. Beneficial Effects
[0052] This disclosure achieves high-precision visual relocalization in complex, non-exposed environments through semantic-geometric joint modeling, voxel differentiable rendering mechanism, and edge-cloud collaborative optimization architecture, breaking through the performance bottleneck of traditional GNSS-based or pure geometric vision methods under conditions of weak texture, lighting variation, and structural repetition.
[0053] Compared with the prior art, this disclosure has the following advantages and beneficial effects:
[0054] 1. Positioning accuracy is effectively improved.
[0055] By jointly modeling semantic and geometric information, dense semantic point clouds and semantic bird's-eye view maps are constructed to achieve centimeter-level positioning accuracy; high-precision pose estimation can still be achieved in GNSS failure scenarios.
[0056] 2. Enhanced robustness and stability.
[0057] A semantic mask-guided reciprocal matching mechanism is adopted to effectively suppress matching drift caused by factors such as weak texture, repetitive structure and lighting changes, and achieve continuous and stable relocation results.
[0058] 3. Improved mapping accuracy and semantic consistency.
[0059] By leveraging the VGGT voxel differentiable rendering method to jointly optimize color, depth, and edge consistency, the accuracy and semantic consistency of 3D mapping are significantly improved compared to traditional methods, ensuring strict alignment between semantics and geometric space.
[0060] 4. Balancing real-time performance with energy efficiency.
[0061] Through an edge-cloud collaborative architecture, data is processed in layers and task load is balanced. While ensuring centimeter-level positioning accuracy, the system has sub-second real-time response capability, effectively reducing energy consumption and improving operating efficiency.
[0062] 5. Semantic interpretability and engineering scalability.
[0063] The output is presented as a two-dimensional semantic bird's-eye view, offering advantages in interpretability and visualization. It directly supports digital twins, train scheduling, and intelligent operation and maintenance systems. The framework is adaptable to various visual sensors and deep semantic models, facilitating expansion to other non-exposed environments (such as underground utility tunnels and underground parking lots). In summary, this disclosure, through multi-layered innovations in semantic enhancement, geometric optimization, and collaborative computing, achieves high-precision, robust, and interpretable visual relocalization in complex non-exposed scenarios, providing a reliable spatial positioning technology foundation for intelligent transportation, underground inspection, emergency navigation, and embodied intelligent systems.
[0064] In the detailed implementation section, the numbering corresponding to the steps can be relabeled to adapt to the description of the embodiments. To better understand the technical solutions of this disclosure, further explanation is provided below in conjunction with the accompanying drawings and embodiments. Attached Figure Description
[0065] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0066] Figure 1 This is a schematic diagram illustrating an application scenario of the semantic bird's-eye view visual relocation method in a non-exposed scenario provided in this embodiment of the disclosure.
[0067] Figure 2 This is a flowchart illustrating a semantic bird's-eye view visual relocation method in a non-exposed scenario provided by an embodiment of the present disclosure, which corresponds to steps S210 to S250 in the invention content.
[0068] Figure 3 A schematic diagram of a structure for visual relocation of semantic bird's-eye view in a non-exposed scenario provided by an embodiment of this disclosure;
[0069] Figure 4 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0070] The specific embodiments of this disclosure will be further described below with reference to the accompanying drawings. Those skilled in the art will understand that this part is a specific example of the foregoing solution, rather than a limitation of the claims.
[0071] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0072] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this application's technical solution, based on the prompt message.
[0073] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0074] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this application. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this application.
[0075] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0076] To make the objectives, technical solutions, and advantages of this disclosure clearer, the principles and spirit of this disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.
[0077] In this article, it is important to understand that any number of elements in the accompanying figures is for illustrative purposes and not for limitation, and any naming is for distinction only and has no limiting meaning.
[0078] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar words used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly. The article "a" or "an" preceding an element does not exclude the existence of multiple such elements.
[0079] The principles and spirit of this disclosure are explained in detail below with reference to several representative embodiments. Existing visual relocalization techniques for non-exposed scenes generally rely on geometric feature mapping and matching. In non-exposed environments such as tunnels and indoor spaces with complex structures, large lighting variations, or occlusions, problems such as insufficient extraction of key features, limited shared field of view, and low matching accuracy are prone to occur, affecting the overall relocalization effect. Especially under conditions of uneven lighting, large changes in target scale, and strong image noise, traditional methods struggle to accurately perceive the position and category of key objects in the scene, resulting in large fluctuations in localization results and poor robustness.
[0080] This disclosure utilizes existing deep learning object detection models and image-to-3D point cloud modeling frameworks to effectively extract key semantic targets from images of non-exposed scenes (such as tunnels and underground utility tunnels). Through geometric reconstruction between multi-view images, a dense 3D point cloud with spatial consistency is generated. Then, by leveraging the projection correspondence between images and the point cloud, 2D semantic information is mapped to 3D space, achieving a fusion of semantic and geometric structural information. The resulting semantically enhanced point cloud not only improves the map's semantic expressiveness but also provides effective spatial prior constraints for subsequent structural recognition and visual relocalization.
[0081] To overcome the limitations of existing technologies, this disclosure proposes a semantically enhanced visual relocalization method for non-exposed scenes. It extracts semantic target regions using a pre-trained object detection model, generates semantic point clouds using image depth reconstruction methods, and further extracts main planar regions such as the ground, left wall, and right wall in non-exposed scenes like tunnels, converting them into semantically annotated bird's-eye view images. During the relocalization stage, the input image can be feature-matched with the known bird's-eye view image through semantic structure, and pose estimation is performed using spatial consistency, effectively improving localization accuracy and stability. This method is robust and can adapt to non-exposed environments with varying structural and textural fineness, showing broad application potential in rail transit scenarios.
[0082] After introducing the basic principles of this disclosure, various non-limiting embodiments of this disclosure will be described in detail below.
[0083] refer to Figure 1 This is an application scenario of the semantic mapping and relocalization method for non-exposed scenes provided by the exemplary embodiments of this disclosure. The application scenario includes a train-side image acquisition device 101, a server 102, a data storage system 103, and a client 104. The train-side image acquisition device 101, server 102, data storage system 103, and client 104 can all be connected via wired or wireless communication networks to realize the transmission of image data in non-exposed scenes, the invocation of model inference services, and the display of relocalization results.
[0084] The train-side image acquisition device 101 can be installed on trains in rail transit systems, such as intercity trains, subway trains, or high-speed trains. It may include sensing modules such as a forward-facing camera, a binocular stereo camera, a fisheye camera, or a depth camera. The device is configured to periodically acquire image sequences and auxiliary sensor data of non-exposed scenes during train operation, and then package and upload the acquired data, along with a timestamp, to the server 102.
[0085] Both server 102 and data storage system 103 can be independent physical servers deployed on the ground, or they can be server clusters or distributed systems composed of multiple physical servers. They can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0086] In some exemplary embodiments, the semantic bird's-eye view visual relocation system in non-exposed scenarios can run on server 102.
[0087] In some exemplary embodiments, the mapping phase of the semantic bird's-eye view visual relocalization system in non-exposed scenarios can be centrally executed on server 102, while preprocessing modules such as image acquisition, compression, and feature extraction can be distributed across image acquisition device 101 to improve overall operating efficiency and reduce transmission bandwidth overhead. In actual system deployment, connections between different modules can be achieved based on technologies such as local area networks, 5G communication, industrial Wi-Fi, and railway private networks.
[0088] The data storage system 103 includes multiple logical modules for semantic model storage, image history caching, point cloud map management, and relocation model parameter updates.
[0089] Server 102 can receive data from train-side image acquisition device 101 and provide bird's-eye view visual repositioning service to user on client 104.
[0090] Server 102 uses a pre-trained target detection model to extract semantic targets from the image sequences acquired by image acquisition device 101;
[0091] Server 102 performs 3D mapping based on the spatiotemporal registration information between the image sequences, and forms a structured semantic point cloud model through the image-point cloud mapping relationship, which is then stored in data storage system 103;
[0092] Server 102 extracts the main structural plane of the non-exposed scene from the semantic point cloud and converts it into a semantic map in the form of a bird's-eye view, which is also stored in the data storage system 103;
[0093] During train operation, server 102 matches the semantic structure of the current image frame received by image acquisition device 101 with the established semantic map, and estimates the pose and relocalization result of the current train camera based on spatial consistency.
[0094] Server 102 sends the relocation result to client 104;
[0095] Users can obtain the results of bird's-eye view visual relocation through client 104 in different non-exposed scenarios (taking tunnels as an example here).
[0096] Client 104 may include a train operation and positioning system, a train dispatching and command center terminal, an operation management console, a maintenance work tablet, or a portable smart terminal. Client 104 is configured to maintain a real-time or periodic connection with server 102, used to receive and display current train pose estimation results, switch between different tunnel section maps, replay positioning trajectories, and receive alarms and model running status information. This client can also be used to upload tag information, annotate mapping areas, and retrieve relocation history records. Data storage system 103 stores a large amount of training data for pre-training the target detection model within the tunnel. This training data includes various visual sensor data and corresponding real-world coordinates, enabling the model to learn how to map environmental features to accurate location coordinate information. The sources of training data include, but are not limited to, existing databases, data crawled from the internet, or data uploaded by users when using the client. When the difference between the location information output by the target detection model and the real coordinates reaches a certain requirement, server 102 can provide users with more accurate visual relocation services for bird's-eye view in non-exposed scenarios based on this model.
[0097] It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this disclosure, and the implementation of this disclosure is not limited in any way. On the contrary, the implementation of this disclosure can be applied to any applicable scenario.
[0098] refer to Figure 2 Semantic bird's-eye view visual relocalization in non-exposed scenarios, applied to servers, the method includes the following steps:
[0099] Step S210: Extract semantic targets and match instances in images from non-exposed scenes.
[0100] In this embodiment, the server receives images of the current non-exposed scene uploaded by the train-end image acquisition device and selects reference images with a certain spatial adjacency relationship from a pre-built image database. The goal of this embodiment is to accurately extract key instance objects that coexist in the current image and the reference image, as the defined region for subsequent structural feature matching. The specific steps are as follows:
[0101] Extracting semantic bounding boxes from the target image: Using a YOLO series deep object detection network to process the image, extracting target regions with semantic categories from the image, resulting in several semantically bounding boxes:
[0102]
[0103] Among them, (x i ,y i ) represents the coordinates of the top-left corner of the i-th bounding box, w i ,h iThese are the width and height, respectively, c i For the target category, σ i The detection confidence level for this target is set. The server sets a threshold θ. d Only retain σ i ≥θ d The target bounding box is used as a candidate semantic region.
[0104] Target image and current image feature matching: The server matches features based on the target category c. i Based on spatial distribution similarity, instance-level pairing is performed between target objects in the current image and reference images to form an instance correspondence table between the current image and reference images. This table is used for subsequent region cropping and saliency enhancement in feature point extraction. Employing instance-level semantic masks not only reduces background interference but also effectively improves the reliability of feature point extraction and the accuracy of region matching, providing accurate priors for subsequent feature extraction and matching.
[0105] Step S220: Extract and match feature points in the current image and the reference image.
[0106] After obtaining the target semantic information through step S210, the server will further extract sparse key points inside the instance based on the semantic instance regions of the current image and the reference image obtained by the information, and perform instance-guided reciprocal matching to construct a local geometric correspondence.
[0107] The server can use differentiable feature extractors such as SuperPoint or ORB to extract differentiable features from each bounding box region (x, y) in the image. i ,y i ,w i ,h i Perform feature extraction to obtain a set of key points:
[0108]
[0109] Among them, (u ij ,v ij Let d be the coordinates of the j-th key point within the i-th target region. ij This is the descriptor vector for that point. For the corresponding semantic regions (c) in the current image and the reference image... i The server employs a reciprocal matching strategy to construct bidirectional nearest neighbor pairs among similar targets:
[0110] M={((u ij ,v ij ),(u′ ij ,v′ ij ))∣match(d ij ,d′ ij )<τ m}
[0111] Where, τ m This is a descriptor distance threshold used to filter out false matches.
[0112] Under semantic guidance, feature extraction and matching performed only within instances can effectively avoid redundant matching and mismatching between different semantic regions, thereby improving the local credibility and matching density of geometric constraints.
[0113] In addition, to improve matching quality, the server can remove abnormal matching points through the RANSAC algorithm and further enhance the consistency between point pairs by using semantic mask constraints and pixel adjacency relationships, and finally construct a robust set of pseudo-corresponding points M, laying the foundation for subsequent pose estimation.
[0114] Step S230: The camera pose is estimated and optimized using an image-to-3D point cloud conversion algorithm. After obtaining the set of pseudo-corresponding points M under semantic constraints in step S220, the server will estimate the camera pose of the current image relative to the reference image using geometric methods in step S230, and perform joint optimization through depth consistency to obtain a high-precision camera pose.
[0115] First, the server uses the two-dimensional corresponding point pairs in the matching point set M to estimate the initial relative pose T using either the Five-point Algorithm or the Efficient PnP algorithm. cr That is, the rotation matrix R∈SO(3) of the current image frame relative to the reference image frame and the translation vector
[0116]
[0117] Subsequently, the server introduces the dense rendering consistency constraint proposed in the VGGT method to optimize the color consistency and edge structure consistency between the current image frame and the reference image frame. That is, given the image intrinsic parameter K, the pixels of the reference image frame are back-projected into 3D points:
[0118] P i =D(u) i ,v i )·K -1 [u i ,v i ,1] T
[0119] Then it is passed through the estimated pose T cr Transform to the current image frame and project back onto the image plane to calculate the predicted color. Compared with actual observation Color residuals between:
[0120]
[0121] In addition, edge consistency terms can be introduced. and geometric depth terms Construct a joint optimization objective function:
[0122]
[0123] Where λ c , λ e , λ d This is the weighting factor.
[0124] The server iteratively optimizes the pose parameter T on multi-scale images using backpropagation and image pyramid optimization strategies. cr This yields accurate relative pose estimation results, providing a spatial alignment basis for subsequent mapping and relocalization.
[0125] Step S240: Generation of semantic 3D point cloud and construction of bird's-eye view.
[0126] In step S240, the server combines the camera pose information obtained in step S230 with the image sequence to construct a dense 3D point cloud model, and maps the previously extracted image semantic information onto the 3D points to generate a semantic point cloud; further, it extracts the non-exposed scene structure master plane from the point cloud and generates a semantic bird's-eye view.
[0127] Specifically, the server relies on the VGGT method to construct a unified three-dimensional voxel mesh {v i The pixels in the image are back-projected into a dense set of points and represented as a three-dimensional Gaussian volume.
[0128]
[0129] in, For spatial location, ∑ i To control the density distribution using the covariance matrix, c i This is an RGB color vector. The server projects each pixel in each image frame into 3D space and establishes its differentiable relationship with the voxels based on the voxel differentiable rendering relationship defined in the VGGT network. A joint optimization function is constructed using the color loss, edge loss, and depth loss between the forward-rendered image and the original image from a multi-image frame sequence.
[0130]
[0131] By optimizing voxel parameters through backpropagation, a dense, structured 3D point cloud model with image alignment accuracy is obtained. Subsequently, the server utilizes the correspondence between the image and the 3D points to map the semantic categories in the 2D bounding boxes to the 3D points, thus obtaining a semantic point cloud. Among them l i For the corresponding semantic tags. To further enable the analysis of structures in non-exposed scenes, the server calculates the normal vectors of the point cloud and uses the RANSAC algorithm to fit the ground below and the left and right walls on both sides, respectively, outputting three principal plane sets. During the bird's-eye view generation phase, the server projects these three planes onto a unified two-dimensional raster map coordinate system (u,v), and uses the semantic label that appears most frequently in each raster as the final semantic category for that raster:
[0132] L(u,v)=argmax l Count (u,v) (l)
[0133] Finally, the server outputs a structured semantic overview that can be used for subsequent relocation tasks. Each pixel location contains a clear semantic type and spatial location label. Step S250: Real-time image relocalization and pose fusion estimation.
[0134] After obtaining the bird's-eye view in step S240, the server receives the current-time image frame uploaded by the train terminal device in step S250, and extracts the semantic target region and feature point set from the image using the trained semantic detection model. This is then combined with the established semantic bird's-eye view. The server performs semantic feature matching between the image and the map to determine the most likely observation location. To achieve robust visual matching, the server employs the following processing procedure:
[0135] For the current image frame I q Perform semantic target detection to obtain a semantic target set. Where b i For a two-dimensional bounding box, l i For semantic tags;
[0136] In the bird's-eye view Perform a sliding window search to find the sub-region that is most similar to the semantic structure of the current image frame. By calculating the weighted similarity of overlapping semantic tags:
[0137]
[0138] Where w l For semantic category weights, IoU represents the intersection-union ratio of the corresponding category in the query graph and the reference region. Finally, the position with the highest score is taken as the initial relocation position (x0, y0, θ0).
[0139] In this step, the server will further integrate the short-time trajectory estimation results provided by the inertial measurement unit. The visual positioning results are fused with the inertial navigation results using a Kalman filter or an optimized backend.
[0140]
[0141] Where T vision The location estimate is obtained by matching the bird's-eye view. The final output is the fused location estimate T. fused This represents the high-precision spatial attitude estimation result of the current train, including planar position and orientation information. This method demonstrates good versatility and robustness in various non-exposed structures, and is particularly suitable for non-exposed urban rail transit environments with limited field of vision and strong structural symmetry.
[0142] refer to Figure 3 The semantic bird's-eye view visual relocalization device in the non-exposed scenario includes:
[0143] The image acquisition module 310 is configured to continuously acquire a sequence of non-exposed scene images along the track direction (taking a tunnel scene as an example) during train (metro or train) operation. This module integrates multiple industrial-grade camera systems at the front of the train, possessing high resolution, wide-angle field of view, and high-speed sampling capabilities, and can operate stably in non-exposed, high-dynamic lighting environments. The image acquisition module supports simultaneous acquisition of images from multiple perspectives to enhance scene structure information. The module uploads images to the processing unit via the train's onboard communication system or performs local processing on an embedded device; its acquisition frequency and processing latency can be dynamically adjusted to adapt to task requirements. Furthermore, the module includes an image caching subsystem for short-term storage of consecutive image frames to prevent communication interruptions or frame loss. The module also integrates image preprocessing functions, including but not limited to: geometric distortion correction, brightness normalization, edge enhancement, spatial filtering, size unification, and image compression, to ensure that the images meet the input requirements of subsequent target detection and mapping modules. The semantic recognition and model inference module 320 is configured to input the non-exposed scene (taking a tunnel as an example) images obtained by the image acquisition module 310 into a pre-trained deep neural network to identify representative tunnel components and structures in the images. The core of this module is an object detection model, which has been trained on a large-scale tunnel image dataset and can accurately identify typical object categories including lighting equipment, ventilation openings, signs, cable trays, surveillance cameras, and emergency exit signs. This module standardizes the images to ensure the regularity of the model input; then, it performs forward inference to generate two-dimensional object detection boxes containing location, category, and confidence level, with each box accompanied by a semantic label and score level. During inference, a non-maximum suppression algorithm is used to remove redundant boxes, and a temporal consistency mechanism can be introduced to enhance the stability of object detection. This module also supports online updates and edge deployment, enabling lightweight inference on in-vehicle devices or uploading images to a backend server for high-precision processing.
[0144] The 3D mapping and semantic projection module 330 is configured to construct a spatially dense point cloud using image sequences and map the aforementioned semantic labels onto the 3D model to achieve semantic mapping of non-exposed scenes. The mapping process involves feature extraction and matching between adjacent image frames to complete viewpoint registration and multi-frame camera pose estimation. The module integrates a structure reconstruction algorithm based on image depth information and disparity recovery to obtain a preliminary 3D spatial point distribution. Subsequently, the system calls a high-fidelity mapping engine to construct a voxelized spatial representation, optimizing the density distribution and color information of each voxel unit in space to achieve dense modeling of structures in the non-exposed scene. The model introduces a forward rendering mechanism during the mapping process, establishing a bidirectional mapping relationship between the 3D point cloud and the image, allowing each spatial point to be projected back to the original image frame.
[0145] Using the target bounding boxes and their semantic labels generated by the semantic recognition module, module 330 can backproject each semantic region in the image onto a 3D model and determine its corresponding point cloud range. Through mechanisms such as label coverage, center overlap, and confidence weighting, label information is assigned to the corresponding point cloud, forming a semantic point cloud model with complete structure and semantic attributes. Finally, the module outputs a dense 3D point cloud scene with object category semantic information, which can be used by subsequent planar extraction and bird's-eye view generation modules.
[0146] The main structure extraction and bird's-eye view generation module 340 is configured to extract the main structural planes and perform two-dimensional semantic mapping on the semantically assigned 3D point cloud model. The core objective of this module is to accurately extract the three main structural planes—the ground, left wall, and right wall—from the non-exposed scene point cloud and visualize their semantic information in a two-dimensional bird's-eye view for subsequent relocalization tasks. This module first uses point cloud normal vector calculation and spatial clustering algorithms to geometrically segment the point cloud and identifies potential main structural regions based on normal directions (including vertical downwards, horizontal leftwards, and rightwards) and spatial projection relationships. Based on this, it optimizes and selects the planar structures with the most points and flat features, and combines structural semantic distribution density as an auxiliary constraint to finally determine the three-dimensional boundaries of the three main plane regions. For each main plane region, the module maps it to a unified two-dimensional coordinate system using orthogonal projection and constructs a unified spatial raster representation model. The final output bird's-eye view not only contains spatial structural information but also includes semantic category encoding for each location, providing the relocalization module with both structural and semantic priors.
[0147] The relocalization and pose estimation module 350 is configured to achieve visual relocalization in non-exposed scenes based on the structural and semantic correspondence between the current image frame information and the constructed semantic bird's-eye view. This module combines image features and spatial structure, employing a multi-level matching strategy to achieve highly robust pose estimation under structural prior guidance. The module extracts the target detection results and their corresponding image feature descriptors from the current image frame, and projects the extracted semantic targets into 3D space through a camera model to construct a local point cloud or semantic structure map of the current frame. Subsequently, the module matches the local representation of the current frame with known semantic structure regions in the bird's-eye view based on semantic similarity and spatial structure distribution. A semantic weighting mechanism is introduced in the matching process, prioritizing the matching of regions containing key objects (such as traffic lights, cable trays, emergency exits, etc.). After matching, the camera's rotation and translation parameters are calculated based on the structural feature point set to complete the spatial pose estimation of the current frame. The final output includes the current camera's six-DOF spatial pose, optional trajectory curves, positioning confidence, and registration status with the semantic map.
[0148] Figure 4 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0149] The processor 1010 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0150] The memory 1020 can be implemented in the form of read-only memory (ROM), random access memory (RAM), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0151] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0152] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (e.g., USB, Ethernet cable) or wireless means (e.g., mobile network, Wi-Fi, Bluetooth). The bus 1050 includes a pathway for transmitting information between the various components of the device (e.g., processor 1010, memory 1020, input / output interface 1030, and communication interface 1040).
[0153] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0154] The electronic devices described above are used to implement the semantic bird's-eye view visual relocation method in the corresponding non-exposed scenario in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0155] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the semantic bird's-eye view visual relocation method in a non-exposed scene as described in any of the above embodiments.
[0156] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0157] The aforementioned non-transitory computer-readable storage media can be any available medium or data storage device that a computer can access, including but not limited to magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MOs), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs)).
[0158] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the semantic bird's-eye view visual relocation method in a non-exposed scene as described in any of the embodiments in the exemplary method section above, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0159] Based on the same inventive concept, corresponding to the semantic bird's-eye view visual relocation method in non-exposed scenarios described in any of the above embodiments, this disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the semantic bird's-eye view visual relocation method in non-exposed scenarios. Corresponding to the execution entity for each step in each embodiment of the semantic bird's-eye view visual relocation method in non-exposed scenarios, the processor executing the corresponding step can belong to the corresponding execution entity.
[0160] The computer program product of the above embodiments is used to enable the computer and / or the processor to execute the semantic bird's-eye view visual relocation method in non-exposed scenarios as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0161] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, method, or computer program product. Therefore, this disclosure can be implemented as entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this disclosure can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.
[0162] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example,, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (not exhaustive) of a computer-readable storage medium may include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.
[0163] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0164] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0165] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Python, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0166] It should be understood that each block of a flowchart and / or block diagram, as well as combinations of blocks in a flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine that, when executed by a computer or other programmable data processing device, creates means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0167] These computer program instructions may also be stored in a computer-readable medium that enables a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce a product comprising an instruction apparatus that implements the functions / operations specified in the boxes of a flowchart and / or block diagram.
[0168] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions that execute on the computer or other programmable apparatus can provide a process for implementing the functions / operations specified in the boxes of a flowchart and / or block diagram.
[0169] Furthermore, although the operations of the methods of this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Rather, the steps depicted in the flowcharts may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0171] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0172] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0173] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0174] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0175] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
[0176] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be interpreted in the broadest sense, thereby encompassing all such modifications and equivalent structures and functions.
Claims
1. A semantic bird's-eye view visual relocalization method in non-exposed scenes, characterized in that, include: Acquire multi-view image sequences in non-exposed scenes (including but not limited to tunnels, underground utility tunnels, or underground parking lots); The image is semantically recognized based on a pre-trained semantic target detection model to obtain a two-dimensional target region containing semantic categories. Semantic-geometric dual-channel fusion mechanism combining semantic mask and geometric constraints is used to extract semantic features with spatial consistency. The image sequence was reconstructed in three dimensions using the Volumetric Gaussian Grid Transformer (VGGT) method to generate a dense three-dimensional semantic point cloud model that integrates semantics and geometric structure. Based on the projection and back-projection relationship between the image and the 3D point cloud, the 2D semantic information is mapped to the 3D space to realize the semantic assignment of the point cloud; Based on the semantic point cloud, the main structural planes such as the ground, left wall and right wall are extracted, and a two-dimensional semantic bird's-eye view with semantic annotations is generated. A reciprocal matching strategy guided by semantic masks is used to perform feature matching and pose fusion estimation in semantic regions of the same type. Finally, the bird's-eye view and the matching strategy are used to perform scene relocalization, resulting in a relocalization result that is highly accurate, robust, and semantically interpretable. Based on the edge-cloud collaborative architecture, the system achieves sub-second real-time response while ensuring centimeter-level positioning accuracy through data layering and load balancing technology.
2. The method according to claim 1, characterized in that, The process of semantic recognition of images based on the object detection model includes: The received reference and query images are preprocessed to obtain images of uniform size and lighting conditions; The image is identified using a pre-trained semantic object detection model, and two-dimensional target regions containing semantic categories are extracted. Introducing semantic confidence constraints during semantic feature extraction improves the accuracy of target region recognition.
3. The method according to claim 1, characterized in that, The construction of the three-dimensional semantic point cloud model includes: The pose parameters of a multi-frame camera are estimated by feature matching and viewpoint registration of multiple image sequences. Voxel modeling networks based on image depth information generate 3D point clouds containing spatial coordinates, depth values, and frame index information; By using a differentiable rendering optimization mechanism, the color, density, and spatial consistency of the point cloud are optimized through backpropagation to obtain a high-fidelity dense semantic point cloud.
4. The method according to claim 3, characterized in that, The 3D modeling model based on image depth information includes: Construct a unified three-dimensional voxel mesh; Each cell in the voxel mesh is represented as a three-dimensional Gaussian body with spatial location, density distribution, and color attributes; Project the pixels in the image sequence into voxel space to establish a differentiable rendering relationship between the image and voxels. A differentiable optimization loss function is defined based on the consistency of the rendered image with the original image in terms of color, edge, and depth. The density, position, and color parameters of the voxel Gaussian volume are optimized using the backpropagation algorithm, resulting in a dense 3D semantic point cloud model that satisfies image constraints.
5. The method according to claim 1, characterized in that, The correspondence between points in the image and the 3D point cloud includes: Using the camera intrinsic and extrinsic parameters obtained during the 3D modeling process, the projection and back-projection relationships between the image pixels and the 3D point cloud points are calculated and established. Each point in the three-dimensional point cloud model is back-projected onto the image plane to determine whether it is within a certain two-dimensional semantic target region. The semantic label is determined based on the semantic confidence in the image corresponding to the point, and the corresponding semantic information is assigned when the confidence is higher than a set threshold.
6. The method according to claim 1, characterized in that, The process of extracting the main planar areas of the ground, left wall, and right wall and generating a semantic bird's-eye view includes: Calculate the normal vector information of the point cloud and perform principal axis clustering based on the normal direction; The random sample consensus algorithm is used to fit the planes in the non-exposed scene to identify the main planes vertically downward, to the left and to the right, namely the ground, the left wall and the right wall. The plane with the most points is selected as the main plane for output, and its semantic annotation information is retained; The three principal planes are projected onto a unified two-dimensional map coordinate system; Each projection point is divided into grids and its semantic category is counted. A semantic annotation layer is generated using the maximum frequency method. Output a structured semantic bird's-eye view that can be used for localization and path planning.
7. The method according to claim 1, characterized in that, The scene relocation process includes: Obtain the image at the current moment and extract the semantic target; Based on a semantic mask-guided reciprocal matching strategy, the positional structure and semantic features in an existing semantic bird's-eye view are matched by constructing local point clouds or image features. Estimate the spatial pose of the current device to obtain high-precision visual repositioning results in a non-exposed environment.
8. A semantic bird's-eye view visual relocalization method in non-exposed scenes, characterized in that, include: The image acquisition module is configured to acquire multi-view image data in non-exposed scenes; The semantic recognition and feature extraction module is configured to perform semantic detection on the image using a pre-trained semantic target detection model and extract target regions containing semantic categories; The 3D mapping and semantic projection module is configured to generate a dense 3D semantic point cloud based on the VGGT model and project the image semantic information onto the corresponding spatial location. The main structure extraction and bird's-eye view generation module is configured to extract the ground and left and right wall main planes from the semantic point cloud and generate a two-dimensional semantic bird's-eye view with semantic annotations. The relocalization and pose estimation module is configured to match the semantic features of the newly acquired image with the bird's-eye view and output the visual relocalization result. The edge-cloud collaborative architecture is configured to collaboratively process the functions of each module, and improves the real-time response capability of the system through data layering and load balancing optimization.
9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method as described in any one of claims 1 to 8.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method as described in any one of claims 1 to 8.
11. A computer program product, characterized in that, It includes computer program instructions that, when run on a computer, cause the computer to perform the method as described in any one of claims 1 to 8.
Citation Information
Cited By
Component identification method and system based on three-dimensional point cloud
CN121686438A
Power equipment three-dimensional fault positioning method and system based on semantic voxel and projection mapping
CN121788602A
Tunnel component identification method and device based on point cloud data and priori knowledge fusion, terminal equipment and storage medium
CN121902281A
Tunnel component recognition method and device based on fusion of point cloud data and prior knowledge, terminal equipment and storage medium
CN121902281B