Positioning method and device based on grid map and feature matching, equipment and medium
By integrating obstacle information with visual features in a map matching method, the problem of insufficient positioning accuracy and robustness of traditional single sensors in complex environments is solved, and high-precision and stable positioning of robots in complex environments is achieved.
Patent Information
- Application Number
- CN202510800442.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional single-sensor positioning methods suffer from reduced accuracy and insufficient robustness in complex or dynamic environments, especially in areas with dense obstacles, scarce texture information, or drastic lighting changes, making it difficult to achieve high-precision and stable positioning.
By fusing obstacle information with environmental visual features, constructing obstacle grid maps and visual feature maps, and using grid unit areas for feature matching, accurate estimation of the robot's current position is achieved. Combining multi-source data and deep learning algorithms improves matching efficiency and accuracy.
The robot's positioning accuracy and system robustness are improved in complex and dynamic environments, and its adaptability and real-time response capabilities in different scenarios are enhanced.
Smart Images

Figure CN120668133A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robotics, and in particular to a positioning method, apparatus, device, and storage medium based on grid map and feature matching. Background Art
[0002] With the widespread application of mobile robots in intelligent manufacturing, autonomous driving, security inspection and other fields, how to achieve high-precision and robust environmental positioning has become a key technology for mobile robots.
[0003] Traditional positioning methods primarily rely on data from a single sensor, such as lidar-based SLAM (Simulation and Localization) or vision-based VSLAM (Visual Localization and Localization). However, single sensors often suffer from reduced accuracy or data loss in complex environments, limiting their adaptability to dynamic or highly occluded scenes. Grid maps, a common approach to environmental modeling, effectively represent the distribution of obstacles in space, providing a foundation for path planning and obstacle avoidance. However, grid maps only contain geometric information about obstacles and lack a description of the visual features of the scene, making it difficult to support high-precision position inference. Summary of the Invention
[0004] This application provides a positioning method, device, equipment, and storage medium based on grid map and feature matching. By integrating obstacle information detected by the robot with the visual features of the current environment, and using a joint matching mechanism between the obstacle grid map and the visual feature map, the robot's current position can be accurately estimated. The technical solution described in this application effectively improves the robustness and accuracy of the positioning system in environments with complex structures or varying lighting conditions, and has excellent versatility and practical value.
[0005] In a first aspect, the present application provides a positioning method based on a grid map and feature matching, comprising:
[0006] Obtaining obstacle information corresponding to an obstacle detected by the robot, and identifying a grid unit area where the obstacle is located in a pre-acquired obstacle grid map based on the obstacle information;
[0007] Acquire a current environment visual image collected by the robot, and extract current environment visual features from the current environment visual image;
[0008] In the grid unit area of the pre-generated visual feature map, the current environment visual features are used for matching to obtain a map matching result of the current environment visual features in the visual feature map, and the current position of the robot is determined based on the map matching result.
[0009] In a second aspect, the present application provides a positioning device based on a grid map and feature matching, comprising:
[0010] A grid unit identification module is used to obtain obstacle information corresponding to an obstacle detected by the robot, and identify the grid unit area where the obstacle is located in a pre-acquired obstacle grid map based on the obstacle information;
[0011] An environmental feature extraction module is used to obtain a current environmental visual image collected by the robot and extract current environmental visual features from the current environmental visual image;
[0012] The robot positioning module is used to match the current environment visual features in the grid unit area of the pre-generated visual feature map using the current environment visual features to obtain a map matching result of the current environment visual features in the visual feature map, and determine the current position of the robot based on the map matching result.
[0013] In a third aspect, the present application provides a positioning device based on a grid map and feature matching, comprising:
[0014] one or more processors;
[0015] A memory stores one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the positioning method based on grid map and feature matching as described in the first aspect.
[0016] In a fourth aspect, the present application provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the positioning method based on grid map and feature matching as described in the first aspect.
[0017] In this application, by combining obstacle information with visual feature information to construct an obstacle grid map and a visual feature map, and then matching the current environment image with local areas of the visual feature map, high-precision positioning of the robot's current position is achieved. This not only integrates the environment's geometric structure and image texture information, but also improves matching efficiency and accuracy by limiting the grid cell area, thereby maintaining the stability and robustness of the positioning system in dynamic or complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flowchart of a positioning method based on grid map and feature matching provided in an embodiment of the present application;
[0019] Figure 2 This is a flowchart of the current environment visual feature extraction provided by the embodiment of the present application;
[0020] Figure 3 This is a flow chart of dynamic region recognition provided by an embodiment of the present application;
[0021] Figure 4 This is a flow chart of dynamic region recognition provided by an embodiment of the present application;
[0022] Figure 5 This is a flowchart of robot positioning provided by an embodiment of the present application;
[0023] Figure 6 This is a structural diagram of a positioning device based on grid map and feature matching provided in an embodiment of the present application;
[0024] Figure 7 This is a structural diagram of a positioning device based on grid map and feature matching provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are merely used to explain the present application and are not intended to limit the present application. It should also be noted that, for ease of description, only portions related to the present application, not all of the contents, are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the various operations (or steps) as being processed sequentially, many of the operations therein can be performed in parallel, concurrently or simultaneously. In addition, the order of the various operations can be rearranged. The process can be terminated when its operations are completed, but may also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0026] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the data used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects connected before and after are in an "or" relationship.
[0027] With the continuous development of mobile robotics technology, robots have been widely used in warehousing and logistics, indoor service, intelligent inspection, and other fields. To ensure their stable operation in complex environments, high-precision positioning systems have become the core foundation for autonomous navigation and environmental interaction. However, the complexity of environmental structures varies significantly across different application scenarios. In particular, in areas with dense obstacle distribution, scarce texture information, or drastic lighting changes, traditional positioning solutions based on single sensors face problems such as reduced accuracy and insufficient robustness.
[0028] Existing positioning methods often rely on geometric maps constructed using lidar or feature maps constructed using purely visual images. Because a single data source cannot fully perceive environmental changes, positioning drift and matching failures are prone to anomalies when environments are highly repetitive or texture features are insufficient. In dynamic environments, such as those with frequent personnel movements and drastically changing scene occlusions, existing systems lack dynamic adaptability, making it difficult to achieve sustained and stable positioning. Furthermore, some methods are computationally intensive and suffer from poor real-time performance, making them unsuitable for resource-constrained scenarios or those requiring high response speeds.
[0029] Therefore, there is an urgent need to propose a positioning method and system that integrates multi-source environmental information and has the ability of area-limited matching and rapid position back-calculation, so as to improve the robot's self-positioning accuracy, stability and real-time response capability in complex dynamic environments, thereby enhancing its task execution efficiency and scene adaptability.
[0030] To address the aforementioned issues, this embodiment provides a positioning method based on grid maps and feature matching. This method aims to build a multidimensional environmental perception model by fusing obstacle information with environmental visual features, and implements feature positioning matching at the grid cell level to improve the robot's positioning accuracy and system robustness in complex and changing environments. This method maps obstacle information perceived by the robot to corresponding grid regions in an obstacle grid map, extracts environmental features from the current visual image, and performs feature comparisons within the corresponding grid regions of the visual feature map. This method quickly and accurately infers the robot's current position, effectively mitigating positioning deviations caused by single sensor errors and improving the stability and adaptability of the overall navigation system.
[0031] The positioning method based on raster map and feature matching provided in this embodiment can be performed by a positioning device based on raster map and feature matching. The positioning device based on raster map and feature matching can be implemented via software and / or hardware. The positioning device based on raster map and feature matching can be composed of two or more physical entities, or a single physical entity. For example, the positioning device based on raster map and feature matching can be an operation and maintenance server used to maintain normal business operations.
[0032] The positioning device based on raster map and feature matching is installed with at least one operating system, including but not limited to Android, Linux, and Windows. The positioning device based on raster map and feature matching can install at least one application based on the operating system. The application can be a native application of the operating system or an application downloaded from a third-party device or server. In this embodiment, the positioning device based on raster map and feature matching has at least one application capable of executing the positioning method based on raster map and feature matching.
[0033] For ease of understanding, this embodiment is described by taking the operation and maintenance server as the main body for executing the positioning method based on the grid map and feature matching as an example.
[0034] Figure 1 A flowchart of a positioning method based on grid map and feature matching provided by an embodiment of the present application is given. Figure 1 , the positioning method based on grid map and feature matching specifically includes:
[0035] S110 : Obtain obstacle information corresponding to an obstacle detected by the robot, and identify a grid unit area where the obstacle is located in a pre-acquired obstacle grid map based on the obstacle information.
[0036] In some embodiments, to fully perceive the spatial structure of the robot's surroundings, a multi-type environmental perception sensor module integrated into the robot body performs joint data acquisition. These environmental perception sensor modules include heterogeneous sensor components such as two-dimensional or three-dimensional lidar, ultrasonic ranging modules, and structured light or time-of-flight depth cameras. A multi-source data fusion strategy is used to acquire real-time spatial distribution information of potential obstacles in the forward environment. Key geometric and dynamic properties associated with each obstacle are extracted to comprehensively construct a feature set describing the obstacles. After acquiring the obstacle information, the obstacle feature data is normalized according to spatial coordinate system conversion rules and mapped based on a preset obstacle grid map corresponding to the robot's physical space. During the mapping process, the spatial location of the obstacle is accurately projected onto the corresponding grid cell in the grid coordinate system based on the grid map's spatial resolution, origin reference system, and directional projection settings. Subsequently, combining the existing static spatial distribution data in the map with the grid cell's status label, algorithms such as Euclidean distance-based nearest neighbor search, minimum error fitting, and shape reconstruction matching are used to compare and analyze the position and morphological characteristics of the obstacle, achieving precise binding identification between the obstacle and the grid cell. In terms of grid division, a regular grid or adaptive resolution division strategy is adopted, dynamically adjusting the spatial scale of each cell based on scene complexity and navigation accuracy requirements. Each grid cell corresponds one-to-one with its unique spatial location in the map projection coordinate system and includes multidimensional attribute fields representing current occupancy status, obstacle type, and historical observation probability. This mechanism achieves a high-precision, multi-dimensional, and continuously updated spatial representation of obstacles in the map model, providing a stable spatial reference foundation for subsequent visual feature matching and positioning calculations.
[0037] S120: Acquire a current environment visual image collected by the robot, and extract current environment visual features from the current environment visual image.
[0038] In some embodiments, the robot's visual perception module may include devices such as a wide-angle camera, a stereo binocular camera, an RGB-D depth camera or a fisheye camera to meet the needs of wide viewing angle, multi-scale and depth perception under different scene conditions. During the operation of the robot, the visual perception module collects the current environment image data in real time at a set frame rate, and transmits the image data to the local processing unit or edge computing module for pre-processing operations, including image denoising, illumination equalization, resolution resampling and parallax correction, etc., to ensure the stability and accuracy of subsequent feature extraction. After the image preprocessing is completed, a feature extraction network based on deep learning or a traditional image processing algorithm is used to perform feature extraction operations on the current environment image. The extracted visual features include local key point information, global descriptors, and semantic segmentation results of the image, texture patterns and edge contour information. For different types of map matching strategies, the appropriate feature representation method is dynamically selected, and the extraction results are normalized, dimensionally reduced and quality evaluated to generate a stable and highly discriminative set of current environment visual features. The visual features extracted in this way are not only highly responsive to salient structures in the image, but also have strong scale invariance and rotation robustness. They can adapt to the robot's perception changes under different angles, distances and lighting conditions, providing a reliable visual perception foundation for subsequent map matching and positioning calculations.
[0039] S130. In the grid unit area of the pre-generated visual feature map, matching is performed using the current environment visual features to obtain a map matching result of the current environment visual features in the visual feature map, and the current position of the robot is determined based on the map matching result.
[0040] In some embodiments, candidate matching regions are first defined within a visual feature map based on the grid cell regions corresponding to identified obstacles, thereby narrowing the search range and improving matching efficiency and robustness. The visual feature map is constructed during deployment or during previous robot operations. It divides the space into multiple grid cells, each associated with a set of typical visual features within that region, including texture features of static structures, keypoint distributions, depth map templates, or semantic label information. During the matching process, feature similarity calculation methods, such as those based on Euclidean distance, Hamming distance, or similarity scoring mechanisms derived from deep neural network outputs, are used to compare the current environment's visual features with historical features stored within the target grid region. When multiple matching candidates are present, spatial consistency verification mechanisms, such as RANSAC verification, geometric consistency filtering, and reprojection error assessment, are further implemented to eliminate anomalous matches and enhance the reliability and accuracy of the positioning results. Ultimately, the candidate location with the highest matching score that satisfies the spatial constraints is used as the estimated robot's current position. This is then fused and corrected with inertial navigation or wheel speedometer data to achieve a more accurate pose solution. By matching the visual features of the current environment in the visual feature map, the robot can quickly and accurately complete the positioning operation in the visual feature map in environments with rich feature information, complex structures or partial occlusion, providing precise spatial reference for its subsequent path planning and task execution.
[0041] Optionally, Figure 2 This is a flowchart of the current environment visual feature extraction provided by the embodiment of this application. Figure 2 As shown, the steps of extracting visual features of the current environment specifically include S1201-S1203:
[0042] S1201: Acquire a cluster of environmental visual images collected by the robot in the same time window as the current environmental visual image.
[0043] For example, the multi-camera visual perception system carried by the robot includes multiple distributed cameras, such as front-view, side-view and rear-view cameras, which ensure the collection of a set of continuous and multi-perspective environmental image data within the same time range through a synchronous trigger mechanism or timestamp alignment technology. The environmental visual image cluster not only covers the spatial information of different perspectives around the robot, but also reflects the dynamic changes of the environment during the time period. After the acquisition is completed, the image cluster is pre-processed, including time synchronization verification, image denoising, distortion correction and resolution standardization, to ensure the quality and consistency of multi-source image data. The environmental visual image cluster data provides a rich information basis for subsequent multi-perspective feature fusion, environmental modeling and dynamic object detection, which helps to improve the perception accuracy and positioning stability of the robot in complex environments.
[0044] S1202: Determine a dynamic area where a dynamic obstacle is located in the current environmental view image based on the environmental view image cluster.
[0045] For example, by analyzing and spatially registering multi-view image data, a method combining optical flow estimation, background modeling, and target detection and tracking algorithms is used to detect moving targets in an environment. First, pixel-level motion vectors are calculated between sequential frames to extract motion information of potentially dynamic regions. Then, deep learning semantic segmentation models or target detection networks, such as YOLO and Mask R-CNN, are combined to identify the categories and boundaries of dynamic objects and precisely locate the specific regions of dynamic obstacles within the image. The spatial geometric relationships between multi-view images are exploited to fuse and correct the 3D positions of dynamic obstacles, mitigating the effects of single-view occlusions and false detections. The identified dynamic regions are then labeled and isolated from the current visual image, providing a basis for subsequent visual feature extraction and matching to distinguish between dynamic and static environments, thereby improving positioning accuracy and robustness. Accurate identification of dynamic regions is crucial for avoiding feature mismatches and positioning drift caused by dynamic obstacles.
[0046] S1203: Determine a current environment static image according to the current environment field of view image and the dynamic area, and extract current environment visual features from the current environment static image.
[0047] For example, regions marked as dynamic obstacles in the current visual image of the environment are segmented and masked. Image inpainting algorithms are then used to complete the content of the masked regions, restoring the continuity and integrity of the background static scene. Simultaneously, by combining the multi-perspective information of the environmental visual image cluster, perspective fusion and image stitching techniques are employed to enhance the spatial consistency and detail restoration of the static image. After obtaining a high-quality static image of the current environment, a multi-scale feature extraction algorithm is used to conduct an in-depth analysis of key static structures within the image. The extracted visual features include local keypoints, texture descriptors, and global feature vectors extracted by a deep convolutional neural network. To improve feature stability and discriminability, the extracted visual features are normalized, feature selected, and dimensionality reduced. Potential noise features introduced by dynamic region processing are also removed, resulting in a visual feature set that accurately reflects the static environment structure. These steps effectively isolate dynamic interference and enhance the representation of static environmental information, providing more reliable basic data for subsequent visual feature map matching and robot positioning, thereby improving positioning accuracy and system robustness.
[0048] Optionally, Figure 3 This is a flow chart of dynamic area recognition provided by the embodiment of this application. Figure 3 As shown, the steps of dynamic area identification specifically include S12021-S12022:
[0049] S12021. Calculate the environmental image frame difference between two temporally adjacent frames of environmental visual images in the environmental visual image cluster, and obtain a plurality of adjacent frame pixel residual maps based on the environmental image frame difference.
[0050] For example, the image frame difference between two temporally adjacent frames in a cluster of environmental visual images is calculated to capture temporal changes in the environment. First, the two adjacent frames are time-synchronized and spatially aligned to ensure precise pixel-to-pixel matching. Subsequently, a pixel-by-pixel subtraction operation is performed to calculate the brightness difference at the same pixel position in the two frames, generating a preliminary image frame difference map. To mitigate the effects of ambient lighting changes and sensor noise on the difference calculation, the frame difference map is filtered using methods such as Gaussian filtering and median filtering to remove isolated noise points and subtle jitter. Based on this processing, a pixel residual map is generated between several adjacent frames, reflecting the motion areas and dynamic change trends of objects in the environment. This residual map not only reveals the motion trajectory of dynamic obstacles but also provides critical motion information for subsequent dynamic region identification and segmentation. Combined with time series analysis techniques, the system can model the change patterns in the residual map, further improving the accuracy and real-time performance of dynamic target detection.
[0051] S12022. Determine a dynamic area where a dynamic obstacle is located in the current environment visual image based on the pixel residual images of the plurality of adjacent frames.
[0052] For example, the dynamic region of the current environment visual image containing dynamic obstacles is identified and determined based on pixel residual maps from several adjacent frames. The residual maps are temporally fused and spatially aggregated, and weighted overlaid across multiple frames to enhance the continuity and saliency of the dynamic region while suppressing isolated noise and short-term flicker errors. Continuous dynamic region candidate blocks are extracted using threshold segmentation and morphological processing methods such as dilation and erosion. To improve the accuracy of the dynamic region, a semantic segmentation model based on a convolutional neural network is combined to perform semantic recognition on the candidate regions, distinguishing moving obstacles from static background structures and eliminating falsely detected static texture changes. Combined with the results of multi-view image registration, the dynamic region is spatially geometrically corrected and 3D reconstructed to precisely locate the range and boundaries of the dynamic obstacle in the 3D environment. Through the computational process of the pixel residual map, the dynamic region containing the dynamic obstacle is accurately extracted, providing reliable dynamic information support for subsequent visual feature screening and location calculations.
[0053] Optionally, the determining, based on the pixel residual images of the plurality of adjacent frames, a dynamic area where a dynamic obstacle is located in the current environmental visual image includes:
[0054] Residual pixel regions whose corresponding pixel values exceed a preset residual threshold are extracted from the pixel residual images of each adjacent frame.
[0055] Exemplarily, a dynamic threshold is first set for the residual image. The threshold is adaptively adjusted according to the changes in ambient light, the sensor noise level and historical dynamic information to balance sensitivity and robustness. Subsequently, each residual image is binarized using a threshold segmentation algorithm, and pixels above the threshold are marked as potential dynamic areas. To reduce the impact of noise, morphological filtering operations are applied to the extracted residual pixel areas, including dilation, erosion, and opening and closing operations, to remove isolated noise points and fill small holes between areas to form coherent and regular dynamic obstacle candidate areas. At the same time, the spatial consistency of the residual areas of multiple frames within the time window is combined and filtered to further improve the continuity and stability of the dynamic area. The residual area extraction strategy effectively enhances the detection capability of small or partially occluded dynamic obstacles, and provides an accurate preliminary judgment basis for subsequent dynamic area positioning and visual feature elimination.
[0056] A dynamic area where a dynamic obstacle is located in the current environmental visual image is determined based on each of the residual pixel areas.
[0057] For example, based on each residual pixel region, spatial aggregation and temporal analysis methods are used to comprehensively determine and identify the dynamic regions within the current environmental visual image where dynamic obstacles reside. First, connectivity analysis is performed on the extracted residual pixel regions to identify candidate dynamic regions that exhibit spatial continuity and satisfy area and shape constraints. Subsequently, a time series aggregation algorithm is used to track and predict the motion trajectories and change trends of these candidate regions, combining multi-frame residual data. This eliminates misclassified regions due to illumination variations, sensor noise, or static texture changes. Furthermore, a deep learning model is used for semantic recognition of dynamic regions. A trained semantic segmentation network is then used to further refine the boundaries of dynamic obstacles and distinguish between moving objects and static parts of the environment. By combining multi-view image correction and 3D geometric constraints, dynamic regions are accurately located and described in both the 2D image space and the 3D environmental model. Finally, the identified dynamic regions serve as the basis for subsequent visual feature extraction and matching, effectively preventing the negative impact of dynamic interference on positioning accuracy and improving the stability and reliability of the overall positioning system.
[0058] Optionally, the determining, based on each of the residual pixel regions, a dynamic region where a dynamic obstacle is located in the current environmental visual image includes:
[0059] A residual region weight of the residual pixel region is obtained based on a time interval between two adjacent frames of environmental visual images corresponding to the residual pixel region and the current environmental visual image.
[0060] For example, we first obtain the acquisition time of the source image frame pair corresponding to each residual pixel region based on the timestamp information, and calculate the time interval Δt between these frame pairs and the current image frame as an important input for the time decay factor. Then, based on the assumption of motion continuity of dynamic targets in visual image sequences, we construct a time decay function model to weight the effectiveness of the residual pixel region. The weight function can be expressed as:
[0061] w=exp(-λ·Δt)
[0062] Among them, λ is the decay rate factor, which is used to control the sensitivity of time to dynamic judgment.
[0063] At the same time, the weights are normalized and comprehensively adjusted based on factors such as the maximum residual amplitude, duration, and spatial stability of the residual region in its corresponding frame pair to improve the accuracy of dynamic region assessment. The resulting residual region weight not only reflects the credibility of the region as a dynamic obstacle indicator in the current frame, but can also be used in subsequent dynamic region fusion, false motion filtering, and abnormal feature removal processes, enhancing the positioning system's robustness to time lag and visual interference.
[0064] Based on each of the residual pixel regions and its corresponding residual region weight, a dynamic region where a dynamic obstacle is located in the current environmental visual image is determined.
[0065] Exemplarily, based on each residual pixel region and its corresponding residual region weight, potential dynamic obstacle regions in the current visual image are weighted and fused, and their credibility is assessed, thereby accurately determining the location and boundaries of dynamic regions. First, residual pixel regions extracted from the residual images of multiple adjacent frames are mapped to the unified coordinate system of the current visual image. Each region's contribution to the overall dynamic region determination is weighted based on its corresponding time-decay weight. Based on this, a residual weighted heat map construction method is employed to overlay the weight distributions of each residual region to generate a dynamic response heat map. In the heat map, the dynamic response value of each pixel reflects its frequency of being labeled as a "changing pixel" across multiple residual regions and its credibility. Subsequently, a multi-threshold segmentation strategy is used to process the heat map, combining edge detection and connected domain analysis to extract high-response regions. Furthermore, dynamic morphological operators and spatial consistency judgment are used to filter out isolated noise regions and misclassified blocks, retaining salient regions with stable spatial structure and continuous temporal features. To improve the accuracy of dynamic region recognition, a dynamic confidence factor based on a spatiotemporal consistency model is introduced. This factor jointly models the residual region weight with the region's motion trajectory coherence and contour stability within the image. The resulting dynamic region is presented as a high-confidence pixel set, accurately locating dynamic obstacles within the current visual image and providing a reliable basis for interference removal in the subsequent static feature extraction and map matching stages.
[0066] Optionally, Figure 4 This is a flow chart of dynamic area recognition provided by the embodiment of this application. Figure 4 As shown, the steps of dynamic area identification specifically include S12023-S12024:
[0067] S12023. Determine, based on the environmental view image cluster, several changing regions where dynamic obstacles are located in the current environmental view image.
[0068] For example, multiple frames in an image cluster are temporally aligned and enhanced to eliminate interference caused by slight viewing angle deviation or imaging noise. Frame differences and pixel-level residuals are calculated between adjacent frames to construct an inter-frame change map sequence, generating initial temporal image response information. Within this change map sequence, a spatial clustering algorithm is introduced to classify regions of significant pixel change in consecutive frames. Combined with a temporal sliding window mechanism, this algorithm identifies change segments that appear consistently across multiple frames and exhibit motion consistency and spatial continuity. Each change segment is considered a candidate change region, and its stability is evaluated through temporal analysis of its boundary structure, area change trends, and center trajectory. This eliminates pseudo-dynamic responses caused by brief flickers or localized jitter. Redundant information within the image cluster is leveraged to perform multi-frame region verification, comparing the dynamic response characteristics of the same physical region from different perspectives. The authenticity of the change regions is determined by region overlap and morphological similarity. A set of change regions is then selected and output, accurately covering the potential distribution range of all dynamic obstacles in the current image, laying the foundation for subsequent static region extraction and feature stability assessment.
[0069] S12024. Perform pixel filling on the plurality of changed areas to obtain a dynamic area where the dynamic obstacle is located.
[0070] For example, after completing the change region extraction, edge contour analysis is performed on each change region to extract its boundary pixels and construct an initial region mask. Because the initial residual detection and change clustering processes may result in holes, broken boundaries, or partial omissions in the region, a pixel filling mechanism is introduced to repair the structure and enhance connectivity of each region. The filling process uses a boundary tracing and closed contour reconstruction method to fit the outer contour of the change region, close irregular boundaries, and generate a closed region model. Morphological operations such as dilation and closing are used to complete missing pixels near the boundary and fill small holes and gaps within the region to improve regional integrity. The filling process considers the consistency characteristics of pixels within the region, using pixel value statistics and spatial consistency to determine whether pixels are misclassified and remove structural noise. Regional connectivity and minimum area constraints are introduced to filter out small isolated areas caused by detection errors, retaining significant change blocks that are stable across multiple frames and whose area and boundary continuity meet preset thresholds. Finally, the change region after pixel filling is defined as the dynamic region of the dynamic obstacle in the current image, which is used for the subsequent static region generation and visual feature screening steps.
[0071] Optionally, Figure 5 This is a flow chart of robot positioning provided by the embodiment of the present application. Figure 5 As shown, the steps of robot positioning specifically include S1301-S1302:
[0072] S1301: Extract the visual coordinate offset of the current environment visual feature in the visual feature map from the map matching result.
[0073] Exemplarily, after completing the matching of the current environment visual features with the target area in the pre-generated visual feature map, the spatial geometric relationship between the matching pairs is analyzed to quantify the relative displacement of the observation point in the current image relative to the map reference feature point. Based on the key point pair information contained in the matching results, a set of visual matching feature points with high confidence is selected, and the position coordinates of these features in the current image coordinate system and the visual feature map coordinate system are extracted respectively. Then, the matching point pairs are spatially aligned and analyzed using least squares fitting, RANSAC robust estimation, or affine / perspective geometric transformation models to obtain the geometric transformation parameters of the overall matching area. The visual coordinate offset is the translation vector between the current environment visual image coordinate system and the local coordinate system of the visual feature map, which represents the degree of change of the robot in the visual domain relative to the map reference position. The offset is multi-dimensionally decomposed (such as horizontal offset Δx, vertical offset Δy, depth drift Δz, etc.) to support subsequent pose estimation and accuracy evaluation. To improve robustness, matching residual weights and feature point distribution density are introduced to optimize the extraction results, ensuring that accurate and stable visual coordinate offset data can be obtained even in cases of uneven feature distribution or partial occlusion. The offset serves as a key intermediate quantity and provides a quantitative input basis for the final precise positioning of the robot.
[0074] S1302: Construct a rigid transformation matrix between the robot and the visual feature map based on the map matching result, and calculate the current position of the robot according to the visual coordinate offset and the rigid transformation matrix.
[0075] Exemplarily, after matching the current environment's visual features with the visual feature map and extracting the visual coordinate offset, a spatial geometric mapping relationship is established between the current visual observation reference system and the global coordinate system of the map. Utilizing the spatial distribution of matching point pairs, the transformation relationship between the visual coordinates is parameterized based on an affine transformation or Euclidean transformation model. By performing minimum mean square error fitting on multiple sets of high-confidence matching point pairs, the rigid transformation parameters, including the rotation matrix R and the translation vector T, are extracted, thereby constructing a three-dimensional rigid transformation matrix [R|T] that maps from the current image coordinate system to the visual map coordinate system. This matrix describes the overall pose relationship of the robot's current visual observation results relative to the visual map, taking into account changes in both direction and position. The visual coordinate offset is used as a local perturbation parameter of the current observation result and substituted into the constructed rigid transformation matrix to complete the inverse calculation of the robot's precise position in the global coordinate system of the map at the current moment. Depth estimation information or auxiliary odometry data of the visual coordinates is introduced into the calculation process to improve the three-dimensional spatial accuracy of the positioning result. To enhance the stability and robustness of positioning results, a Kalman filter or extended particle filter mechanism is introduced to dynamically smooth the position estimate obtained based on visual matching. This offsets offset errors caused by noise, occlusion, or illumination changes in single-frame matching, thereby obtaining a more accurate estimate of the robot's current position. This position information can ultimately be used for subsequent tasks such as navigation path planning, map updates, or multi-sensor fusion positioning.
[0076] Based on the above embodiments, Figure 6 This is a schematic diagram of the structure of a positioning device based on grid map and feature matching provided in an embodiment of the present application. Figure 6 The positioning device based on grid map and feature matching provided in this embodiment specifically includes: a grid unit recognition module 21, an environment feature extraction module 22, and a robot positioning module 23.
[0077] The grid unit identification module 21 is configured to obtain obstacle information corresponding to an obstacle detected by the robot, and identify the grid unit area where the obstacle is located in the pre-acquired obstacle grid map based on the obstacle information;
[0078] The environment feature extraction module 22 is configured to obtain the current environment visual image collected by the robot and extract the current environment visual features from the current environment visual image;
[0079] The robot positioning module 23 is configured to match the current environment visual features in the grid unit area of the pre-generated visual feature map using the current environment visual features to obtain a map matching result of the current environment visual features in the visual feature map, and determine the current position of the robot based on the map matching result.
[0080] Based on the above embodiment, the environmental feature extraction module 22 includes: a visual image cluster acquisition unit, configured to obtain an environmental visual image cluster collected by the robot in the same time window as the current environmental visual image; a dynamic obstacle recognition unit, configured to determine the dynamic area where the dynamic obstacles are located in the current environmental visual image based on the environmental visual field image cluster; an environmental visual feature extraction unit, configured to determine the current environmental static image based on the current environmental visual field image and the dynamic area, and extract the current environmental visual features from the current environmental static image.
[0081] Based on the above embodiment, the dynamic obstacle identification unit includes: an adjacent frame pixel residual map subunit, which is configured to calculate the environmental image frame difference between two temporally adjacent frames of environmental visual images in the environmental visual image cluster, and obtain a number of adjacent frame pixel residual maps based on the environmental image frame difference; and a dynamic area positioning subunit, which is configured to determine the dynamic area where the dynamic obstacle is located in the current environmental visual image based on the several adjacent frame pixel residual maps.
[0082] Based on the above embodiment, the dynamic area positioning subunit includes: a residual pixel area extraction component, which is configured to extract residual pixel areas whose corresponding pixel values exceed a preset residual threshold from each adjacent frame pixel residual map; a dynamic area positioning component, which is configured to determine the dynamic area where the dynamic obstacle is located in the current environment visual image based on each of the residual pixel areas.
[0083] Based on the above embodiment, the dynamic area positioning component includes: a residual area weight sub-component, which is configured to obtain the residual area weight of the residual pixel area based on the time interval between the two adjacent frames of environmental visual images corresponding to the residual pixel area and the current environmental visual image; and a dynamic area positioning sub-component, which is configured to determine the dynamic area where the dynamic obstacle is located in the current environmental visual image based on each of the residual pixel areas and its corresponding residual area weight.
[0084] Based on the above embodiment, the dynamic obstacle identification unit includes: a change area positioning subunit, configured to determine several change areas where dynamic obstacles are located in the current environmental view image based on the environmental view image cluster; and a dynamic area filling subunit, configured to fill pixels in several of the change areas to obtain the dynamic area where the dynamic obstacle is located.
[0085] Based on the above embodiment, the robot positioning module 23 includes: a visual coordinate offset unit, configured to extract the visual coordinate offset of the current environment visual feature in the visual feature map from the map matching result; a robot positioning unit, configured to construct a rigid transformation matrix between the robot and the visual feature map based on the map matching result, and calculate the current position of the robot according to the visual coordinate offset and the rigid transformation matrix.
[0086] As described above, the positioning device based on grid map and feature matching provided by the embodiment of the present application improves the accuracy and robustness of autonomous positioning of mobile robots in complex environments. By mapping the obstacle perception results to a preset grid map, and combining the stable visual features extracted from the current environmental visual image, a high-confidence feature matching operation is performed within the local grid range, which can effectively suppress the influence of dynamic interference on the positioning results and improve the ability to extract static structural information. At the same time, by constructing the spatial matching relationship and rigid transformation matrix of visual features, the robot can be accurately positioned in the visual feature map, overcoming the high dependence of traditional positioning systems on environmental staticity and sensor synchronization, and significantly enhancing the adaptability and real-time performance of the system in actual deployment. This device is suitable for a variety of application scenarios such as indoor and outdoor multi-scene navigation, target recognition and path planning in dynamic environments, and has strong engineering promotion value and practical prospects.
[0087] The positioning device based on grid map and feature matching provided in the embodiment of the present application can be used to execute the positioning method based on grid map and feature matching provided in the above embodiment, and has corresponding functions and beneficial effects.
[0088] Figure 7 This is a schematic diagram of a positioning device based on grid map and feature matching provided by an embodiment of the present application. Figure 7 The positioning device based on grid map and feature matching includes: a processor 31, a memory 32, a communication device 33, an input device 34, and an output device 35. The number of processors 31 in the positioning device based on grid map and feature matching can be one or more, and the number of memories 32 in the positioning device based on grid map and feature matching can be one or more. The processor 31, memory 32, communication device 33, input device 34, and output device 35 of the positioning device based on grid map and feature matching can be connected via a bus or other means.
[0089] The memory 32, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the positioning method based on grid map and feature matching in any embodiment of the present application (for example, the grid unit identification module 21, the environmental feature extraction module 22, and the robot positioning module 23 in the positioning device based on grid map and feature matching). The memory 32 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the device. In addition, the memory 32 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories may be connected to the device via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0090] The communication device 33 is used for data transmission.
[0091] The processor 31 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 32, that is, realizes the above-mentioned positioning method based on grid map and feature matching.
[0092] The input device 34 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 35 may include a display device such as a display screen.
[0093] The positioning device based on grid map and feature matching provided above can be used to execute the positioning method based on grid map and feature matching provided in the above embodiment, and has corresponding functions and beneficial effects.
[0094] An embodiment of the present application also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to execute a positioning method based on grid map and feature matching. The positioning method based on grid map and feature matching includes: obtaining obstacle information corresponding to an obstacle detected by the robot, and identifying the grid unit area where the obstacle is located in a pre-acquired obstacle grid map based on the obstacle information; obtaining a current environment visual image collected by the robot, and extracting current environment visual features from the current environment visual image; matching using the current environment visual features in the grid unit area of a pre-generated visual feature map to obtain a map matching result of the current environment visual features in the visual feature map, and determining the current position of the robot based on the map matching result.
[0095] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media, such as CD-ROMs, floppy disks, or tape drives; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. Storage media may also include other types of memory or combinations thereof. In addition, the storage medium may be located in the first computer system in which the program is executed, or may be located in a different second computer system that is connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). The storage medium may store program instructions (e.g., embodied as a computer program) that can be executed by one or more processors.
[0096] Of course, the storage medium containing computer-executable instructions provided in an embodiment of the present application, whose computer-executable instructions are not limited to the above-mentioned positioning method based on raster map and feature matching, can also execute related operations in the positioning method based on raster map and feature matching provided in any embodiment of the present application.
[0097] The positioning device based on raster map and feature matching, storage medium and positioning equipment based on raster map and feature matching provided in the above embodiments can execute the positioning method based on raster map and feature matching provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, please refer to the positioning method based on raster map and feature matching provided in any embodiment of the present application.
[0098] The above are only preferred embodiments of the present application and the technical principles employed. The present application is not limited to the specific embodiments described herein, and any obvious changes, readjustments, and substitutions that are apparent to those skilled in the art will not depart from the scope of protection of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present application. The scope of the present application is determined by the scope of the claims.
Claims
1. A positioning method based on grid map and feature matching, characterized in that: include: Obtaining obstacle information corresponding to an obstacle detected by the robot, and identifying a grid unit area where the obstacle is located in a pre-acquired obstacle grid map based on the obstacle information; Acquire a current environment visual image collected by the robot, and extract current environment visual features from the current environment visual image; In the grid unit area of the pre-generated visual feature map, the current environment visual features are used for matching to obtain a map matching result of the current environment visual features in the visual feature map, and the current position of the robot is determined based on the map matching result.
2. The positioning method based on grid map and feature matching according to claim 1, characterized in that: The extracting the current environment visual features from the current environment visual image includes: Acquire a cluster of environmental visual images collected by the robot and in the same time window as the current environmental visual image; Determining a dynamic area where a dynamic obstacle is located in the current environmental view image based on the environmental view image cluster; A current environment static image is determined according to the current environment visual field image and the dynamic area, and current environment visual features are extracted from the current environment static image.
3. The positioning method based on grid map and feature matching according to claim 2, characterized in that: The determining, based on the environmental view image cluster, a dynamic area where a dynamic obstacle is located in the current environmental view image includes: Calculating the environmental image frame difference between two temporally adjacent frames of environmental visual images in the environmental visual image cluster, and obtaining a plurality of adjacent frame pixel residual maps based on the environmental image frame difference; A dynamic area where a dynamic obstacle is located in the current environmental visual image is determined based on the pixel residual images of the plurality of adjacent frames.
4. The positioning method based on grid map and feature matching according to claim 3, characterized in that: The determining, based on the pixel residual images of the plurality of adjacent frames, of a dynamic area where a dynamic obstacle is located in the current environmental visual image comprises: Extracting residual pixel regions whose corresponding pixel values exceed a preset residual threshold from the pixel residual maps of each adjacent frame; A dynamic area where a dynamic obstacle is located in the current environmental visual image is determined based on each of the residual pixel areas.
5. The positioning method based on grid map and feature matching according to claim 4, characterized in that: The determining, based on each of the residual pixel regions, of a dynamic region where a dynamic obstacle is located in the current environmental visual image includes: Obtaining a residual region weight of the residual pixel region based on a time interval between two adjacent frames of environmental visual images corresponding to the residual pixel region and the current environmental visual image; Based on each of the residual pixel regions and its corresponding residual region weight, a dynamic region where a dynamic obstacle is located in the current environmental visual image is determined.
6. The positioning method based on grid map and feature matching according to claim 2, characterized in that: The determining, based on the environmental view image cluster, a dynamic area where a dynamic obstacle is located in the current environmental view image includes: Determining, based on the environmental view image cluster, a plurality of change regions where dynamic obstacles are located in the current environmental view image; Pixel filling is performed on a plurality of the changed areas to obtain a dynamic area where the dynamic obstacle is located.
7. The positioning method based on grid map and feature matching according to claim 1, characterized in that: Determining the current position of the robot based on the map matching result includes: Extracting the visual coordinate offset of the current environment visual feature in the visual feature map from the map matching result; A rigid transformation matrix between the robot and the visual feature map is constructed based on the map matching result, and a current position of the robot is calculated according to the visual coordinate offset and the rigid transformation matrix.
8. A positioning device based on grid map and feature matching, characterized in that: include: A grid unit identification module is used to obtain obstacle information corresponding to an obstacle detected by the robot, and identify the grid unit area where the obstacle is located in a pre-acquired obstacle grid map based on the obstacle information; An environmental feature extraction module is used to obtain a current environmental visual image collected by the robot and extract current environmental visual features from the current environmental visual image; The robot positioning module is used to match the current environment visual features in the grid unit area of the pre-generated visual feature map using the current environment visual features to obtain a map matching result of the current environment visual features in the visual feature map, and determine the current position of the robot based on the map matching result.
9. A positioning device based on grid map and feature matching, characterized in that: include: one or more processors; A memory storing one or more programs, when the one or more programs are executed by the one or more processors, enables the one or more processors to implement the positioning method based on grid map and feature matching as described in any one of claims 1-7.
10. A storage medium containing computer-executable instructions, characterized in that: When executed by a computer processor, the computer executable instructions are used to perform the positioning method based on grid map and feature matching as described in any one of claims 1 to 7.