Multi-camera mapping method and device based on spherical PnP

By constructing a spherical PnP-based multi-camera mapping method, a spherical coordinate system is built, and ORB features are corrected using a neural network. Kalman filtering is combined to optimize pose, which solves the problems of computational complexity and feature matching robustness of traditional multi-camera SLAM in complex scenes, and achieves efficient multi-camera mapping.

CN121120363APending Publication Date: 2025-12-12STATE GRID JIANGSU ELECTRIC POWER CO XUZHOU POWER SUPPLY CO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511219850.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Traditional multi-camera SLAM methods suffer from high computational complexity in complex scenes and insufficient robustness in feature matching under dynamic environments, especially in panoramic images where the accuracy of feature point matching decreases. Existing PnP methods also lack adaptability to spherical images.

Method used

A multi-camera mapping method based on spherical PnP is adopted. By constructing a spherical coordinate system, ORB features are extracted and feature correction is performed through a neural network model. Combined with the Kalman filter pose prediction mechanism, image keyframes are optimized to achieve multi-camera mapping.

Benefits of technology

It improves the stability and robustness of cross-camera feature matching, realizes unified optimization of multi-camera observation data, and improves the accuracy and real-time performance of multi-camera mapping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120363A_ABST
    Figure CN121120363A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-camera mapping method and device based on spherical PnP, and the method specifically comprises the steps: constructing a spherical coordinate system, and giving the spherical coordinates of all cameras; oRB features of an original image are extracted, feature correction is carried out through a neural network model, and image features are given; transforming the image features, and determining direction vectors of the image features; a Kalman filtering pose prediction mechanism and the historical state of each camera are fused, and the initial pose of each camera is given; taking a minimum direction vector included angle error as an optimization target, performing PnP analysis of a spherical feature space, optimizing the initial pose, and determining the pose of each camera; and repositioning the image key frame by global optimization constraint, and realizing multi-camera mapping based on spherical PnP in combination with an original image. On the aspect of improving the stability of cross-camera feature matching, multi-camera observation data is integrated into a unified optimization framework, the camera poses and map points are jointly optimized, and good real-time performance and robustness are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer vision, and particularly relates to a multi-camera mapping method and device based on spherical PnP. BACKGROUND

[0002] Traditional visual SLAM (Visual Simultaneous Localization and Mapping) relies on inter-frame feature matching and geometric constraints, but the small error of each step of pose estimation will accumulate over time, resulting in map drift; monocular SLAM is more difficult to directly measure depth, and needs to estimate scale through triangulation, and the triangulation error will increase with the increase of feature point depth. The core goal of a multi-camera SLAM system is to establish geometric constraints between multiple camera coordinate systems and unify them to a global reference coordinate system for joint optimization.

[0003] Existing multi-view SLAM methods face significant challenges in complex scenes: for example, MultiCol-SLAM, as a typical method of multi-view collaborative SLAM, integrates multi-camera observations by constructing a multi-keyframe (MKF) structure, and has unique advantages. However, it still has certain limitations: including the optimization model complexity increases quadratically with the number of cameras, the computational complexity is too high, and the feature matching robustness in dynamic environments is insufficient, and the descriptor it relies on will greatly decrease the matching accuracy across the edge distortion region of the fisheye camera.

[0004] Therefore, a method for distributed multi-camera systems is given, namely Multicam-SLAM, which proposes a non-overlapping field of view SLAM framework that independently constructs sub-maps for each camera and aligns poses across sub-maps through pose graph optimization. As patent CN116309801A gives a camera pose estimation method, device, equipment and storage medium, the method includes: capturing an image of a real scene through a camera, obtaining image coordinate points corresponding to pre-calculated three-dimensional coordinate points in the real scene from the image; projecting the image coordinate points onto a standardized sphere of the camera to obtain first projection coordinates, and the center of the standardized sphere is the optical center of the camera; projecting the three-dimensional coordinate points onto the standardized sphere to obtain second projection coordinates; based on the three-dimensional coordinate points, the first projection coordinates and the second projection coordinates, calculating the camera pose. This scheme projects points onto the standardized sphere of the camera, and after projection, all points are the same distance from the optical center of the camera, and all points have the same contribution to the pose, making the optimization iteration of the pose more easily convergent, the calculation process more reasonable, reducing the occurrence of the algorithm falling into a local minimum value during iterative calculation of the camera pose, and improving the accuracy of calculating the pose.

[0005] But the limitation is that the feature matching in the submap depends on the re-projection error threshold, which may be large in the repeated area, and the effect is not as good as the unified feature coordinate system.

[0006] In the field of computer vision and robot navigation, the PnP problem (Perspective-n-Point) is a typical geometric optimization problem, which estimates the pose of the camera (i.e., the camera position and the camera attitude) through the known 3D space points and their 2D projections in the image.

[0007] The traditional PnP method has certain limitations, for example, the PnP method requires that the feature points be uniformly distributed on the image plane and have obvious depth difference, otherwise the degeneration problem is easy to occur; at the same time, the pinhole model used by the traditional PnP method is not suitable for panoramic images.

[0008] Therefore, around the core idea of multi-camera SLAM, in order to solve the shortcomings of the traditional method and the existing multi-camera SLAM, how to make the multi-camera fusion SLAM based on the spherical PnP technology better adapt to the spherical image and better estimate the pose of the camera is a problem to be solved by those skilled in the art. SUMMARY

[0009] In view of the defects in the prior art, the present application provides a multi-camera mapping method and device based on spherical PnP, which specifically comprises the following steps: constructing a spherical coordinate system and giving the spherical coordinates of each camera; extracting the ORB features of the original images, correcting the features through a neural network model, and giving the image features; transforming the image features to determine the direction vectors of the image features; fusing the Kalman filter pose prediction mechanism and the historical state of each camera to give the preliminary pose of each camera; determining the image key frame based on the preliminary pose of each camera, taking the minimum angle error of the direction vectors as the optimization objective, performing PnP analysis on the spherical feature space, optimizing the preliminary pose, and determining the pose of each camera; repositioning the image key frame with global optimization constraints, and combining the original images of each camera to realize multi-camera mapping based on spherical PnP.

[0010] The present application improves the stability of cross-camera feature matching, integrates multi-camera observation data into a unified optimization framework, and jointly optimizes the camera pose and the map point, which has good real-time performance and robustness.

[0011] In the first aspect, the present application provides a multi-camera mapping method based on spherical PnP, which specifically comprises the following steps: constructing a spherical coordinate system and giving the spherical coordinates of each camera; acquiring the original images of each camera, extracting the ORB features, correcting the features through a neural network model, and giving the image features; Transform the image features according to the spherical coordinates of the cameras to determine the direction vectors of the image features; According to the direction vectors of the image features, a Kalman filter pose prediction mechanism and historical states of the cameras are fused to give preliminary poses of the cameras; Based on the preliminary poses of the cameras, image key frames are determined, a PnP analysis in a spherical feature space is performed to optimize the preliminary poses and determine the poses of the cameras, with minimizing the angle error of the direction vectors as an optimization objective; The image key frames are repositioned with global optimization constraints, and the original images of the cameras are combined to realize multi-camera mapping based on spherical PnP.

[0012] Further, the multiple cameras form a rigid vision system, the number of cameras is 4-6, adjacent cameras have a common viewing area and the center positions of the cameras are different, and the relative poses between the cameras are fixed; A spherical coordinate system is constructed to give the spherical coordinates corresponding to the cameras, specifically including the following steps: Three-dimensional space points of each camera are obtained; The three-dimensional space points are projected to a virtual unit sphere in a linear manner to form unit sphere points; Based on the common viewing area of adjacent cameras, the relative poses between the cameras are given; Taking the center of any camera as the center of the spherical coordinate system, the unit sphere points of each camera are calibrated by fusing the relative poses between the cameras to give the spherical coordinates corresponding to each camera.

[0013] Further, the original images of the cameras are collected, ORB features are extracted, and the features are corrected through a neural network model to give image features, specifically including the following steps: The original images collected by the cameras are obtained, Gaussian blur and downsampling are fused, and an image pyramid is constructed according to a preset number of layers; Based on a local neighborhood corner detection method, a FAST detector is used to extract feature points from each layer of the image pyramid, and non-maximum suppression is combined to give candidate corners; According to the gradient information in the neighborhood of each candidate corner, the candidate corners are directionally labeled to generate labeled candidate corners; The labeled candidate corners are processed based on a pre-constructed neural network model, the labeled candidate corners and the local image blocks corresponding to the neighborhoods are taken as inputs for feature extraction to determine high-dimensional feature descriptors, and the image features are given.

[0014] Further, the construction of the neural network model includes the following steps: The initial neural network model is built, wherein the initial neural network model comprises an encoder, a decoder, a first output head and a second output head, the encoder comprises at least two convolutional layers, the convolutional layers of the encoder are arranged in a standard convolutional layer and a depth separable convolutional layer alternately combined manner, the decoder combines deconvolution upsampling and bilinear interpolation processing, the first output head outputs a confidence degree, and the second output head outputs a high-bit feature descriptor; A joint loss function of the neural network model is given through a binary cross-entropy loss for the first output head and a contrast loss for the second output head. The initial neural network model is trained and iterated based on the joint loss function until convergence, and the construction of the neural network model is completed.

[0015] Further, based on the direction vector of the image feature, a Kalman filter pose prediction mechanism and historical states of each camera are fused to give a preliminary pose of each camera, and the preliminary pose of each camera comprises the following steps: The state vector is determined by the image feature; The state vector of the previous frame is obtained, and the pose prior of the current frame is calculated by combining the state transition relationship; The image feature of the current frame is extracted to obtain the observed pose of the current frame; The preliminary pose and the corresponding covariance are generated by fusing the pose prior and the observed pose.

[0016] Further, the preliminary pose and the corresponding covariance are generated by fusing the pose prior and the observed pose, and the preliminary pose and the corresponding covariance comprise the following steps: The residual of the pose prior and the observed pose is given; The pose prior is updated by calculating the Jacobian matrix and the Kalman gain to generate the preliminary pose and the corresponding covariance.

[0017] Further, based on the preliminary pose of each camera, the image key frame is determined, the angle error between the direction vector of the image feature and the projection direction based on the current preliminary pose is minimized as an optimization target, the PnP analysis of the spherical feature space is performed, the preliminary pose is optimized, and the pose of each camera is determined, and the pose of each camera comprises the following steps: Based on the preliminary pose and the corresponding covariance, the image key frame of each camera is given; The angle error between the direction vector of the image feature and the projection direction based on the current preliminary pose is minimized to obtain an optimization target; Based on the optimization target, the extrinsic parameters of each camera are updated and the preliminary pose and the corresponding covariance are iterated by solving the nonlinear least squares; Until convergence, the pose of each camera is determined.

[0018] Further, the image key frame is repositioned by a global optimization constraint, and the multi-camera mapping based on the spherical PnP is realized by combining the original images of each camera, and the multi-camera mapping based on the spherical PnP is specifically represented as: Visual word sequence transformation is performed on the image features of keyframes to generate bag-of-words vectors; Based on the bag-of-words vectors of image keyframes, search for historical frames that match the image keyframes to determine the set of candidate frames for loop closure. Constraints are applied to the closed-loop candidate frame set to form the target candidate frame; Optimize image keyframes using image features of target candidate frames; For the feature points optimized for keyframes of the image, multiple affine transformation blocks are fitted and formed, and the geometric deviations between each affine transformation block are given. By integrating the affine transformation blocks and the geometric deviations between them, a global optimization framework is formed. Combined with the intrinsic and extrinsic parameters of each camera, the original images of each camera for the corresponding key frames are adjusted. The feature point error vectors of the original images before and after adjustment are analyzed until the threshold requirement is met, thus realizing multi-camera mapping based on spherical PnP.

[0019] Furthermore, the fitting process forms multiple affine transformation blocks, specifically including the following steps: The feature points of the image keyframes are divided to form multiple initial affine transformation blocks; Boundary points are extracted at the intersection of adjacent initial affine transformation blocks, and nonlinear boundary fitting is performed using cubic spline curves to form multiple affine transformation blocks.

[0020] Secondly, the present invention also provides a multi-camera mapping device based on spherical PnP, employing the multi-camera mapping method based on spherical PnP as described above, specifically including: The coordinate processing unit is used to construct a spherical coordinate system and provide the spherical coordinates of each camera. The image feature analysis unit is used to acquire raw images from each camera, extract ORB features, and perform feature correction through a neural network model to provide image features. It also transforms the image features by combining the spherical coordinates of each camera to determine the direction vector of the image features. The pose estimation unit is used to give the preliminary pose of each camera by fusing the Kalman filter pose prediction mechanism and the historical state of each camera based on the direction vector of the image features. Based on the preliminary pose of each camera, the image keyframe is determined. The optimization objective is to minimize the angle error between the direction vectors. PnP analysis in the spherical feature space is performed to optimize the preliminary pose and determine the pose of each camera. The mapping unit is used to relocate keyframes of the image with global optimization constraints and combine them with the original images from each camera to realize multi-camera mapping based on spherical PnP.

[0021] The present invention provides a multi-camera mapping method and apparatus based on spherical PnP, which has at least the following beneficial effects: (1) In terms of improving the stability of feature matching across cameras, this invention integrates multi-camera observation data into a unified optimization framework, jointly optimizing camera pose and map points, which has good real-time performance and robustness.

[0022] (2) ORB feature extraction is crucial for the initial pose estimation of each camera and the determination of image keyframes. Furthermore, obtaining more semantically meaningful and robust ORB features can improve the quality of subsequent PnP analysis in the spherical feature space. During ORB feature extraction, candidate corner points are generated using FAST and BRIEF, and then corrected using a pre-built neural network model from deep learning to extract more semantically meaningful and robust keypoints from the image.

[0023] (3) This alternating layer setting in the editor of the neural network model can ensure high feature extraction capability while also being applicable to resource-constrained devices. Setting the joint loss function can guide the neural network model to learn corner features with orientation labels and descriptors of local image patches, thereby improving the performance of feature extraction and matching.

[0024] (4) Relocating image keyframes using global optimization constraints and multi-camera mapping based on spherical PnP can improve the quality of image feature expression, enhance the grasp of global geometric structure, improve the fusion effect of multi-camera images, and ultimately achieve accurate multi-camera mapping, which is suitable for various application scenarios that require accurate mapping. Attached Figure Description

[0025] Figure 1 A flowchart illustrating a multi-camera mapping method based on spherical PnP provided by the present invention; Figure 2 A flowchart illustrating the spherical coordinates of each camera according to one embodiment of the present invention; Figure 3 A schematic diagram illustrating the positional relationship between a three-dimensional spatial point and a unit spherical point according to a certain embodiment of the present invention; Figure 4 A flowchart illustrating the image features provided in one embodiment of the present invention is shown. Figure 5 A schematic diagram of the architecture of a neural network model according to one embodiment of the present invention; Figure 6 A flowchart illustrating the initial pose of each camera is provided for one embodiment of the present invention. Figure 7 A schematic diagram illustrating the process of determining the pose of each camera according to one embodiment of the present invention; Figure 8 A schematic diagram illustrating the process of implementing multi-camera mapping based on spherical PnP according to one embodiment of the present invention; Figure 9 This is a schematic diagram of a multi-camera mapping device based on spherical PnP provided by the present invention. Detailed Implementation

[0026] To better understand the above technical solutions, a detailed description of the solutions will be provided below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0027] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms, and “multiple” generally includes at least two unless the context clearly indicates otherwise.

[0028] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that an article or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device that includes said element.

[0029] For a multi-camera SLAM system, camera pose is predicted using a Kalman filter mechanism within a unified spherical coordinate system. Then, PnP (Progressive Perspective) is solved based on the spherical coordinate system, enabling the multi-camera fused SLAM to better adapt to spherical reconstruction and estimate camera pose. Finally, the pose estimation results are combined to relocalize keyframes of the image, achieving multi-camera mapping.

[0030] Based on PnP solving in spherical coordinates, spatial point orientations can be directly modeled on the sphere, avoiding complex transformations and improving solution accuracy. Furthermore, determining image keyframes through pose estimation and achieving relocalization of the original image under global optimization constraints effectively compresses and corrects errors, improving the accuracy and structural consistency of multi-camera mapping.

[0031] like Figure 1 As shown, this invention provides a multi-camera mapping method based on spherical PnP, which specifically includes the following steps: Construct a spherical coordinate system and give the spherical coordinates of each camera; The system acquires raw images from each camera, extracts ORB features, and performs feature correction using a neural network model to provide image features. By combining the spherical coordinates of each camera, the image features are transformed to determine the direction vector of the image features; Based on the orientation vectors of image features, the Kalman filter pose prediction mechanism and the historical states of each camera are fused to give the preliminary pose of each camera. Image keyframes are determined based on the preliminary poses of each camera. The optimization objective is to minimize the angle error between the direction vectors. PnP analysis is performed in the spherical feature space to optimize the preliminary poses and determine the poses of each camera. Global optimization constraints are used to relocate keyframes of images, and combined with the original images from each camera, to achieve multi-camera mapping based on spherical PnP.

[0032] Among them, ORB (Oriented FAST and Rotated BRIEF) features belong to local image features, which are fused together for detection and description. The detection is based on FAST corner detection, and the description is the BRIEF binary descriptor.

[0033] This invention improves the stability of feature matching across cameras by integrating multi-camera observation data into a unified optimization framework, jointly optimizing camera pose and map points, and exhibiting good real-time performance and robustness.

[0034] In this invention, the camera type can be a fisheye camera, a pinhole camera, etc., and multiple cameras form a rigid vision system. This rigid vision system can realize the capture of panoramic images. The number of cameras in the rigid vision system can be 4-6. Adjacent cameras have a common viewing area and the center positions of each camera are different. The relative poses between each camera are fixed.

[0035] like Figure 2 As shown, a spherical coordinate system is constructed, and the spherical coordinates of each camera are given. The specific steps include the following: Obtain the 3D spatial points of each camera; Three-dimensional spatial points are projected linearly onto a virtual unit sphere to form unit sphere points; Based on the shared viewing area of ​​adjacent cameras, the relative poses between each camera are given; Using the center of any camera as the center of the spherical coordinate system, the relative poses of each camera are fused, and the unit spherical point of each camera is calibrated to give the corresponding spherical coordinates of each camera.

[0036] In one embodiment, taking a four-fisheye camera as an example, three-dimensional spatial points are linearly projected onto a virtual unit sphere. To facilitate subsequent optimization in the spherical coordinate system, the projection onto the image plane is not performed; instead, unit spherical points (i.e., spatial coordinate points represented by unit vectors) are used directly. For example... Figure 3 As shown, where p i Let P be a point in three-dimensional space for a certain camera. i For a unit spherical point.

[0037] Subsequently, since the relative poses of the four cameras remain fixed, the four cameras move as a whole to acquire the relative poses of each camera.

[0038] Because the center positions of the four cameras are inconsistent, the centers of the spherical coordinate systems obtained by the four cameras do not coincide, making it impossible to simply unify the obtained unit spherical points. Therefore, based on the relative poses of the cameras and utilizing the common field of view of the four fisheye cameras, the center of one camera is taken as the center of the entire rigid vision system. The centers of the other three cameras are then transformed to coincide with the center of the rigid vision system, resulting in a complete spherical coordinate system. This joint calibration of the four fisheye cameras solves the problem of inconsistent camera center positions. The final spherical coordinates of each camera are the unit direction vectors with the unified spherical center as the origin.

[0039] Rigid vision systems rely on certain geometric constraints to determine the relative poses of each camera, mapping the three-dimensional spatial points of all cameras to a unified spherical coordinate system, thus constructing a globally consistent spherical feature space.

[0040] Extracting ORB features is crucial for estimating the initial pose of each camera and determining keyframes in the image. Furthermore, obtaining more semantically meaningful and robust ORB features can improve the quality of subsequent PnP analysis in the spherical feature space.

[0041] like Figure 4 As shown, the process involves acquiring raw images from each camera, extracting ORB features, and then refining these features using a neural network model to provide the image features. The specific steps include: The system acquires raw images from each camera, fuses them with Gaussian blur and downsampling, and constructs an image pyramid based on a preset number of layers. Based on the local neighborhood corner detection method, the FAST detector is used to extract feature points for each layer of the image pyramid, and non-maximum suppression is combined to give candidate corner points; Based on the gradient information in the neighborhood of each candidate corner point, the candidate corner points are labeled with directions to generate labeled candidate corner points; Based on a pre-built neural network model, the candidate corner points are processed. The candidate corner points and their corresponding local image patches are used as input to extract features, determine high-dimensional feature descriptors, and give image features.

[0042] Typically, in ORB feature extraction, the keypoint detector used is FAST, and the descriptor is BRIEF. However, the FAST detector itself lacks scale invariance, and keypoints may be lost when the image is scaled. Therefore, to enhance its robustness, an image pyramid is constructed, and FAST feature point extraction is performed independently on each layer of the image. Since FAST also lacks rotation invariance, oFAST (oriented FAST) is introduced to calculate the principal orientation of keypoints while detecting them, performing orientation labeling to achieve robustness against rotation.

[0043] However, when applied to rigid vision systems, relying solely on the FAST detector presents new challenges. Panoramic images exhibit significant nonlinear distortions, particularly in the upper and lower polar regions where geometric distortion is more severe. Furthermore, the BRIEF descriptor relies on a fixed pixel-pair sampling template to generate binary feature vectors, which is based by default on the local Euclidean structure of the image. In images with severe distortion, this sampling method of BRIEF causes the descriptor to fail to maintain stable relative geometric relationships, resulting in a significant decrease in the accuracy of feature point matching.

[0044] Therefore, during ORB feature extraction, candidate corner points are generated by FAST and BRIEF, and then a pre-built neural network model in deep learning is used for correction to extract key points in the image that are more semantically meaningful and robust.

[0045] like Figure 5 As shown, the architecture of the neural network model includes an input layer, an editor, a decoder, and an output head. The output head includes a first output head and a second output head. The first output head outputs the confidence score, and the second output head outputs the high-order feature descriptor.

[0046] The editor includes multiple convolutional layers, including standard convolutional layers and depthwise separable convolutional layers. The editor uses an alternating pattern of standard and depthwise separable convolutional layers. Standard convolutional layers use multi-channel kernels to convolve the input. Depthwise separable convolutional layers consist of depthwise convolution and pointwise convolution. The depthwise convolution applies a kernel independently to each input channel without mixing with other channels. The pointwise convolution uses a 1×1 convolution to change the number of channels while maintaining the spatial dimension. After the input layer, standard convolutional layers are used to extract basic features, followed by depthwise separable convolution to reduce computation, and then standard convolution to expand the feature dimension. This pattern is repeated, and pooling is added. This alternating layer setup ensures high feature extraction capabilities while also allowing application on resource-constrained devices.

[0047] The decoder uses lightweight deconvolution to progressively upsample and restore image resolution, and uses bilinear interpolation to smooth the features. The first output head, also known as the confidence output head, can use a standard convolutional layer with a sigmoid activation function. The second output head, also known as the descriptor output head, can adopt a bottleneck structure design, first using 1×1 convolution to map the features to a high-dimensional space, and then using global average pooling to compress them to the descriptor dimension, outputting a high-dimensional feature descriptor.

[0048] The construction of a neural network model includes the following steps: An initial neural network model is constructed, which includes an encoder, a decoder, a first output head, and a second output head. The encoder includes at least two convolutional layers, and the combination of the convolutional layers of the encoder adopts an alternation of standard convolutional layers and depthwise separable convolutional layers. The decoder combines deconvolution upsampling and bilinear interpolation processing. The first output head outputs the confidence score, and the second output head outputs the high-bit feature descriptor. The joint loss function of the neural network model is given by using the binary cross-entropy loss for the first output head and the contrast loss for the second output head. The binary cross-entropy loss of the first output head is used to evaluate the accuracy of the neural network model in predicting the presence or absence of corner points, specifically expressed as follows: ; Among them, L CE Let y be the binary cross-entropy loss function, where N is the total number of labeled candidate corner points, and y is the cross-entropy loss function. i Let y be the true label of the i-th candidate corner point. i p can be 0 or 1 i The confidence level predicted by the neural network model. The value of y is redefined in the binary cross-entropy loss. iThis allows it to indicate not only whether a corner exists, but also whether the direction of the corner is as expected. Thus, this binary cross-entropy loss function can simultaneously measure the accuracy of predicting whether a corner exists and the rationality of its direction.

[0049] The contrast loss of the second output head is used to optimize the discriminativeness of descriptors, making descriptors of similar candidate corner points closer together and those of dissimilar points further apart, specifically expressed as: ; Among them, L contrast To compare the loss functions, M is the total number of (a,p,n) triples, a is the candidate corner point to be corrected, p is the similar candidate corner point related to a, n is the dissimilar candidate corner point unrelated to a, m is the margin hyperparameter, and f() is the descriptor function. θ is the Euclidean distance, γ is the weighting coefficient controlling the influence of directional differences, and θ is the distance between the two points. diff (a, n) represents the directional difference between a and n, and θ diff (a, p) represents the directional difference between a and p. In the contrastive loss function, exp(γ•θ) diff (a, n)), exp (γ·θ diff (a, p) are all penalty terms for directional differences. When calculating descriptor matching, directional information is considered at the same time to improve the distinguishability of feature points.

[0050] The joint loss function is achieved by weighting the binary cross-entropy loss and the contrastive loss, as shown below: ; Among them, L total Here, α is the weight of the binary cross-entropy loss, and β is the weight of the contrastive loss. The joint loss function can simultaneously improve the accuracy of corner detection and the discriminative power of descriptors.

[0051] The initial neural network model is trained iteratively based on the joint loss function until convergence, thus completing the construction of the neural network model.

[0052] Setting a joint loss function can guide the neural network model to learn corner features with orientation labels and descriptors of local image patches, thereby improving the performance of feature extraction and matching.

[0053] By combining the spherical coordinates of each camera, the image features are transformed to determine the direction vector of the image features. The specific steps include the following: Extract the pixel coordinates of image features and calculate the radial distance and azimuth angle from the image center; Based on the camera imaging model, the polar angle is obtained, where the polar angle is specifically expressed as: θ=r / f, where r is the radial distance from the image feature to the image center, and f is the focal length of the camera; The polar angle and azimuth angle are converted into three-dimensional direction vectors to obtain the direction vectors of the image features.

[0054] Based on the constructed unified spherical coordinate system and the extracted image features, the image features acquired by each camera are converted into direction vectors with the unified spherical coordinate system as the reference, thereby realizing a unified representation of image features in the spherical coordinate system, laying the foundation for the subsequent estimation of the initial pose of each camera and the determination of image keyframes.

[0055] like Figure 6 As shown, based on the direction vectors of image features, the Kalman filter pose prediction mechanism and the historical states of each camera are fused to give the preliminary pose of each camera. The specific steps include the following: The state vector is determined by image features, that is, the three-dimensional position and orientation of the camera are combined into a state vector; Obtain the state vector of the previous frame, combine it with the state transition relationship, and calculate the pose prior of the current frame. The state transition relationship can be constructed based on the uniform velocity model or the uniform acceleration model. Using the state vector of the previous frame and the state transition relationship, the pose prior of the current frame can be predicted. Image features are extracted from the current frame to obtain the observed pose of the current frame; By fusing prior pose and observed pose, a preliminary pose and its corresponding covariance are generated.

[0056] Specifically, the process of fusing prior pose information and observed pose to generate a preliminary pose and its corresponding covariance includes the following steps: Give the pose prior and the residual of the observed pose, where the residual is the difference between the pose prior and the observed pose. The pose prior is updated by calculating the Jacobian matrix and Kalman gain, generating the initial pose and corresponding covariance.

[0057] By calculating the residual between the observed pose and the prior pose, the consistency between the prior estimate and the observation can be measured. Using the Jacobian matrix and Kalman gain, the observation information can be effectively fused into the pose prior, thereby correcting the prior and generating a more accurate preliminary pose estimate and the corresponding covariance, which reflects the uncertainty of the estimate.

[0058] like Figure 7 As shown, image keyframes are determined based on the preliminary poses of each camera. Minimizing the angular error between direction vectors is used as the optimization objective. PnP analysis in the spherical feature space is performed to optimize the preliminary poses and determine the poses of each camera. The specific steps include: Based on the initial pose and the corresponding covariance, keyframes of the images from each camera are given. The optimization objective is obtained by minimizing the angular error between the orientation vector of the image features and the projection direction obtained based on the current preliminary pose. Based on the optimization objective, the extrinsic parameters of each camera are updated and the initial pose and corresponding covariance are iterated through nonlinear least squares solution; The positions of each camera are determined after convergence.

[0059] The direction vector of the image feature, given by the constructed unified spherical coordinate system, represents the ideal direction of the image feature with respect to the corresponding camera optical center. Under this projection, all points are at the same distance from the camera optical center, thus eliminating scale differences caused by camera distortion or position changes, ensuring that each point contributes equally to pose optimization. The projection direction is obtained based on the current preliminary pose. According to the currently estimated preliminary camera pose (including initial camera rotation and translation), the projection direction vector is obtained by transforming the 3D spatial points into the spherical coordinate system. It reflects the direction of the spatial point in the camera's field of view under the currently estimated preliminary camera pose. By minimizing the angular error between the direction vector of the image feature and the projection direction obtained based on the current preliminary pose, the camera pose can be optimized, making the projection direction closer to the ideal direction based on the unified spherical coordinate system, thereby improving the accuracy of pose estimation.

[0060] like Figure 8 As shown, global optimization constraints are used to relocalize keyframes of the image, and combined with the original images from each camera, multi-camera mapping based on spherical PnP is achieved, specifically as follows: Visual word sequence transformation is performed on the image features of the keyframes of the image to give a bag-of-words vector. The visual word sequence is obtained by sorting and encoding the image features according to their positions in the keyframes of the image. The bag-of-words vector is binary, with each visual word corresponding to a binary bit, indicating whether the visual word appears in the keyframe of the image. Based on the bag-of-words vectors of image keyframes, search for historical frames that match the image keyframes to determine the set of candidate frames for loop closure. Constraints are applied to the closed-loop candidate frame set to form the target candidate frame; Optimize image keyframes using image features of target candidate frames; For the optimized feature points of image keyframes, multiple affine transformation blocks are fitted and given the geometric deviation between each affine transformation block. Among them, multiple affine transformation blocks can effectively correct the local geometric distortion of feature points in image keyframes, improve the expression quality of feature points, and provide a more accurate data basis for subsequent global deviation verification. By integrating the various affine transformation blocks and the geometric deviations between them, a global optimization framework is formed. Combined with the intrinsic and extrinsic parameters of each camera, the original images of each camera in the corresponding keyframes are adjusted. Through the global optimization framework, the original images of each camera can be adjusted so that they can be seamlessly stitched together in the virtual scene.

[0061] Analyze the feature point error vectors of the original images before and after adjustment until a threshold requirement is met to achieve multi-camera mapping based on spherical PnP. Comparing the feature point error vectors of the images before and after adjustment can quantify the effect of the adjustment, ensure that the error is within an acceptable range, and also help verify the effectiveness of the entire optimization process.

[0062] Specifically, the process of fitting multiple affine transformation blocks includes the following steps: The feature points of the image keyframes are divided to form multiple initial affine transformation blocks; Boundary points are extracted at the intersection of adjacent initial affine transformation blocks, and nonlinear boundary fitting is performed using cubic spline curves to form multiple affine transformation blocks.

[0063] In one embodiment, such as a robot navigation scenario or an autonomous driving scenario, after triggering an image keyframe, the image features of the keyframe are transformed, providing an information basis for subsequent feature matching with historical frames and quickly identifying potential loop-closing candidate frames with high similarity. For the selected candidate frame set, further geometric verification is performed to narrow down the range of potential loop-closing candidate frames, forming target candidate frames. The image keyframes are optimized using the target candidate frames, and then relocated through global optimization constraints. Combined with the original images from each camera, multi-camera mapping based on spherical PnP is achieved.

[0064] In the above scenarios, affine transformation block division and nonlinear boundary fitting for the geometric deviations between affine transformation blocks can enhance image consistency. Specifically, dividing the image keyframe into multiple affine transformation blocks makes the target within each block more closely match the actual target. The geometric deviations between the various affine transformation blocks include differences in rotation, translation, and scaling. Using the affine transformation blocks and the geometric deviations between them as variables, and combining camera intrinsic and extrinsic parameters with scene geometric constraints, a global optimization objective function is constructed. By solving this global optimization objective function, the original images of each camera for the corresponding image keyframe are adjusted.

[0065] The above methods, which use global optimization constraints to relocate keyframes of images and perform multi-camera mapping based on spherical PnP, can improve the quality of image feature representation, enhance the understanding of global geometric structure, improve the fusion effect of multi-camera images, and ultimately achieve accurate multi-camera mapping, which is suitable for various application scenarios that require accurate mapping.

[0066] like Figure 9 As shown, the present invention also provides a multi-camera mapping device based on spherical PnP, which employs the multi-camera mapping method based on spherical PnP as described above, specifically including: The coordinate processing unit is used to construct a spherical coordinate system and provide the spherical coordinates of each camera. The image feature analysis unit is used to acquire raw images from each camera, extract ORB features, and perform feature correction through a neural network model to provide image features. It also transforms the image features by combining the spherical coordinates of each camera to determine the direction vector of the image features. The pose estimation unit is used to give the preliminary pose of each camera by fusing the Kalman filter pose prediction mechanism and the historical state of each camera based on the direction vector of the image features. Based on the preliminary pose of each camera, the image keyframe is determined. The optimization objective is to minimize the angle error between the direction vectors. PnP analysis in the spherical feature space is performed to optimize the preliminary pose and determine the pose of each camera. The mapping unit is used to relocate keyframes of the image with global optimization constraints and combine them with the original images from each camera to realize multi-camera mapping based on spherical PnP.

[0067] The present invention provides a multi-camera mapping method and apparatus based on spherical PnP, which has at least the following beneficial effects: (1) In terms of improving the stability of feature matching across cameras, this invention integrates multi-camera observation data into a unified optimization framework, jointly optimizing camera pose and map points, which has good real-time performance and robustness.

[0068] (2) ORB feature extraction is crucial for the initial pose estimation of each camera and the determination of image keyframes. Furthermore, obtaining more semantically meaningful and robust ORB features can improve the quality of subsequent PnP analysis in the spherical feature space. During ORB feature extraction, candidate corner points are generated using FAST and BRIEF, and then corrected using a pre-built neural network model from deep learning to extract more semantically meaningful and robust keypoints from the image.

[0069] (3) This alternating layer setting in the editor of the neural network model can ensure high feature extraction capability while also being applicable to resource-constrained devices. Setting the joint loss function can guide the neural network model to learn corner features with orientation labels and descriptors of local image patches, thereby improving the performance of feature extraction and matching.

[0070] (4) Relocating image keyframes using global optimization constraints and multi-camera mapping based on spherical PnP can improve the quality of image feature expression, enhance the grasp of global geometric structure, improve the fusion effect of multi-camera images, and ultimately achieve accurate multi-camera mapping, which is suitable for various application scenarios that require accurate mapping.

[0071] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.

Claims

1. A multi-camera mapping method based on spherical PnP, characterized in that, Specifically, the steps include the following: Construct a spherical coordinate system and give the spherical coordinates of each camera; The system acquires raw images from each camera, extracts ORB features, and performs feature correction using a neural network model to provide image features. By combining the spherical coordinates of each camera, the image features are transformed to determine the direction vector of the image features; Based on the orientation vectors of image features, the Kalman filter pose prediction mechanism and the historical states of each camera are fused to give the preliminary pose of each camera. Image keyframes are determined based on the preliminary poses of each camera. The optimization objective is to minimize the angle error between the direction vectors. PnP analysis is performed in the spherical feature space to optimize the preliminary poses and determine the poses of each camera. Global optimization constraints are used to relocate keyframes of images, and combined with the original images from each camera, to achieve multi-camera mapping based on spherical PnP.

2. The multi-camera mapping method based on spherical PnP as described in claim 1, characterized in that, A rigid vision system is formed by multiple cameras, with 4-6 cameras. Adjacent cameras share a common field of view and each camera has a different center position. The relative poses of each camera are fixed. Construct a spherical coordinate system and give the spherical coordinates of each camera. The specific steps include the following: Obtain the 3D spatial points of each camera; Three-dimensional spatial points are projected linearly onto a virtual unit sphere to form unit sphere points; Based on the shared viewing area of ​​adjacent cameras, the relative poses between each camera are given; Using the center of any camera as the center of the spherical coordinate system, the relative poses of each camera are fused, and the unit spherical point of each camera is calibrated to give the corresponding spherical coordinates of each camera.

3. The multi-camera mapping method based on spherical PnP as described in claim 1, characterized in that, The process involves acquiring raw images from each camera, extracting ORB features, and then refining these features using a neural network model to provide the image features. The specific steps include: The system acquires raw images from each camera, fuses them with Gaussian blur and downsampling, and constructs an image pyramid based on a preset number of layers. Based on the local neighborhood corner detection method, the FAST detector is used to extract feature points for each layer of the image pyramid, and non-maximum suppression is combined to give candidate corner points; Based on the gradient information in the neighborhood of each candidate corner point, the candidate corner points are labeled with directions to generate labeled candidate corner points; Based on a pre-built neural network model, the candidate corner points are processed. The candidate corner points and their corresponding local image patches are used as input to extract features, determine high-dimensional feature descriptors, and give image features.

4. The multi-camera mapping method based on spherical PnP as described in claim 3, characterized in that, The construction of a neural network model includes the following steps: An initial neural network model is constructed, which includes an encoder, a decoder, a first output head, and a second output head. The encoder includes at least two convolutional layers, and the combination of the convolutional layers of the encoder adopts an alternation of standard convolutional layers and depthwise separable convolutional layers. The decoder combines deconvolution upsampling and bilinear interpolation processing. The first output head outputs the confidence score, and the second output head outputs the high-bit feature descriptor. The joint loss function of the neural network model is given by using the binary cross-entropy loss for the first output head and the contrast loss for the second output head. The initial neural network model is trained iteratively based on the joint loss function until convergence, thus completing the construction of the neural network model.

5. The multi-camera pose determination method based on spherical PnP as described in claim 1, characterized in that, Based on the orientation vectors of image features, the Kalman filter pose prediction mechanism and the historical states of each camera are fused to give the preliminary pose of each camera. The specific steps include the following: Determine the state vector using image features; Obtain the state vector of the previous frame, and calculate the pose prior of the current frame by combining the state transition relationship; Image features are extracted from the current frame to obtain the observed pose of the current frame; By fusing prior pose and observed pose, a preliminary pose and its corresponding covariance are generated.

6. The multi-camera mapping method based on spherical PnP as described in claim 5, characterized in that, By fusing prior pose information and observed pose, a preliminary pose and its corresponding covariance are generated. This process includes the following steps: Give the pose prior and the residual of the observed pose; The pose prior is updated by calculating the Jacobian matrix and Kalman gain, generating the initial pose and corresponding covariance.

7. The multi-camera mapping method based on spherical PnP as described in claim 1, characterized in that, Based on the preliminary poses of each camera, keyframes of the image are determined. Minimizing the angular error between direction vectors is used as the optimization objective. PnP analysis is performed in the spherical feature space to optimize the preliminary poses and determine the poses of each camera. The specific steps include: Based on the initial pose and the corresponding covariance, keyframes of the images from each camera are given. The optimization objective is obtained by minimizing the angular error between the orientation vector of the image features and the projection direction obtained based on the current preliminary pose. Based on the optimization objective, the extrinsic parameters of each camera are updated and the initial pose and corresponding covariance are iterated through nonlinear least squares solution; The positions of each camera are determined after convergence.

8. The multi-camera mapping method based on spherical PnP as described in claim 1, characterized in that, Global optimization constraints are used to relocalize keyframes of the image, and combined with the original images from each camera, multi-camera mapping based on spherical PnP is achieved, specifically as follows: Visual word sequence transformation is performed on the image features of keyframes to generate bag-of-words vectors; Based on the bag-of-words vectors of image keyframes, search for historical frames that match the image keyframes to determine the set of candidate frames for loop closure. Constraints are applied to the closed-loop candidate frame set to form the target candidate frame; Optimize image keyframes using image features of target candidate frames; For the feature points optimized for keyframes of the image, multiple affine transformation blocks are fitted and formed, and the geometric deviations between each affine transformation block are given. By integrating the affine transformation blocks and the geometric deviations between them, a global optimization framework is formed. Combined with the intrinsic and extrinsic parameters of each camera, the original images of each camera for the corresponding key frames are adjusted. The feature point error vectors of the original images before and after adjustment are analyzed until the threshold requirement is met, thus realizing multi-camera mapping based on spherical PnP.

9. The multi-camera mapping method based on spherical PnP as described in claim 8, characterized in that, The process of fitting multiple affine transformation blocks includes the following steps: The feature points of the image keyframes are divided to form multiple initial affine transformation blocks; Boundary points are extracted at the intersection of adjacent initial affine transformation blocks, and nonlinear boundary fitting is performed using cubic spline curves to form multiple affine transformation blocks.

10. A multi-camera mapping device based on spherical PnP, employing the multi-camera mapping method based on spherical PnP as described in any one of claims 1 to 9, characterized in that, Specifically, it includes: The coordinate processing unit is used to construct a spherical coordinate system and provide the spherical coordinates of each camera. The image feature analysis unit is used to acquire raw images from each camera, extract ORB features, and perform feature correction through a neural network model to provide image features. It also transforms the image features by combining the spherical coordinates of each camera to determine the direction vector of the image features. The pose estimation unit is used to give the preliminary pose of each camera by fusing the Kalman filter pose prediction mechanism and the historical state of each camera based on the direction vector of the image features. Based on the preliminary pose of each camera, the image keyframe is determined. The optimization objective is to minimize the angle error between the direction vectors. PnP analysis in the spherical feature space is performed to optimize the preliminary pose and determine the pose of each camera. The mapping unit is used to relocate keyframes of the image with global optimization constraints and combine them with the original images from each camera to realize multi-camera mapping based on spherical PnP.

Citation Information

Patent Citations

  • Camera pose estimation method and device, equipment and storage medium

    CN116309801A