Visual positioning method, device and equipment based on panoramic camera and storage medium

By combining dual fisheye cameras and advanced algorithms, the problems of image distortion and depth information acquisition in the panoramic camera visual positioning system are solved, achieving more comprehensive environmental information acquisition and high-precision pose estimation.

CN120655704APending Publication Date: 2025-09-16SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410296066.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In traditional visual positioning systems, panoramic cameras cannot be truly panoramic cameras when their images are severely distorted and their field of view is less than 180°. This leads to interference in feature matching and difficulty in obtaining depth information, affecting positioning accuracy and robustness.

Method used

A dual fisheye camera is used to acquire panoramic images. Feature points are extracted using the SPHORB algorithm and matched using an optimized feature matching algorithm. The motion estimation algorithm and the bundle adjustment optimization algorithm are combined to calculate the pose transformation matrix. The spherical coordinate system is introduced for modeling, and the self-supervised depth estimation algorithm is used to obtain depth information.

Benefits of technology

It expands the camera's visual range, improves the accuracy and robustness of visual positioning, prevents the problem of field of view reduction caused by distortion correction, enhances the stability and repeatability of feature points under different viewing angles and lighting changes, and improves the accuracy of pose estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655704A_ABST
    Figure CN120655704A_ABST
Patent Text Reader

Abstract

The invention relates to a visual positioning method and device based on a panoramic camera, equipment and a storage medium. The method comprises the following steps: acquiring a panoramic image through a double-fisheye camera; performing SPHORB feature point extraction on the panoramic image, and matching feature points between two continuous frames of panoramic images by using an optimized feature matching algorithm to obtain matched feature point pairs; based on the feature point pairs, calculating relative motion of the double-fisheye camera between two continuous frames of panoramic images by using a motion estimation algorithm to obtain a pose transformation matrix of the double-fisheye camera; and optimizing the pose transformation matrix by using a bundle adjustment optimization algorithm to obtain an optimized pose transformation matrix. According to the embodiment of the invention, the double-fisheye camera is used for shooting the panoramic image, the visual range of the camera is expanded, more comprehensive environment information is obtained, feature point extraction is carried out on the panoramic image, and the visual positioning precision of equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of visual positioning technology, and in particular relates to a visual positioning method, device, equipment and storage medium based on a panoramic camera. Background Art

[0002] Visual positioning technology is an important technology that enables the system to obtain texture information in all directions and achieve accurate positioning in complex environments. This technology plays a very critical role in autonomous driving, robot navigation, and augmented and virtual reality systems.

[0003] Visual positioning typically estimates and optimizes camera pose by correlating features or optical flows between camera frames. Traditional visual positioning systems primarily consist of a camera image information receiving module, a front-end visual odometry module, and a back-end nonlinear optimization module. The camera image information receiving module primarily uses the camera to capture a series of image data containing various texture information about the environment. The front-end visual odometry module extracts salient features from the image data and correlates the same features across different images to estimate the camera pose transformation matrix between the two frames. The back-end nonlinear optimization module further globally optimizes the series of pose transformation matrices to reduce errors and improve positioning accuracy. In traditional visual positioning systems, the most widely used cameras are typically monocular, binocular, and depth cameras. Because traditional cameras have a horizontal field of view of only approximately 60° and a vertical field of view of only approximately 45°, they struggle to capture a global field of view and thus fail to fully perceive environmental information. In weakly textured indoor environments, the lack of features can even lead to positioning loss, significantly limiting the performance of visual positioning systems.

[0004] To address these issues, existing technologies employ panoramic cameras for visual positioning. However, existing panoramic camera-based visual positioning systems typically use fisheye lenses for panoramic capture, resulting in images with significant nonlinear perspective distortion and severe image distortion, which significantly interferes with feature matching between frames. Some positioning methods use distortion correction to address this distortion problem. However, correcting severe distortion is not only difficult but also inevitably results in a reduction in the field of view. When the field of view is less than 180°, it cannot constitute a truly panoramic camera. Furthermore, panoramic camera-based visual positioning also faces the problem of difficulty in acquiring depth information. Summary of the Invention

[0005] The present application provides a panoramic camera-based visual positioning method, apparatus, device, and storage medium, which aims to solve at least one of the above-mentioned technical problems in the prior art to a certain extent.

[0006] In order to solve the above problems, this application provides the following technical solutions:

[0007] A visual positioning method based on a panoramic camera, comprising:

[0008] Acquire panoramic images through dual fisheye cameras;

[0009] Extracting SPHORB feature points from the panoramic image using the SPHORB algorithm, and matching the feature points between two consecutive panoramic image frames using an optimized feature matching algorithm to obtain matched feature point pairs;

[0010] Based on the feature point pairs, a motion estimation algorithm is used to calculate the relative motion of the dual fisheye camera between two consecutive panoramic image frames to obtain a pose transformation matrix of the dual fisheye camera;

[0011] The pose transformation matrix is ​​optimized using a bundle adjustment optimization algorithm to obtain an optimized pose transformation matrix.

[0012] The technical solution adopted in the embodiment of the present application also includes: obtaining a panoramic image by using a dual fisheye camera, specifically:

[0013] Acquire two fisheye images using the dual fisheye camera;

[0014] Using the same calibration plate to perform internal parameter calibration on the dual fisheye camera;

[0015] Placing the dual fisheye cameras back to back in a calibration environment with distinct features, and performing extrinsic calibration on the dual fisheye cameras using a stereo calibration algorithm;

[0016] Calculating the relative position and orientation of the dual fisheye cameras according to the intrinsic parameter calibration results and the extrinsic parameter calibration results of the dual fisheye cameras, and transforming the two fisheye images into the same spherical coordinate system through coordinate transformation to obtain two fisheye images after coordinate transformation;

[0017] A self-supervised depth estimation algorithm is used to obtain depth information of the two fisheye images after the coordinate transformation, and the two fisheye images are combined according to the depth information to generate a panoramic image; wherein the depth information is represented by a spherical coordinate system, and the depth information includes longitude, latitude and distance from the center of the sphere.

[0018] The technical solution adopted in the embodiment of the present application also includes: using the optimized feature matching algorithm to match the feature points between two consecutive panoramic image frames, specifically:

[0019] A hexagonal grid approximation is performed on the panoramic image, and a spherical FAST detector and a spherical rBRIEF descriptor are constructed on the geodesic line. The rBRIEF descriptor is smoothed using a hexagonal Gaussian kernel. Feature point matching is performed by comparing the Hamming distance between the rBRIEF descriptors corresponding to two feature points of two consecutive panoramic images.

[0020] The technical solution adopted in the embodiment of the present application further includes: based on the feature point pairs, using a motion estimation algorithm to calculate the relative motion of the dual fisheye camera between two consecutive panoramic image frames to obtain a pose transformation matrix of the dual fisheye camera, specifically:

[0021] The motion of the dual fisheye camera is estimated using a direct linear transformation method of the PnP algorithm, the two-dimensional image coordinates and the three-dimensional coordinates of the map points are obtained according to the image information and depth information of the panoramic image, and the position transformation matrix of the dual fisheye camera is calculated according to the two-dimensional image coordinates and the three-dimensional coordinates of the map points.

[0022] The technical solution adopted in the embodiment of the present application further includes: after calculating the relative motion of the dual-fisheye camera between two consecutive panoramic image frames based on the feature point pairs using a motion estimation algorithm and obtaining the pose transformation matrix of the dual-fisheye camera, the following steps are further included:

[0023] A representative frame is selected from the panoramic image as a key frame, the key frame is placed in a sliding window, a new key frame is created according to a key frame selection strategy, and the sliding window is optimized according to the new key frame.

[0024] The technical solution adopted in the embodiment of the present application further includes: creating a new key frame according to the key frame selection strategy, and optimizing the sliding window according to the new key frame, specifically:

[0025] Determine whether the number of feature point pairs matched between the current frame and the previous key frame is less than a set threshold. If so, create a new key frame, put the new key frame into the sliding window, and remove the last key frame in the sliding window to optimize the sliding window.

[0026] The technical solution adopted in the embodiment of the present application also includes: optimizing the posture transformation matrix using the bundle adjustment optimization algorithm, specifically:

[0027] Based on the feature point pairs and the pose transformation matrix, the key frames in the optimized sliding window are subjected to bundle adjustment optimization to obtain an optimized pose transformation matrix; wherein, the bundle adjustment optimization is specifically as follows: a graph comprising vertices and edges is used to represent the SLAM problem, wherein each vertex represents a state variable, and each edge represents a relationship or constraint between state variables; a reprojection cost function is constructed, and the cost function is minimized by the Levenberg-Marquardt algorithm to find a set of optimal poses and map points so as to minimize the observed reprojection error.

[0028] Another technical solution adopted in the embodiment of the present application is: a visual positioning device based on a panoramic camera, comprising:

[0029] Image acquisition module: used to acquire panoramic images through dual fisheye cameras;

[0030] Feature point extraction module: used to extract SPHORB feature points from the panoramic image using the SPHORB algorithm, and match the feature points between two consecutive panoramic image frames using an optimized feature matching algorithm to obtain matched feature point pairs;

[0031] A pose estimation module is configured to calculate the relative motion of the dual fisheye camera between two consecutive panoramic images using a motion estimation algorithm based on the feature point pairs, and obtain a pose transformation matrix of the dual fisheye camera;

[0032] Posture optimization module: used to optimize the posture transformation matrix using the bundle adjustment optimization algorithm to obtain the optimized posture transformation matrix.

[0033] Another technical solution adopted by the embodiment of the present application is: a device, the device comprising a processor and a memory coupled to the processor, wherein:

[0034] The memory stores program instructions for implementing the panoramic camera-based visual positioning method;

[0035] The processor is configured to execute the program instructions stored in the memory to control a visual positioning method based on a panoramic camera.

[0036] Another technical solution adopted in the embodiment of the present application is: a storage medium storing program instructions executable by a processor, wherein the program instructions are used to execute the panoramic camera-based visual positioning method.

[0037] Compared with the prior art, the beneficial effects of the embodiments of the present application are as follows: the panoramic camera-based visual positioning method, device, equipment and storage medium of the embodiments of the present application use dual fisheye cameras for image capture, which expands the camera's visual range and obtains more comprehensive environmental information. By introducing a spherical coordinate system to model the pose estimation problem, the problem of reduced field of view and depth uncertainty caused by distortion correction is prevented. At the same time, a self-supervised depth estimation algorithm is used to obtain depth information, which optimizes the accuracy and robustness of visual positioning, further improves the pose estimation accuracy and robustness, and uses an advanced SPHORB algorithm to extract feature points from panoramic images, so that the feature points have high stability and repeatability under different viewing angles and lighting changes, thereby improving the visual positioning accuracy of the device. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flow chart of a visual positioning method based on a panoramic camera according to an embodiment of the present application;

[0039] Figure 2 is a schematic diagram of a panoramic image generation process according to an embodiment of the present application;

[0040] Figure 3 This is a schematic structural diagram of a panoramic camera-based visual positioning device according to an embodiment of the present application;

[0041] Figure 4 This is a schematic diagram of the device structure of an embodiment of the present application;

[0042] Figure 5 A schematic diagram of the structure of the storage medium of an embodiment of the present application. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0044] The terms "first," "second," and "third" in this application are used only for descriptive purposes and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of such features. In the description of this application, "multiple" means at least two, for example, two, three, etc., unless otherwise specifically defined. All directional indications in the embodiments of this application (such as up, down, left, right, front, back...) are only used to explain the relative positional relationship, movement, etc. between the components under a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications also change accordingly. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products, or devices.

[0045] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0046] Specifically, see Figure 1 , is a flow chart of a panoramic camera-based visual positioning method according to an embodiment of the present application. The panoramic camera-based visual positioning method according to an embodiment of the present application comprises the following steps:

[0047] S100: Acquires panoramic images through dual fisheye cameras;

[0048] In this step, two fisheye images containing various texture information in the environment are obtained by using dual fisheye cameras, and the two fisheye images are combined to generate a panoramic image, which expands the camera's visual range, obtains more comprehensive environmental information, and saves equipment costs.

[0049] For details, please refer to Figure 2 , is a schematic diagram of a panoramic image generation process according to an embodiment of the present application, wherein the generation process specifically includes:

[0050] S101: Acquire two fisheye images using a dual fisheye camera, and perform internal reference calibration on the dual fisheye camera using the same calibration board;

[0051] S102: Place the two fisheye cameras back to back in a calibration environment with distinct features, and perform extrinsic calibration using a stereo calibration algorithm; wherein the calibration environment can be a checkerboard calibration board with multiple different angles or a large environment with clear features and a fixed position.

[0052] S103: Calculating the relative position and orientation of the dual fisheye cameras based on the intrinsic parameter calibration results and the extrinsic parameter calibration results, and transforming the two fisheye images into the same spherical coordinate system through coordinate transformation to obtain two fisheye images after coordinate transformation;

[0053] S104: Use a self-supervised depth estimation algorithm to obtain the depth information of the two fisheye images after coordinate transformation, combine the two fisheye images according to the depth information, and generate a panoramic image; wherein, the self-supervised depth estimation algorithm uses the sfm-learner method, and uses the constraints of reprojection between consecutive frames of video on posture and depth to realize the network's learning of depth and posture; the depth information is represented by a spherical coordinate system, and the main parameters include longitude, latitude, and distance from the center of the sphere. The embodiment of the present application obtains the depth information of the panoramic image by using a self-supervised depth estimation algorithm, so that the fisheye camera can also obtain the depth information of the feature points for subsequent posture estimation calculations. The embodiment of the present application uses a spherical coordinate system instead of a plane coordinate system to represent the depth information, avoiding the problem of reduced field of view due to distortion correction, and improving the accuracy and stability of subsequent feature point extraction and matching.

[0054] S110: Using the SPHORB algorithm to extract SPHORB feature points from the panoramic image, and using an optimized feature matching algorithm to find the correspondence between the feature points of two consecutive panoramic image frames, to obtain matched feature point pairs;

[0055] In this step, SPHORB feature points are an improvement on the traditional ORB (Oriented FAST and Rotated BRIEF, a commonly used image feature) features. The embodiment of this application uses the advanced SPHORB algorithm to extract feature points from panoramic images, making the feature points highly stable and repeatable under different viewing angles and lighting changes, thereby improving the visual positioning accuracy of the device. The feature matching algorithm specifically involves performing a hexagonal grid approximation on the panoramic image, constructing a spherical FAST detector and a spherical rBRIEF descriptor on the geodesic line, smoothing the rBRIEF descriptor using a hexagonal Gaussian kernel, and matching feature points by comparing the Hamming distance between the rBRIEF descriptors corresponding to two feature points in two consecutive panoramic image frames.

[0056] S120: Based on the feature point pairs, a motion estimation algorithm is used to calculate the relative motion of the dual fisheye camera between two consecutive panoramic image frames to obtain a pose transformation matrix of the dual fisheye camera;

[0057] In this step, the direct linear transformation method of the PnP (Perspective-n-Point, a method for solving 3D to 2D point correspondence) algorithm is used to estimate the motion of the dual fisheye camera. The 2D image coordinates and 3D map point coordinates are obtained based on the image information and depth information of the panoramic image. The pose transformation matrix of the dual fisheye camera is calculated based on the 2D image coordinates and the 3D map point coordinates.

[0058] S130: Selecting representative frames from the panoramic image as key frames, placing the key frames in a sliding window, creating new key frames according to a key frame selection strategy, and optimizing the sliding window based on the new key frames;

[0059] In this step, the sliding window optimization method specifically determines whether the number of feature point pairs matched between the current frame and the previous keyframe is less than a set threshold. If so, a new keyframe is created and placed into the sliding window. The old keyframe at the end of the sliding window is removed to optimize the sliding window and reduce the amount of computation. For example, if the threshold is set to 10, if the number of feature point pairs matched between the current frame and the previous keyframe is less than 10, a new keyframe is created.

[0060] S140: performing bundle adjustment optimization on the key frames in the optimized sliding window based on the matched feature point pairs and the pose transformation matrix to obtain an optimized pose transformation matrix;

[0061] In this step, bundle adjustment optimization involves representing the SLAM (Simultaneous Localization and Mapping) problem using a graph consisting of vertices and edges. Each vertex represents a state variable, such as the robot's pose or the location of a feature point on a map, and each edge represents the relationship or constraint between the state variables, such as the observational relationship between the robot and a landmark. A reprojection cost function is constructed and minimized using the Levenberg-Marquardt algorithm to find the optimal set of poses and map points that minimizes the observed reprojection error.

[0062] Based on the above, the panoramic camera-based visual positioning method of the embodiment of the present application uses a dual fisheye camera for image capture, which expands the camera's visual range and obtains more comprehensive environmental information. By introducing a spherical coordinate system to model the pose estimation problem, the problem of reduced field of view and depth uncertainty caused by distortion correction is prevented. At the same time, a self-supervised depth estimation algorithm is used to obtain depth information, which optimizes the accuracy and robustness of visual positioning, further improves the pose estimation accuracy and robustness, and uses an advanced SPHORB algorithm to extract feature points from panoramic images, so that the feature points have high stability and repeatability under different viewing angles and lighting changes, thereby improving the visual positioning accuracy of the device.

[0063] See also Figure 3 , is a schematic diagram of the structure of a panoramic camera-based visual positioning device according to an embodiment of the present application. The panoramic camera-based visual positioning device 40 according to an embodiment of the present application includes:

[0064] Image acquisition module 41: used to acquire panoramic images through dual fisheye cameras;

[0065] Feature point extraction module 42: used to extract SPHORB feature points from the panoramic image using the SPHORB algorithm, and match the feature points between two consecutive panoramic image frames using an optimized feature matching algorithm to obtain matched feature point pairs;

[0066] A pose estimation module 43 is configured to calculate the relative motion of the dual fisheye camera between two consecutive panoramic image frames using a motion estimation algorithm based on the feature point pairs, and obtain a pose transformation matrix of the dual fisheye camera;

[0067] The posture optimization module 44 is used to optimize the posture transformation matrix using a bundle adjustment optimization algorithm to obtain an optimized posture transformation matrix.

[0068] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0069] The device provided in the embodiment of the present application can be applied in the aforementioned method embodiment. For details, please refer to the description of the aforementioned method embodiment, which will not be repeated here.

[0070] See also Figure 4 , is a schematic diagram of the device structure of an embodiment of the present application. The device 50 includes:

[0071] A memory 51 storing executable program instructions;

[0072] a processor 52 connected to the memory 51;

[0073] The processor 52 is used to call the executable program instructions stored in the memory 51 and perform the following steps: acquiring a panoramic image through a dual fisheye camera; using the SPHORB algorithm to extract SPHORB feature points from the panoramic image, and using an optimized feature matching algorithm to match the feature points between two consecutive panoramic image frames to obtain matched feature point pairs; based on the feature point pairs, using a motion estimation algorithm to calculate the relative motion of the dual fisheye camera between two consecutive panoramic image frames to obtain a pose transformation matrix of the dual fisheye camera; and using a bundle adjustment optimization algorithm to optimize the pose transformation matrix to obtain an optimized pose transformation matrix.

[0074] The processor 52 may also be referred to as a CPU (Central Processing Unit). The processor 52 may be an integrated circuit chip having signal processing capabilities. The processor 52 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The general-purpose processor may be a microprocessor or any conventional processor.

[0075] See also Figure 5, is a structural diagram of the storage medium of an embodiment of the present application. The storage medium of the embodiment of the present application stores program instructions 61 that can implement the following steps: acquiring a panoramic image through a dual fisheye camera; using the SPHORB algorithm to extract SPHORB feature points from the panoramic image, and using the optimized feature matching algorithm to match the feature points between two consecutive frames of panoramic images to obtain matched feature point pairs; based on the feature point pairs, using the motion estimation algorithm to calculate the relative motion of the dual fisheye camera between two consecutive frames of panoramic images to obtain the pose transformation matrix of the dual fisheye camera; using the bundle adjustment optimization algorithm to optimize the pose transformation matrix to obtain the optimized pose transformation matrix. Among them, the program instructions 61 can be stored in the above-mentioned storage medium in the form of a software product, including a number of instructions for enabling a device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the methods of each embodiment of the present application. The aforementioned storage media include: various media that can store program instructions, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks, or terminal devices such as computers, servers, mobile phones, and tablets. Among them, the server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0076] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0077] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. The above is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the content of the description and drawings of this application, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A visual positioning method based on a panoramic camera, characterized in that: include: Acquire panoramic images through dual fisheye cameras; Extracting SPHORB feature points from the panoramic image using the SPHORB algorithm, and matching the feature points between two consecutive panoramic image frames using an optimized feature matching algorithm to obtain matched feature point pairs; Based on the feature point pairs, a motion estimation algorithm is used to calculate the relative motion of the dual fisheye camera between two consecutive panoramic image frames to obtain a pose transformation matrix of the dual fisheye camera; The pose transformation matrix is ​​optimized using a bundle adjustment optimization algorithm to obtain an optimized pose transformation matrix.

2. The visual positioning method based on a panoramic camera according to claim 1, characterized in that: The panoramic image is obtained by using a dual fisheye camera, specifically: Acquire two fisheye images using the dual fisheye camera; Using the same calibration plate to perform internal parameter calibration on the dual fisheye camera; Placing the dual fisheye cameras back to back in a calibration environment with distinct features, and performing extrinsic calibration on the dual fisheye cameras using a stereo calibration algorithm; Calculating the relative position and orientation of the dual fisheye cameras according to the intrinsic parameter calibration results and the extrinsic parameter calibration results of the dual fisheye cameras, and transforming the two fisheye images into the same spherical coordinate system through coordinate transformation to obtain two fisheye images after coordinate transformation; A self-supervised depth estimation algorithm is used to obtain depth information of the two fisheye images after the coordinate transformation, and the two fisheye images are combined according to the depth information to generate a panoramic image; wherein the depth information is represented by a spherical coordinate system, and the depth information includes longitude, latitude and distance from the center of the sphere.

3. The visual positioning method based on a panoramic camera according to claim 2, characterized in that: The optimized feature matching algorithm is used to match the feature points between two consecutive panoramic images, specifically: A hexagonal grid approximation is performed on the panoramic image, and a spherical FAST detector and a spherical rBRIEF descriptor are constructed on the geodesic line. The rBRIEF descriptor is smoothed using a hexagonal Gaussian kernel. Feature point matching is performed by comparing the Hamming distance between the rBRIEF descriptors corresponding to two feature points of two consecutive panoramic images.

4. The visual positioning method based on a panoramic camera according to claim 3, characterized in that: Based on the feature point pairs, the relative motion of the dual fisheye camera between two consecutive panoramic images is calculated using a motion estimation algorithm to obtain the pose transformation matrix of the dual fisheye camera, specifically: The motion of the dual fisheye camera is estimated using a direct linear transformation method of the PnP algorithm, the two-dimensional image coordinates and the three-dimensional coordinates of the map points are obtained according to the image information and depth information of the panoramic image, and the position transformation matrix of the dual fisheye camera is calculated according to the two-dimensional image coordinates and the three-dimensional coordinates of the map points.

5. The panoramic camera-based visual positioning method according to any one of claims 1 to 4, characterized in that: After calculating the relative motion of the dual-fisheye camera between two consecutive panoramic image frames using a motion estimation algorithm based on the feature point pairs and obtaining the pose transformation matrix of the dual-fisheye camera, the method further includes: A representative frame is selected from the panoramic image as a key frame, the key frame is placed in a sliding window, a new key frame is created according to a key frame selection strategy, and the sliding window is optimized according to the new key frame.

6. The visual positioning method based on a panoramic camera according to claim 5, characterized in that: The step of creating a new key frame according to the key frame selection strategy and optimizing the sliding window according to the new key frame is as follows: Determine whether the number of feature point pairs matched between the current frame and the previous key frame is less than a set threshold. If so, create a new key frame, put the new key frame into the sliding window, and remove the last key frame in the sliding window to optimize the sliding window.

7. The visual positioning method based on a panoramic camera according to claim 6, characterized in that: The bundle adjustment optimization algorithm is used to optimize the posture transformation matrix, specifically: Based on the feature point pairs and the pose transformation matrix, the key frames in the optimized sliding window are subjected to bundle adjustment optimization to obtain an optimized pose transformation matrix; wherein, the bundle adjustment optimization is specifically as follows: a graph comprising vertices and edges is used to represent the SLAM problem, wherein each vertex represents a state variable, and each edge represents a relationship or constraint between state variables; a reprojection cost function is constructed, and the cost function is minimized by the Levenberg-Marquardt algorithm to find a set of optimal poses and map points so as to minimize the observed reprojection error.

8. A visual positioning device based on a panoramic camera, characterized in that: include: Image acquisition module: used to acquire panoramic images through dual fisheye cameras; Feature point extraction module: used to extract SPHORB feature points from the panoramic image using the SPHORB algorithm, and match the feature points between two consecutive panoramic image frames using an optimized feature matching algorithm to obtain matched feature point pairs; A pose estimation module is configured to calculate the relative motion of the dual fisheye camera between two consecutive panoramic images using a motion estimation algorithm based on the feature point pairs, and obtain a pose transformation matrix of the dual fisheye camera; Posture optimization module: used to optimize the posture transformation matrix using the bundle adjustment optimization algorithm to obtain the optimized posture transformation matrix.

9. A device, characterized in that The device includes a processor and a memory coupled to the processor, wherein: The memory stores program instructions for implementing the panoramic camera-based visual positioning method according to any one of claims 1 to 7; The processor is configured to execute the program instructions stored in the memory to control a visual positioning method based on a panoramic camera.

10. A storage medium, characterized in that: Program instructions executable by a processor are stored, and the program instructions are used to execute the visual positioning method based on a panoramic camera as described in any one of claims 1 to 7.