Map construction method, device and equipment

The small hole image of the virtual camera is generated through panoramic images, and a three-dimensional visual map is constructed, which solves the problem of poor signal of GPS or Beidou satellite navigation system in indoor environments, and realizes accurate positioning and efficient data acquisition of terminal equipment. It is suitable for indoor positioning in coal, electricity, petrochemical and other industries.

CN114187344BActive Publication Date: 2025-08-22GUANGDONG GAOHANG INTELLECTUAL PROPERTY OPERATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111348552.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-15
Publication Date
2025-08-22
Estimated Expiration
2041-11-15

AI Technical Summary

Technical Problem

In indoor environments, the signal of GPS or Beidou satellite navigation system is poor, resulting in the inability to accurately locate terminal equipment, especially in the energy industries such as coal, electricity, and petrochemicals that cannot meet the positioning needs.

Method used

Panoramic images are used to generate small hole images of the virtual camera. By determining the rotation matrix and external parameter matrix of the virtual camera, a three-dimensional visual map is constructed, and the global positioning of the terminal device is performed based on the three-dimensional visual map.

Benefits of technology

It realizes accurate positioning in an indoor environment, improves data collection efficiency and map construction robustness, and is suitable for indoor positioning in energy industries such as coal, electricity, and petrochemicals, ensuring personnel safety and efficient management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187344B_ABST
    Figure CN114187344B_ABST
Patent Text Reader

Abstract

The present application provides a map construction method, apparatus and equipment, which includes: obtaining a panoramic image of a target scene, generating a first pinhole image corresponding to a first virtual camera based on the panoramic image; determining a rotation matrix between a target pose and an initial pose of a second virtual camera, and determining an extrinsic parameter matrix between the target pose and the initial pose based on the rotation matrix; determining a second pinhole image corresponding to the second virtual camera based on the rotation matrix, selecting two-dimensional feature points corresponding to the actual position of the target scene from the first pinhole image and the second pinhole image, and determining a three-dimensional map point corresponding to the actual position based on the two-dimensional feature points and the extrinsic parameter matrix; and constructing a three-dimensional visual map of the target scene based on multiple three-dimensional map points of the target scene. Through the technical solution of the present application, a terminal device of the target scene can be globally positioned based on the three-dimensional visual map, and the terminal device can be accurately positioned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision, and in particular to a map construction method, apparatus, and device. Background Art

[0002] GPS (Global Positioning System) is a high-precision radio navigation and positioning system based on artificial Earth satellites. It provides accurate geographic location, vehicle speed, and precise time information anywhere in the world, including near-Earth space. The Beidou satellite navigation system, consisting of three segments: the space segment, the ground segment, and the user segment, provides users with high-precision, highly reliable positioning, navigation, and timing services around the world, 24 / 7, and possesses regional navigation, positioning, and timing capabilities.

[0003] Because terminal devices are equipped with GPS or BeiDou satellite navigation systems, they can be used to locate the terminal device when positioning is required. In outdoor environments, due to the relatively strong GPS or BeiDou signals, the GPS or BeiDou satellite navigation systems can be used to accurately locate the terminal device. However, indoors, due to the relatively poor GPS or BeiDou signals, the GPS or BeiDou satellite navigation systems cannot accurately locate the terminal device. For example, in energy industries such as coal, electricity, and petrochemicals, there is an increasing demand for positioning. These positioning needs are generally in indoor environments, where accurate positioning of the terminal device is difficult due to signal obstruction and other issues. Summary of the Invention

[0004] The present application provides a map construction method, the method comprising:

[0005] Acquire a panoramic image of the target scene, and generate a first pinhole image corresponding to a first virtual camera based on the panoramic image; wherein the position of the first virtual camera is the center position of a sphere in a spherical coordinate system, and the initial posture of the first virtual camera is any posture centered at the center position of the sphere;

[0006] Determining a rotation matrix between a target pose of a second virtual camera and the initial pose, and determining an extrinsic parameter matrix between the target pose and the initial pose based on the rotation matrix; wherein the position of the second virtual camera is the center position of a sphere in the viewing spherical coordinate system, and the target pose is obtained by rotating the initial pose around the coordinate axis of the viewing spherical coordinate system;

[0007] Determining a second pinhole image corresponding to a second virtual camera based on the rotation matrix, selecting two-dimensional feature points corresponding to an actual position of a target scene from the first pinhole image and the second pinhole image, and determining a three-dimensional map point corresponding to the actual position based on the two-dimensional feature points and the extrinsic parameter matrix;

[0008] A three-dimensional visual map of the target scene is constructed based on a plurality of three-dimensional map points of the target scene.

[0009] The present application provides a map construction device, the device comprising:

[0010] An acquisition module, used to acquire a panoramic image of a target scene;

[0011] a generating module configured to generate a first pinhole image corresponding to a first virtual camera based on the panoramic image; wherein the position of the first virtual camera is the center position of a sphere in a spherical coordinate system, and the initial posture of the first virtual camera is any posture centered at the center position of the sphere;

[0012] a determination module, configured to determine a rotation matrix between a target pose of a second virtual camera and the initial pose, and determine an extrinsic parameter matrix between the target pose and the initial pose based on the rotation matrix; wherein the position of the second virtual camera is the center position of a sphere in the viewing spherical coordinate system, and the target pose is obtained by rotating the initial pose around the coordinate axis of the viewing spherical coordinate system;

[0013] The generation module is further used to determine the second pinhole image corresponding to the second virtual camera based on the rotation matrix; the determination module is further used to select two-dimensional feature points corresponding to the actual position of the target scene from the first pinhole image and the second pinhole image, and determine the three-dimensional map points corresponding to the actual position based on the two-dimensional feature points and the extrinsic parameter matrix; and, construct a three-dimensional visual map of the target scene based on multiple three-dimensional map points of the target scene.

[0014] The present application provides a map construction device, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is configured to execute the machine-executable instructions to implement the map construction method of an embodiment of the present application.

[0015] It can be seen from the above technical solutions that in the embodiment of the present application, a three-dimensional visual map of the target scene can be constructed, and the terminal device of the target scene can be globally positioned based on the three-dimensional visual map, and the terminal device can be accurately positioned. The target scene can be an indoor environment, realizing a vision-based indoor positioning function, which can be applied in energy industries such as coal, electricity, and petrochemicals to realize indoor positioning of personnel (such as workers, patrol personnel, etc.), quickly obtain personnel location information, ensure personnel safety, and realize efficient personnel management. A panoramic image of the target scene can be used to construct a three-dimensional visual map. The panoramic image has a large field of view, which avoids repeated data collection of the target scene and improves data collection efficiency. When determining the three-dimensional map point, the pinhole image can be obtained by using the projection method of the virtual camera, and the posture constraints of the virtual camera can be added. The three-dimensional map point is determined by using the virtual camera binding optimization strategy to improve the robustness, efficiency and accuracy of the map. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flowchart of a map construction method in one embodiment of the present application;

[0017] Figure 2 is a schematic diagram of a scene reconstruction and positioning method based on panoramic images in this application;

[0018] Figure 3 is a schematic diagram of expanding a panoramic image into a pinhole image in one embodiment of the present application;

[0019] Figure 4A It is a schematic diagram of the latitude and longitude coordinates of the spherical image and the rectangular coordinates of the panoramic image;

[0020] Figure 4B is a schematic diagram of the rectangular coordinates of the first pinhole image and the latitude and longitude coordinates of the visual spherical image;

[0021] Figure 4C It is a schematic diagram between the coordinates of the pinhole image and the coordinates of the viewing sphere;

[0022] Figure 5 It is a structural diagram of a map construction device in one embodiment of the present application. DETAILED DESCRIPTION

[0023] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a," "the," and "the" used in this application and claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more associated listed items.

[0024] It should be understood that although the terms first, second, third, etc. may be used to describe various information in the embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" used may also be interpreted as "at the time of" or "when" or "in response to determining".

[0025] In the embodiment of the present application, a map construction method is proposed, which is used to construct a three-dimensional visual map of a target scene, and then use the three-dimensional visual map to perform global positioning of a terminal device in the target scene. That is, when the terminal device moves in the target scene, the three-dimensional visual map is used to perform global positioning of the terminal device. Figure 1 FIG. 1 is a flow chart of a map construction method, which may include:

[0026] Step 101: Acquire a panoramic image of a target scene and generate a first pinhole image corresponding to a first virtual camera based on the panoramic image. Exemplarily, the position of the first virtual camera is the center of a sphere in a spherical coordinate system, and the initial pose of the first virtual camera is any pose centered at the center of the sphere.

[0027] Exemplarily, generating a first pinhole image corresponding to the first virtual camera based on the panoramic image may include but is not limited to: generating a visual spherical image corresponding to the visual spherical coordinate system based on the panoramic image, and generating a first pinhole image corresponding to the first virtual camera based on the visual spherical image.

[0028] Generating the visual spherical image corresponding to the visual spherical coordinate system based on the panoramic image may include, but is not limited to: determining a mapping relationship between the longitude and latitude coordinates in the visual spherical image and the rectangular coordinates in the panoramic image based on the width and height of the panoramic image; for each longitude and latitude coordinate in the visual spherical image, determining the rectangular coordinate corresponding to the longitude and latitude coordinate from the panoramic image based on the mapping relationship, and determining the pixel value of the longitude and latitude coordinate based on the pixel value of the rectangular coordinate. On this basis, the visual spherical image may be generated based on the pixel value of each longitude and latitude coordinate in the visual spherical image.

[0029] Generating a first pinhole image corresponding to the first virtual camera based on the visual spherical image may include, but is not limited to: determining the center point coordinates of the first pinhole image based on the width and height of the first pinhole image, and determining a mapping relationship between rectangular coordinates in the first pinhole image and latitude and longitude coordinates in the visual spherical image based on the center point coordinates and a target distance, where the target distance is the distance between the center point of the first pinhole image and the center position of a sphere in the visual spherical coordinate system. For each rectangular coordinate in the first pinhole image, the latitude and longitude coordinates corresponding to the rectangular coordinate are determined from the visual spherical image based on the mapping relationship, and the pixel value of the rectangular coordinate is determined based on the pixel value of the latitude and longitude coordinate. On this basis, the first pinhole image can be generated based on the pixel value of each rectangular coordinate in the first pinhole image.

[0030] Step 102: determine the rotation matrix between the target posture of the second virtual camera and the initial posture of the first virtual camera, and determine the extrinsic parameter matrix between the target posture and the initial posture based on the rotation matrix; illustratively, the position of the second virtual camera can be the center position of the sphere of the viewing spherical coordinate system, and the target posture is obtained by rotating the initial posture around the coordinate axis of the viewing spherical coordinate system.

[0031] Exemplarily, determining the rotation matrix between the target pose of the second virtual camera and the initial pose of the first virtual camera may include, but is not limited to: determining a first rotation angle between the target pose and the initial pose in the direction of a first coordinate axis, and determining a first sub-rotation matrix in the direction of the first coordinate axis based on the first rotation angle; determining a second rotation angle between the target pose and the initial pose in the direction of a second coordinate axis, and determining a second sub-rotation matrix in the direction of the second coordinate axis based on the second rotation angle; determining a third rotation angle between the target pose and the initial pose in the direction of a third coordinate axis, and determining a third sub-rotation matrix in the direction of the third coordinate axis based on the third rotation angle. On this basis, determining the rotation matrix between the target pose and the initial pose based on the first sub-rotation matrix, the second sub-rotation matrix, and the third sub-rotation matrix.

[0032] In one possible implementation, determining an extrinsic parameter matrix between the target pose and the initial pose based on the rotation matrix may include, but is not limited to: determining a translation matrix between the first virtual camera and the second virtual camera, and determining the extrinsic parameter matrix based on the rotation matrix and the translation matrix.

[0033] Step 103: Determine a second pinhole image corresponding to the second virtual camera based on the rotation matrix, select two-dimensional feature points corresponding to the actual position of the target scene from the first pinhole image and the second pinhole image, and determine a three-dimensional map point corresponding to the actual position based on the two-dimensional feature points and the extrinsic parameter matrix.

[0034] In one possible implementation, determining the three-dimensional map point corresponding to the actual position based on the two-dimensional feature points and the extrinsic parameter matrix may include but is not limited to: determining a target loss value of a configured loss function; determining a projection function value between the coordinate system of the virtual camera and the coordinate system of the pinhole image based on the target loss value and the two-dimensional feature points corresponding to the actual position; and determining the three-dimensional map point corresponding to the actual position based on the extrinsic parameter matrix and the projection function value, that is, as a three-dimensional map point in a three-dimensional visual map.

[0035] Step 104: construct a three-dimensional visual map of the target scene based on the multiple three-dimensional map points of the target scene.

[0036] Exemplarily, the three-dimensional visual map may include but is not limited to: a sample global descriptor corresponding to a sample image, a three-dimensional map point corresponding to the sample image, and a sample local descriptor corresponding to the three-dimensional map point; wherein the sample image is a pinhole image selected from the first pinhole image and the second pinhole image.

[0037] In one possible implementation, after step 104, during the global positioning process of the terminal device, a target image of the terminal device in the target scene is obtained; based on the similarity between the target image and the multiple frames of sample images corresponding to the three-dimensional visual map, a candidate sample image is selected from the multiple frames of sample images; a plurality of feature points are obtained from the target image; for each feature point, a target three-dimensional map point corresponding to the feature point is determined from the three-dimensional map points corresponding to the candidate sample images; and based on the multiple feature points and the target three-dimensional map points corresponding to the multiple feature points, a global positioning pose in the three-dimensional visual map corresponding to the target image is determined.

[0038] The method of selecting a candidate sample image from the multiple sample images based on the similarity between the target image and the multiple sample images corresponding to the three-dimensional visual map may include, but is not limited to, determining a global descriptor to be tested corresponding to the target image, and determining the distance between the global descriptor to be tested and the sample global descriptor corresponding to each sample image frame corresponding to the three-dimensional visual map. The method of selecting a candidate sample image from the multiple sample images based on the distance between the global descriptor to be tested and each sample global descriptor is performed; wherein the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is a minimum distance; or the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold.

[0039] Among them, determining the target three-dimensional map point corresponding to the feature point from the three-dimensional map points corresponding to the candidate sample image may include but is not limited to: determining the local descriptor to be tested corresponding to the feature point, the local descriptor to be tested can be used to represent the feature vector of the image block where the feature point is located, and the image block can be located in the target image; determining the distance between the local descriptor to be tested and the sample local descriptor corresponding to each three-dimensional map point corresponding to the candidate sample image; on this basis, based on the distance between the local descriptor to be tested and each sample local descriptor, selecting the target three-dimensional map point from the three-dimensional map points corresponding to the candidate sample image; wherein the distance between the local descriptor to be tested and the sample local descriptor corresponding to the target three-dimensional map point can be a minimum distance, and the minimum distance is less than a distance threshold.

[0040] Determining the global descriptor to be tested corresponding to the target image may include, but is not limited to: determining a bag-of-words vector corresponding to the target image based on a trained dictionary model, and determining the bag-of-words vector as the global descriptor to be tested; or inputting the target image into a trained deep learning model to obtain a target vector corresponding to the target image, and determining the target vector as the global descriptor to be tested. Of course, the above are just two examples of determining the global descriptor to be tested, and there is no limitation on the method of determining the global descriptor to be tested.

[0041] It can be seen from the above technical solutions that in the embodiment of the present application, a three-dimensional visual map of the target scene can be constructed, and the terminal device of the target scene can be globally positioned based on the three-dimensional visual map, and the terminal device can be accurately positioned. The target scene can be an indoor environment, realizing a vision-based indoor positioning function, which can be applied in energy industries such as coal, electricity, and petrochemicals to realize indoor positioning of personnel (such as workers, patrol personnel, etc.), quickly obtain personnel location information, ensure personnel safety, and realize efficient personnel management. A panoramic image of the target scene can be used to construct a three-dimensional visual map. The panoramic image has a large field of view, which avoids repeated data collection of the target scene and improves data collection efficiency. When determining the three-dimensional map point, the pinhole image can be obtained by using the projection method of the virtual camera, and the posture constraints of the virtual camera can be added. The three-dimensional map point is determined by using the virtual camera binding optimization strategy to improve the robustness, efficiency and accuracy of the map.

[0042] The map construction method of the embodiment of the present application is described below with reference to specific embodiments.

[0043] Data sources used to implement positioning functions can include GPS, LiDAR, millimeter-wave radar, and visual sensors (such as cameras), and positioning functions can be implemented using these data sources. GPS is easily affected by satellite conditions, weather conditions, and data transmission conditions, and cannot be used indoors. LiDAR and millimeter-wave radar have the advantages of low computational complexity, providing depth information, and being unaffected by light. However, LiDAR and millimeter-wave radar provide sparse information, are expensive, and are not yet widely used. In comparison, although the visual information provided by visual sensors is affected by light and weather, visual sensors are low-cost, small in size, easy to install, and rich in content, and have great application prospects in map positioning.

[0044] In order to use visual sensors to achieve positioning functions, it is usually necessary to build a high-precision three-dimensional visual map. The three-dimensional visual map is a perception of prior knowledge of the environment. Global positioning can be performed based on the three-dimensional visual map to obtain the global positioning posture of the terminal device in the three-dimensional visual map, thereby realizing the positioning function.

[0045] When building a 3D visual map, a monocular camera is typically used to capture images of the target scene. Based on these images, a 3D reconstruction algorithm (Structure From Motion) is used to obtain the resulting 3D visual map. However, the monocular camera's narrow field of view results in insufficient scene coverage, requiring repeated capture of target scene images for reconstruction, resulting in relatively low mapping efficiency.

[0046] To address the above issues, the present invention proposes a method for scene reconstruction and positioning based on panoramic images. A panoramic camera can be used to capture panoramic images, which are then used for 3D reconstruction to produce a 3D visual map. 3D reconstruction is the process of converting a physical scene into a 3D point cloud using the SFM algorithm. Because panoramic cameras have a very large field of view, they can avoid repeated image acquisition of the target scene, improving data acquisition efficiency and resolving the issues associated with the narrow field of view of monocular cameras. The quality of the 3D visual map is significantly higher than that of a monocular camera, resulting in higher efficiency and improved accuracy and robustness.

[0047] See also Figure 2The figure shows a schematic diagram of a scene reconstruction and positioning method based on panoramic images, which can include an offline mapping process and an online positioning process. For the offline mapping process, a panoramic image of the target scene (i.e., the scene for which a three-dimensional visual map needs to be constructed) can be collected, and the panoramic image can be expanded into a pinhole image. Based on the pinhole image and the fixed constraint (i.e., the extrinsic parameter matrix), a multi-camera bound SFM reconstruction is performed to obtain a three-dimensional visual map, and the map information corresponding to the three-dimensional visual map is stored. For the online positioning process, a target image of the target scene can be collected, and the target image can be pose-solved based on the three-dimensional visual map of the target scene to obtain the global positioning pose in the three-dimensional visual map corresponding to the target image, thereby completing the positioning process.

[0048] 1. You can capture a panoramic image of the target scene and expand the panoramic image into at least two pinhole images, see Figure 3 FIG. 1 is a schematic diagram of expanding a panoramic image into a pinhole image. The process includes:

[0049] Step 301: Obtain a panoramic image of the target scene. For example, a panoramic video of the target scene may be obtained and converted into a multi-frame panoramic image. There is no limitation to this process.

[0050] Step 302: Generate a spherical image corresponding to a spherical coordinate system based on the panoramic image.

[0051] For example, assuming that the mapping relationship between the spherical image and the panoramic image is a longitude-latitude mapping, the longitude-latitude coordinates of the spherical image are The relationship with the rectangular coordinates (x, y) of the panoramic image is shown in formula (1):

[0052]

[0053] In formula (1), λ and They represent the longitude and latitude coordinates of the spherical image (i.e., the image in the spherical coordinate system), and x and y represent the horizontal and vertical coordinates of the panoramic image, respectively.

[0054] See also Figure 4A The figure shows the relationship between the longitude and latitude coordinates of the spherical image and the rectangular coordinates of the panoramic image. The left side is the schematic diagram of the spherical image, and the right side is the schematic diagram of the panoramic image. In the spherical image, the range of the longitude coordinate is λ∈[0,2π], and the range of the latitude coordinate is In the panoramic image, the aspect ratio of the panoramic image is W:H=2:1, and the center position of the panoramic image is the origin (0,0).

[0055] Continue to see Figure 4A As shown, the latitude and longitude coordinates in the spherical image and the rectangular coordinates (up ,v p ) can be seen in formula (2).

[0056]

[0057] From formula (2), it can be seen that based on the width W and height H of the panoramic image, the mapping relationship between the longitude and latitude coordinates in the visual spherical image and the rectangular coordinates in the panoramic image can be determined. Obviously, for each longitude and latitude coordinate in the visual spherical image, as Based on this mapping relationship, the longitude and latitude coordinates can be determined from the panoramic image. The corresponding rectangular coordinates (u p ,v p ), and based on the rectangular coordinates (u p ,v p ) pixel values ​​to determine the latitude and longitude coordinates The pixel value, that is, the rectangular coordinate (u p ,v p ) as the latitude and longitude coordinates On this basis, the pixel values ​​of each latitude and longitude coordinate in the visual spherical image can be combined into a visual spherical image, and thus a visual spherical image can be obtained.

[0058] In summary, a spherical image can be obtained under the spherical coordinate system. The spherical coordinate system can be a spherical coordinate system. There is no restriction on the spherical coordinate system. The image under the spherical coordinate system is the spherical image. The relationship between the spherical image and the panoramic image can be seen in Figure 4A shown.

[0059] Step 303: Generate a first pinhole image corresponding to the first virtual camera based on the spherical image. The spherical image is an image in a spherical coordinate system. The position of the first virtual camera is the center of the sphere in the spherical coordinate system. The initial posture of the first virtual camera is any posture centered at the center of the sphere.

[0060] For example, the viewing spherical image can be expanded into pinhole images under multiple viewpoints, and each pinhole image corresponds to a virtual camera. The virtual camera does not exist in the real scene, but is a virtual camera at the center position of the sphere in the viewing spherical coordinate system, that is, the position of the virtual camera coincides with the center position of the sphere in the viewing spherical coordinate system.

[0061] The viewing spherical coordinate system can be a three-dimensional coordinate system, i.e., it has an X-axis, a Y-axis, and a Z-axis. For each virtual camera, the pose of the virtual camera can be any pose centered on the center of the sphere. For example, the virtual camera corresponds to three poses: the pose in the first direction coincides with the X-axis of the viewing spherical coordinate system (i.e., a rotation of 0 degrees around the X-axis), the pose in the second direction coincides with the Y-axis of the viewing spherical coordinate system (i.e., a rotation of 0 degrees around the Y-axis), and the pose in the third direction coincides with the Z-axis of the viewing spherical coordinate system (i.e., a rotation of 0 degrees around the Z-axis). For another example, the pose in the first direction of the virtual camera is rotated 60 degrees around the X-axis, the pose in the second direction coincides with the Y-axis, and the pose in the third direction coincides with the Z-axis. For another example, the pose in the first direction of the virtual camera is rotated 120 degrees around the X-axis, the pose in the second direction is rotated 60 degrees around the Y-axis, and the pose in the third direction coincides with the Z-axis, and so on. In summary, the posture of the virtual camera can be obtained by rotating around the coordinate axes (such as the X-axis, Y-axis, and Z-axis) of the viewing spherical coordinate system, and there is no restriction on the rotation angle.

[0062] For example, a virtual camera in any posture can be called the first virtual camera, and the posture of the first virtual camera can be called the initial posture. Obviously, the position of the first virtual camera is the center position of the sphere in the visual spherical coordinate system, and the initial posture of the first virtual camera is any posture centered on the center position of the sphere.

[0063] In this embodiment, the first direction of the initial posture coincides with the X-axis of the viewing spherical coordinate system, the second direction of the initial posture coincides with the Y-axis of the viewing spherical coordinate system, and the third direction of the initial posture coincides with the Z-axis of the viewing spherical coordinate system. Of course, the above is only an example of the initial posture and there is no limitation to this.

[0064] See also Figure 4B As shown, it is a schematic diagram of the relationship between the rectangular coordinates of the first pinhole image (i.e., the pinhole image corresponding to the first virtual camera) and the latitude and longitude coordinates of the visual spherical image. The left side is a schematic diagram of the visual spherical image, and the right side is a schematic diagram of the first pinhole image. The visual spherical image can be an image of a unit sphere.

[0065] See also Figure 4B As shown, a pinhole camera is virtualized at the center of the sphere in the viewing spherical coordinate system, namely the first virtual camera, and the image corresponding to the first virtual camera is the first pinhole image under the viewing plane. Assuming that the width of the first pinhole image is w and the height of the first pinhole image is h, the coordinates of the center point of the first pinhole image are (u0,v0)=(w / 2,h / 2). Assuming that the rectangular coordinates of point Q on the first pinhole image are (u q ,v q ), the line connecting point Q and the center position O of the spherical coordinate system intersects the spherical surface at point Q s, the distance between the center position O of the spherical coordinate system and the center point of the first pinhole image is d. Based on this, when the rectangular coordinates of point Q on the first pinhole image are converted into three-dimensional rectangular coordinates with the unit sphere center as the origin, it can be seen from formula (3):

[0066]

[0067] from Figure 4B It can be seen that the three-dimensional coordinates of point Qs and point Q are in a certain ratio. This ratio is based on the similarity principle of triangles, and the relationship of this ratio can be seen in formula (4):

[0068]

[0069] Due to OQ s =1, Therefore, click Q s The coordinates of can be seen in formula (5):

[0070]

[0071] In practical applications, the coordinates of points on the visual sphere can also be expressed using longitude and latitude, as shown in formula (6):

[0072]

[0073] Combining formula (3), formula (5) and formula (6), we can get the relationship between the rectangular coordinates of the first pinhole image and the latitude and longitude coordinates of the spherical image. See formula (7) for an example of this relationship.

[0074]

[0075] In formula (7), λ and Respectively, they represent the longitude and latitude coordinates of the visual spherical image, u0 = w / 2, v0 = h / 2, w represents the width of the first pinhole image, and h represents the height of the first pinhole image. Both w and h are known values. d represents the distance between the center position O of the sphere in the visual spherical coordinate system and the center point of the first pinhole image. It is a known value and can be configured based on experience or calculated using an algorithm. There is no restriction on this. Based on the above formula (7), the rectangular coordinates on the first pinhole image and the longitude and latitude coordinates of the visual spherical image can be projected, that is, the conversion from the visual spherical image to the first pinhole image can be performed.

[0076] In summary, based on the width w and height h of the first pinhole image, the center point coordinates of the first pinhole image (u0, v0) = (w / 2, h / 2) can be determined. Based on the center point coordinates and the target distance d, the rectangular coordinates (u, v) in the first pinhole image and the longitude and latitude coordinates in the spherical image can be determined. The mapping relationship between them can be seen in formula (7). Obviously, for each rectangular coordinate in the first pinhole image, the longitude and latitude coordinates corresponding to the rectangular coordinate can be determined from the visual spherical image based on the mapping relationship, and the pixel value of the rectangular coordinate can be determined based on the pixel value of the longitude and latitude coordinates, that is, the pixel value of the longitude and latitude coordinates is used as the pixel value of the rectangular coordinate. On this basis, the pixel value of each rectangular coordinate in the first pinhole image can be combined into a first pinhole image, and thus a first pinhole image is obtained.

[0077] In one possible embodiment, after obtaining a panoramic image of the target scene, the panoramic image can be directly converted into a first pinhole image corresponding to the first virtual camera, rather than converting the panoramic image into a visual spherical image corresponding to a visual spherical coordinate system, and the first pinhole image can be generated based on the visual spherical image. For example, the relationship between the rectangular coordinates of the first pinhole image and the rectangular coordinates of the panoramic image can be obtained, as shown in formula (8). The derivation process of formula (8) can be seen in formulas (2), (3), (5), and (6). The meaning of each letter in formula (8) can be seen in the above formula and will not be repeated here.

[0078]

[0079] It can be seen from formula (8) that for each rectangular coordinate (u, v) in the first pinhole image, the rectangular coordinate (u) corresponding to the rectangular coordinate (u, v) can be determined from the panoramic image based on the mapping relationship. p ,v p ), and based on the rectangular coordinate (u p ,v p ) determines the pixel value of the rectangular coordinate (u, v), that is, the rectangular coordinate (u p ,v p ) as the pixel value of the rectangular coordinate (u, v). On this basis, the pixel value of each rectangular coordinate in the first pinhole image can be combined into a first pinhole image, thereby obtaining the first pinhole image.

[0080] Step 304: determine the rotation matrix between the target posture of the second virtual camera and the initial posture of the first virtual camera. Exemplarily, the position of the second virtual camera can be the center position of the sphere of the viewing spherical coordinate system, and the target posture is obtained by rotating the initial posture around the coordinate axis of the viewing spherical coordinate system.

[0081] For example, the spherical image can be expanded into pinhole images from multiple viewpoints, with each pinhole image corresponding to a virtual camera. The virtual camera is a virtual camera located at the center of the sphere in the spherical coordinate system, that is, the position of the virtual camera coincides with the center of the sphere in the spherical coordinate system. The pose of the virtual camera can be obtained by rotating around the coordinate axes of the spherical coordinate system (such as the X-axis, Y-axis, and Z-axis), and there is no limit on the rotation angle.

[0082] Given the initial pose of the first virtual camera, the target pose of the second virtual camera is obtained by rotating the initial pose around the coordinate axes of the viewing spherical coordinate system. For example, the first direction of the target pose is obtained by rotating the initial pose 60 degrees around the X-axis, the second direction of the target pose is obtained by rotating the initial pose 60 degrees around the Y-axis, and the third direction of the target pose is obtained by rotating the initial pose 0 degrees around the Z-axis. Assuming that the first direction of the initial pose coincides with the X-axis, the second direction of the initial pose coincides with the Y-axis, and the third direction of the initial pose coincides with the Z-axis, then the first direction of the target pose is rotated 60 degrees around the X-axis, the second direction of the target pose is rotated 60 degrees around the Y-axis, and the third direction of the target pose coincides with the Z-axis. For another example, the first direction of the target pose is obtained by rotating the initial pose 120 degrees around the X-axis, the second direction of the target pose is obtained by rotating the initial pose 90 degrees around the Y-axis, and the third direction of the target pose is obtained by rotating the initial pose 0 degrees around the Z-axis. This rotational relationship between the target pose and the initial pose is not restricted.

[0083] In summary, based on the target posture of the second virtual camera and the initial posture of the first virtual camera, the first rotation angle between the target posture and the initial posture in the direction of the first coordinate axis can be determined. For example, for the rotation angle of the initial posture around the X axis, the first rotation angle is recorded as A. X , the second rotation angle between the target posture and the initial posture in the direction of the second coordinate axis can be determined, that is, the rotation angle of the initial posture around the Y axis, the second rotation angle is recorded as A Y , the third rotation angle between the target posture and the initial posture in the direction of the third coordinate axis can be determined, that is, the rotation angle of the initial posture around the Z axis, the third rotation angle is recorded as A Z .

[0084] Based on the first rotation angle A X , the second rotation angle A Y and the third rotation angle A Z , the rotation matrix between the target pose of the second virtual camera and the initial pose of the first virtual camera can be determined. For example, based on the first rotation angle A X Determine the first sub-rotation matrix R of the first coordinate axis direction x , based on the second rotation angle A YDetermine the second sub-rotation matrix R of the second coordinate axis direction y , based on the third rotation angle A Z Determine the third sub-rotation matrix R of the third coordinate axis direction z Then, based on the first sub-rotation matrix R x , the second sub-rotation matrix R y and the third sub-rotation matrix R z The rotation matrix between the target pose and the initial pose can be determined.

[0085] Since the rotation around the Z axis does not conform to the shooting habit, that is, the rotation angle around the Z axis for the initial posture is usually 0 degrees, the third sub-rotation matrix R z Usually 1, for this third sub-rotation matrix R z Without restriction, based on the first sub-rotation matrix R x and the second sub-rotation matrix R y This is explained by taking the determination of the rotation matrix as an example.

[0086] Refer to formula (9), which is based on the first rotation angle A X Determine the first sub-rotation matrix R x , see formula (10), based on the second rotation angle A Y Determine the second sub-rotation matrix R y , based on the first sub-rotation matrix R x and the second sub-rotation matrix R y The rotation matrix R can be obtained, as shown in formula (11).

[0087]

[0088]

[0089] R=R y R x Formula (11)

[0090] In summary, based on the first rotation angle A between the target posture of the second virtual camera and the initial posture of the first virtual camera X and the second rotation angle A Y , we can determine the first sub-rotation matrix R x and the second sub-rotation matrix R y , and then obtain the rotation matrix R between the target posture and the initial posture.

[0091] For example, the number of the second virtual camera can be at least one, and for each second virtual camera, the target posture corresponding to the second virtual camera can be known, and then the rotation matrix R between the target posture of the second virtual camera and the initial posture of the first virtual camera can be obtained. For example, see Figure 4C The figure shows the relationship between the coordinates of the pinhole image and the coordinates of the viewing sphere. Six pinhole camera viewpoints are virtualized at equal angles around the Y axis at the center of the sphere in the viewing sphere coordinate system. The second rotation angle A Y The set is {0, 1 / 3*π, 2 / 3*π, π, 4 / 3*π, 5 / 3*π}, the first rotation angle A X 0. Based on this, 6 virtual cameras can be obtained, and the first rotation angle A of the first virtual camera is X is 0, the second rotation angle A Y is 0, this virtual camera is also the first virtual camera. The first rotation angle A of the second virtual camera X is 0, the second rotation angle A Y =1 / 3*π, this virtual camera is recorded as the second virtual camera 1, and the rotation matrix R between the target posture of the second virtual camera 1 and the initial posture of the first virtual camera is determined based on formula (9)-formula (11). Similarly, the first rotation angle A of the sixth virtual camera is X is 0, the second rotation angle A Y =5 / 3*π, this virtual camera is recorded as the second virtual camera 5, and the rotation matrix R between the target posture of the second virtual camera 5 and the initial posture of the first virtual camera is determined based on formula (9)-formula (11).

[0092] In summary, when there are multiple second virtual cameras, the target postures of different second virtual cameras may be different, and the rotation matrix R corresponding to each second virtual camera may be determined.

[0093] Step 305: Determine an extrinsic parameter matrix between the target pose of the second virtual camera and the initial pose of the first virtual camera based on the rotation matrix between the target pose and the initial pose. For example, a translation matrix between the first virtual camera and the second virtual camera may be determined, and then the extrinsic parameter matrix between the target pose and the initial pose may be determined based on the rotation matrix and the translation matrix.

[0094] For example, since the second virtual camera and the first virtual camera are fixedly connected and have the same optical center, that is, the position of the second virtual camera is the center position of the sphere in the spherical coordinate system, and the position of the first virtual camera is also the center position of the sphere in the spherical coordinate system, the translation matrix between the first virtual camera and the second virtual camera can be That is, there is no translation between the positions of the first virtual camera and the second virtual camera.

[0095] For example, after obtaining the rotation matrix and the translation matrix, the extrinsic parameter matrix can be obtained based on the rotation matrix and the translation matrix. See formula (12), which is an example of determining the extrinsic parameter matrix.

[0096]

[0097] In formula (12), c0 represents the first virtual camera (also known as the reference camera), c i represents the i-th second virtual camera, Represents the rotation matrix between the i-th second virtual camera and the first virtual camera, as shown in formula (9) to formula (11), Represents the translation matrix between the i-th second virtual camera and the first virtual camera, that is, Represents the external parameter matrix between the i-th second virtual camera and the first virtual camera. This will be used in the subsequent multi-camera binding optimization process, see the subsequent examples.

[0098] In summary, when there are multiple second virtual cameras, an extrinsic parameter matrix between each second virtual camera and the first virtual camera may be determined. The extrinsic parameter matrix may include a rotation matrix and a translation matrix.

[0099] Step 306: Determine a second pinhole image corresponding to the second virtual camera based on the rotation matrix between the target pose of the second virtual camera and the initial pose of the first virtual camera. For example, the second pinhole image may be determined based on the rotation matrix and the first pinhole image, or based on the rotation matrix and the viewing spherical image. Of course, the above methods are merely examples and are not limiting in this embodiment.

[0100] For example, see Figure 4B As shown, based on the rotation matrix R between the target posture of the second virtual camera and the initial posture of the first virtual camera, the point Q on the spherical image can be s Rotate and get the rotated point Q s ', and the rotated point Q s 'With point Q before rotation s The relationship can be seen in formula (13):

[0101]

[0102] Obviously, point Q on the visual spherical image s The corresponding point is Q on the first pinhole image, while the corresponding point is Q on the spherical image. s ' corresponds to point Q' on the second pinhole image. Based on this, the relationship between point Q' on the second pinhole image and point Q on the first pinhole image can be shown in formula (14):

[0103]

[0104] In summary, it can be seen that based on the rotation matrix R, the mapping relationship between the coordinates on the first pinhole image and the coordinates on the second pinhole image can be determined. The mapping relationship can be shown in formula (14). Based on this mapping relationship, the first pinhole image can be converted into the second pinhole image, which will not be repeated here.

[0105] In another possible implementation, formula (13) can be substituted into formula (6) and formula (7) to obtain a mapping relationship between the coordinates on the visual spherical image and the coordinates on the second pinhole image. Based on this mapping relationship, the visual spherical image can be converted into the second pinhole image, which will not be repeated here.

[0106] For example, when generating pixels of a pinhole image (such as the first pinhole image or the second pinhole image), interpolation calculations can also be performed. For example, the coordinates of the panoramic image can be floating-point values, and bilinear interpolation can be used to perform interpolation calculations to obtain the pixels of the pinhole image. There is no restriction on this process.

[0107] In summary, the extrinsic parameter matrix between the first and second pinhole images and the virtual cameras can be obtained. For example, assuming there is a first virtual camera, a second virtual camera 1, and a second virtual camera 2, the first pinhole image corresponding to the first virtual camera, the second pinhole image 1 corresponding to the second virtual camera 1, and the second pinhole image 2 corresponding to the second virtual camera 2 are obtained. Furthermore, the extrinsic parameter matrix 11 between the second virtual camera 1 and the first virtual camera and the extrinsic parameter matrix 21 between the second virtual camera 2 and the first virtual camera are obtained.

[0108] 2. Multi-camera binding mapping, for example, based on the first pinhole image, the second pinhole image and the fixed constraint (i.e., the extrinsic parameter matrix), multi-camera binding SFM reconstruction is performed to obtain a three-dimensional visual map.

[0109] Exemplarily, for the first pinhole image and the second pinhole image (the number of the first pinhole image is one, and the number of the second pinhole image is at least one), two-dimensional feature points corresponding to the actual position of the target scene can be selected from the first pinhole image and the second pinhole image (in actual applications, multiple frames of panoramic images at different times can be obtained, and each frame of the panoramic image corresponds to the first pinhole image and the second pinhole image. Multiple two-dimensional feature points corresponding to the actual position can be selected from these first pinhole images and the second pinhole image, that is, these two-dimensional feature points are feature points in all pinhole images at different times). For example, for a certain actual position (that is, the actual physical position) of the target scene, the two-dimensional feature points corresponding to the actual position can be selected from the first pinhole image, and the two-dimensional feature points corresponding to the actual position can be selected from the second pinhole image, that is, multiple two-dimensional feature points are obtained. Based on these two-dimensional feature points and the extrinsic parameter matrix, the three-dimensional map point corresponding to the actual position can be determined. For example, determine the target loss value of the configured loss function; based on the target loss value and the two-dimensional feature point corresponding to the actual position, determine the projection function value between the coordinate system of the virtual camera and the coordinate system of the pinhole image; based on the extrinsic parameter matrix and the projection function value, determine the three-dimensional map point corresponding to the actual position, that is, as the three-dimensional map point in the three-dimensional visual map.

[0110] For example, in the SFM reconstruction process of multi-camera binding, considering the rigid connection constraints between cameras (i.e., the constraints of the external parameter matrix of the cameras in a group), the rigid connection constraints between virtual cameras (i.e., the external parameter matrix) can be added to the optimization. The camera bundling optimization mainly adds the rigid connection constraints between cameras to the reconstruction process for bundling optimization. For each camera rigid connection group set, it is composed of several snapshots with the same camera rigid connection constraints. Each snapshot is composed of images taken by each camera on the camera group at the same time. The external parameters between the cameras in the group are the above-mentioned embodiments. This fixed constraint will be added to the subsequent reconstruction process.

[0111] For example, the reprojection error of a fixed group can be expressed as shown in formula (15):

[0112]

[0113] In formula (15), S k Represents a fixed group. Assuming there are 5 second virtual cameras, then S kIt represents the fixed group composed of these five second virtual cameras. The value range of i is 1-5, that is, c1 represents the first second virtual camera, c2 represents the second second virtual camera, and so on. j represents the two-dimensional feature point, that is, the two-dimensional feature point corresponding to the actual position (the two-dimensional feature point determined from all pinhole images). When j is 1, it represents the first two-dimensional feature point corresponding to the actual position. When j is 2, it represents the second two-dimensional feature point corresponding to the actual position, and so on. ρ j Represents the kernel function of the j-th two-dimensional feature point, which is used to suppress large external point errors. j represents the j-th 2D feature point, i.e., the 2D coordinates of the j-th 2D feature point in the pinhole image. π represents the projection function between the coordinate system of the virtual camera and the coordinate system of the pinhole image, i.e., the projection function from the camera system to the image system. represents the external parameter matrix in the fixed group, c0 represents the first virtual camera, c i represents the i-th second virtual camera, that is, represents the extrinsic parameter matrix between the i-th second virtual camera and the first virtual camera. When i is 1, it represents the extrinsic parameter matrix between the first second virtual camera and the first virtual camera. When i is 2, it represents the extrinsic parameter matrix between the second second virtual camera and the first virtual camera. And so on. The above extrinsic parameter matrices are all obtained extrinsic parameter matrices. T wc Represents the reference camera in the fixed group, usually the pose matrix of camera No. 0, that is, the pose matrix of the first virtual camera. k Indicates the 3D map point corresponding to the actual position, that is, the 3D map point to be obtained, which represents the 3D point of the scene. For any camera in the fixed group, it can be obtained through T wc and Get, that is

[0114] In formula (15), the parameters that need to be optimized include the pose matrix of the reference camera (i.e., the pose matrix T of the first virtual camera) wc ) and 3D map point X k , and the external parameter matrix and feature point x j is known, so in the SFM reconstruction process of multi-camera binding, the above loss function can be minimized, such as by minimizing the above loss function through the LM (Levenberg arquardt) algorithm, and finally the pose matrix T of the reference camera is obtained wc and 3D map point X k , that is, we can get the three-dimensional map point X corresponding to the actual position k .

[0115] In summary, the function of the reprojection error (shown in formula (15)) can be used as the configured loss function, and the above loss function is minimized by the LM algorithm, and the minimized value is used as the target loss value of the loss function. From formula (15), it can be seen that based on the target loss value and the two-dimensional feature point x corresponding to the actual position j , the projection function value between the coordinate system of the virtual camera and the coordinate system of the pinhole image can be determined, that is, Then, based on the external parameter matrix and the projection function value The pose matrix T of the reference camera can be derived wc and 3D map point X k , that is, we can get the three-dimensional map point X corresponding to the actual position k , that is, as a three-dimensional map point in the three-dimensional visual map.

[0116] For multiple actual locations of the target scene, the above method can be used to obtain the three-dimensional map point corresponding to each actual location, thereby constructing a three-dimensional visual map of the target scene based on the multiple three-dimensional map points of the target scene. In other words, the three-dimensional visual map can include multiple three-dimensional map points of the target scene.

[0117] 3. Storing map information corresponding to the 3D visual map, for example, storing map information corresponding to the 3D visual map in a visual feature database. Exemplarily, the map information corresponding to the 3D visual map may include, but is not limited to: a sample global descriptor corresponding to a sample image, a 3D map point corresponding to the sample image, and a sample local descriptor corresponding to the 3D map point. The sample image is a pinhole image selected from the first pinhole image and the second pinhole image. For example, all pinhole images may be used as sample images, or a portion of all pinhole images may be selected as sample images. This is not limited to the above.

[0118] After obtaining the 3D visual map of the target scene, the 3D visual map may include the following information:

[0119] Sample image pose: The sample image is an image used when constructing a three-dimensional visual map, i.e., the first pinhole image and the second pinhole image, etc., that is, a three-dimensional visual map can be constructed based on the sample image, and the pose matrix of the sample image (which can be simply referred to as the sample image pose) can be stored in the three-dimensional visual map, that is, the three-dimensional visual map can include the sample image pose. Referring to the above embodiment, the sample image pose can be the pose matrix T of the reference camera corresponding to the first pinhole image corresponding to the sample image. wc .

[0120] Sample global descriptor: For each frame of sample image, the sample image can correspond to an image global descriptor, which is recorded as the sample global descriptor. The sample global descriptor uses a high-dimensional vector to represent the sample image. The sample global descriptor is used to distinguish the image features of different sample images.

[0121] For each sample image, a bag-of-words vector corresponding to the sample image can be determined based on the trained dictionary model, and the bag-of-words vector can be determined as the sample global descriptor corresponding to the sample image. For example, the bag-of-words method is a method for determining a global descriptor. In the bag-of-words method, a bag-of-words vector can be constructed. The bag-of-words vector is a vector representation method used for image similarity detection, and the bag-of-words vector can be used as the sample global descriptor corresponding to the sample image.

[0122] In the bag-of-visual-words method, a "dictionary" needs to be trained in advance, which can also be called a dictionary model. Generally, feature point descriptors in a large number of images are clustered to obtain a classification tree through training. Each classification tree can represent a visual "word", and these visual "words" constitute the dictionary model.

[0123] For a sample image, all feature point descriptors in the sample image can be classified into "words" and the frequency of occurrence of all words can be counted. In this way, the frequency of each word in the dictionary can form a vector, which is the bag-of-words vector corresponding to the sample image. The bag-of-words vector can be used to measure the similarity between two images and the bag-of-words vector is used as the sample global descriptor corresponding to the sample image.

[0124] For each sample image frame, the sample image can be input into a trained deep learning model to obtain a target vector corresponding to the sample image, and the target vector is determined as the sample global descriptor corresponding to the sample image. For example, a deep learning method is a method for determining a global descriptor. In a deep learning method, a sample image can be subjected to multi-layer convolution by a deep learning model to ultimately obtain a high-dimensional target vector, which is used as the sample global descriptor corresponding to the sample image.

[0125] Deep learning methods require pre-training of deep learning models, such as CNN (Convolutional Neural Networks) models. These models are typically trained using a large number of images, and there are no restrictions on the training method. For a sample image, the sample image can be input into the deep learning model, which processes the sample image to obtain a high-dimensional target vector, which is used as the sample global descriptor corresponding to the sample image.

[0126] Sample local descriptors corresponding to feature points of the sample image: For a sample image, it may include multiple feature points. A feature point is a specific pixel position in the sample image. The feature point may correspond to an image local descriptor, which is recorded as a sample local descriptor. The sample local descriptor uses a vector to describe the features of the image block in the vicinity of the feature point (i.e., the pixel position). The vector may also be called the descriptor of the feature point. In summary, the sample local descriptor is a feature vector used to represent the image block where the feature point is located, and the image block may be located in the sample image. For each feature point of the sample image, the feature point corresponds to a three-dimensional map point in the three-dimensional visual map, that is, the sample local descriptor corresponding to the feature point may be called the sample local descriptor corresponding to the three-dimensional map point.

[0127] Among them, algorithms such as ORB (Oriented FAST and Rotated BRIEF), SIFT (Scale-Invariant Feature Transform), and SURF (Speeded Up Robust Features) can be used to extract feature points from the sample image and determine the sample local descriptors corresponding to the feature points. Deep learning algorithms (such as SuperPoint, DELF, D2-Net, etc.) can also be used to extract feature points from the sample image and determine the sample local descriptors corresponding to the feature points. There is no restriction on this, as long as feature points can be obtained and the sample local descriptors can be determined.

[0128] Map point information of a three-dimensional map point: The map point information may include, but is not limited to, the 3D spatial position of the three-dimensional map point, all observed sample images, and the numbers of the corresponding 2D feature points.

[0129] 4. Online positioning process. For example, the target image of the target scene can be collected, and the posture of the target image can be solved based on the three-dimensional visual map of the target scene to obtain the global positioning posture in the three-dimensional visual map corresponding to the target image, and the positioning process is completed. For example, when it is necessary to globally position the terminal device based on the three-dimensional visual map, the terminal device can download the three-dimensional visual map from the server, and the terminal device can be positioned based on the three-dimensional visual map during the movement of the terminal device in the target scene. The target scene can be an indoor environment, that is, when the terminal device moves in the indoor environment, the global positioning posture of the terminal device in the three-dimensional visual map can be determined. Of course, the target scene can also be an outdoor environment, etc., and there is no restriction on this target scene. In the above embodiment, the posture (such as the global positioning posture, etc.) can be position and posture, which are generally represented by a rotation matrix and a translation vector, and there is no restriction on this.

[0130] The terminal device may include a visual sensor, such as a camera, etc. The visual sensor is used to collect a target image of a target scene (ie, a real-time image during the movement) during the movement of the terminal device.

[0131] Among them, the terminal device can be a wearable device (such as a video helmet, smart watch, smart glasses, etc.), and the visual sensor is deployed on the wearable device; or, the terminal device is a recorder (such as a device carried by staff when performing work, which has real-time audio and video acquisition, photography, recording, intercom, positioning and other functions in one), and the visual sensor is deployed on the recorder; or, the terminal device is a camera (such as a split camera, etc.), and the visual sensor is deployed on the camera. Or, the terminal device is a robot, and the visual sensor is deployed on the robot. Or, the terminal device is an autonomous driving vehicle, and the visual sensor is deployed on the autonomous driving vehicle. Of course, the above are just a few examples, and there is no limitation to this. For example, it can also be a smart phone, etc., as long as the terminal device is deployed with a visual sensor.

[0132] Exemplarily, based on the three-dimensional visual map of the target scene, after obtaining the target image, the target three-dimensional map point corresponding to the target image can be determined from the three-dimensional visual map of the target scene, and the global positioning posture of the terminal device in the three-dimensional visual map can be determined based on the target three-dimensional map point.

[0133] For example, based on a three-dimensional visual map of a target scene, in one possible implementation, the following steps may be used to determine the global positioning pose of the terminal device in the three-dimensional visual map:

[0134] Step S11: During the global positioning process of the terminal device, a target image of the terminal device in the target scene is obtained, such as by collecting the target image in the target scene through a visual sensor, that is, a real-time video image.

[0135] Step S12: Determine the global descriptor to be tested corresponding to the target image.

[0136] Exemplarily, the target image may correspond to a global image descriptor, which may be recorded as a global descriptor to be tested. The global descriptor to be tested represents the target image using a high-dimensional vector, and the global descriptor to be tested is used to distinguish image features of different target images. For example, a bag-of-words vector corresponding to the target image is determined based on a trained dictionary model, and the bag-of-words vector is determined as the global descriptor to be tested corresponding to the target image. Alternatively, the target image is input into a trained deep learning model to obtain a target vector corresponding to the target image, and the target vector is determined as the global descriptor to be tested corresponding to the target image.

[0137] In summary, the global descriptor to be tested corresponding to the target image can be determined based on the bag-of-visual-words method or the deep learning method. The determination method refers to the method for determining the sample global descriptor, which will not be repeated here.

[0138] Step S13: determining the distance between the global descriptor to be tested (ie, the global descriptor to be tested corresponding to the target image) and the sample global descriptor corresponding to each frame of the sample image corresponding to the three-dimensional visual map.

[0139] Referring to the above embodiment, the three-dimensional visual map may include a sample global descriptor corresponding to each frame of the sample image. Therefore, the distance between the global descriptor to be tested and the sample global descriptor corresponding to each frame of the sample image can be determined, such as the Euclidean distance, that is, the Euclidean distance between the two feature vectors is calculated.

[0140] Step S14: Based on the distance between the global descriptor to be tested and each sample global descriptor, a candidate sample image is selected from the multiple frames of sample images corresponding to the three-dimensional visual map; wherein the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is the minimum distance; or, the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold.

[0141] For example, assuming that the three-dimensional visual map corresponds to sample image 1, sample image 2 and sample image 3, the distance 1 between the global descriptor to be tested and the sample global descriptor corresponding to sample image 1 can be calculated, and the distance 2 between the global descriptor to be tested and the sample global descriptor corresponding to sample image 2 can be calculated, and the distance 3 between the global descriptor to be tested and the sample global descriptor corresponding to sample image 3 can be calculated.

[0142] In one possible implementation, if distance 1 is the minimum distance, sample image 1 is selected as the candidate sample image. Alternatively, if distance 1 is less than a distance threshold (which can be configured based on experience), and distance 2 is less than the distance threshold, but distance 3 is not less than the distance threshold, both sample image 1 and sample image 2 are selected as candidate sample images. Alternatively, if distance 1 is the minimum distance and is less than the distance threshold, sample image 1 is selected as the candidate sample image. However, if distance 1 is the minimum distance and is not less than the distance threshold, no candidate sample image can be selected, i.e., global positioning fails.

[0143] In summary, candidate sample images can be selected from multiple frames of sample images corresponding to the three-dimensional visual map.

[0144] Step S15: Acquire multiple feature points from the target image. Exemplarily, for each feature point, a local descriptor to be tested corresponding to the feature point can be determined. The local descriptor to be tested can be used to represent a feature vector of an image block where the feature point is located, and the image block can be located in the target image.

[0145] For example, a target image may include multiple feature points. A feature point can be a specific pixel location in the target image. Each feature point can correspond to a local image descriptor, which is denoted as the local descriptor to be tested. The local descriptor to be tested uses a vector to describe the characteristics of the image block within the vicinity of the feature point (i.e., the pixel location). This vector can also be called the descriptor of the feature point. In summary, the local descriptor to be tested is a feature vector used to represent the image block where the feature point is located.

[0146] Among them, ORB, SIFT, SURF and other algorithms can be used to extract feature points from the target image and determine the local descriptors to be tested corresponding to the feature points. Deep learning algorithms (such as SuperPoint, DELF, D2-Net, etc.) can also be used to extract feature points from the target image and determine the local descriptors to be tested corresponding to the feature points. There is no restriction on this, as long as the feature points can be obtained and the local descriptors to be tested can be determined.

[0147] Step S16: For each feature point corresponding to the target image, determine the distance between the local descriptor to be tested corresponding to the feature point and the sample local descriptor corresponding to each three-dimensional map point corresponding to the candidate sample image, such as the Euclidean distance, that is, calculate the Euclidean distance between the two feature vectors.

[0148] Referring to the above embodiment, for each frame of sample image, the three-dimensional visual map includes a sample local descriptor corresponding to each three-dimensional map point corresponding to the sample image. After obtaining the candidate sample image, the sample local descriptor corresponding to each three-dimensional map point corresponding to the candidate sample image can be obtained from the three-dimensional visual map.

[0149] After obtaining each feature point corresponding to the target image, the distance between the local descriptor to be tested corresponding to the feature point and the sample local descriptor corresponding to each three-dimensional map point corresponding to the candidate sample image is determined.

[0150] Step S17: For each feature point, based on the distance between the local descriptor to be tested corresponding to the feature point and the sample local descriptor corresponding to each three-dimensional map point corresponding to the candidate sample image, select a target three-dimensional map point from the three-dimensional map points corresponding to the candidate sample image; the distance between the local descriptor to be tested and the sample local descriptor corresponding to the target three-dimensional map point is a minimum distance, and the minimum distance is less than the distance threshold.

[0151] For example, assuming that the candidate sample image corresponds to 3D map point 1, 3D map point 2, and 3D map point 3, then the distance 1 between the local descriptor to be tested and the sample local descriptor corresponding to 3D map point 1 can be calculated, the distance 2 between the local descriptor to be tested and the sample local descriptor corresponding to 3D map point 2 can be calculated, and the distance 3 between the local descriptor to be tested and the sample local descriptor corresponding to 3D map point 3 can be calculated.

[0152] In one possible implementation, if Distance 1 is the minimum distance, 3D map point 1 is selected as the target 3D map point. Alternatively, if Distance 1 is less than the distance threshold, and Distance 2 is less than the distance threshold, but Distance 3 is not less than the distance threshold, then both 3D map point 1 and 3D map point 2 may be selected as the target 3D map point. Alternatively, if Distance 1 is the minimum distance and Distance 1 is less than the distance threshold, then 3D map point 1 may be selected as the target map point. However, if Distance 1 is the minimum distance and Distance 1 is not less than the distance threshold, then the target 3D map point cannot be selected, i.e., global positioning fails.

[0153] For each feature point of the target image, a target 3D map point corresponding to the feature point is selected from the candidate sample images corresponding to the target image to obtain a matching relationship between the feature point and the target 3D map point.

[0154] Step S18: Determine a global positioning pose in the three-dimensional visual map corresponding to the target image based on the multiple feature points corresponding to the target image and the target three-dimensional map points corresponding to the multiple feature points.

[0155] For example, feature point 1 of the target image corresponds to three-dimensional map point 1, feature point 2 of the target image corresponds to three-dimensional map point 2, and so on, thereby obtaining multiple matching relationship pairs, each matching relationship pair includes a two-dimensional feature point and a three-dimensional map point, the two-dimensional feature point represents the two-dimensional position in the target image, and the three-dimensional map point represents the three-dimensional position in the three-dimensional visual map, that is, the matching relationship pair includes a mapping relationship from a two-dimensional position to a three-dimensional position, that is, a mapping relationship from a two-dimensional position in the target image to a three-dimensional position in the three-dimensional visual map.

[0156] If the total number of multiple matching relationship pairs does not meet the quantity requirement, it means that the global positioning pose in the three-dimensional visual map corresponding to the target image cannot be determined based on the multiple matching relationship pairs. If the total number of multiple matching relationship pairs meets the quantity requirement (i.e., the total number reaches the preset quantity value), it means that the global positioning pose in the three-dimensional visual map corresponding to the target image can be determined based on the multiple matching relationship pairs, that is, the global positioning pose in the three-dimensional visual map corresponding to the target image can be determined based on the multiple matching relationship pairs.

[0157] For example, a PnP (Perspective NPoint) algorithm can be used to calculate the global positioning pose of a target image within a 3D visual map, with no restrictions on the calculation method. For example, the input data of the PnP algorithm is multiple matching relationship pairs, each of which includes a 2D position in the target image and a 3D position in the 3D visual map. Based on these multiple matching relationship pairs, the PnP algorithm can be used to calculate the pose of the target image within the 3D visual map, i.e., the global positioning pose.

[0158] In one possible implementation, after obtaining multiple matching pairs, valid matching pairs can be found from the multiple matching pairs. Based on these valid matching pairs, a PnP algorithm can be used to calculate the global positioning pose of the target image in the three-dimensional visual map. For example, a RANSAC (RANdom SAmple Consensus) detection algorithm can be used to find valid matching pairs from all matching pairs, and there is no restriction on this process.

[0159] In summary, the global positioning pose of the terminal device in the three-dimensional visual map can be determined.

[0160] As can be seen from the above technical solutions, in the embodiments of the present application, a three-dimensional visual map of the target scene can be constructed, and the terminal device of the target scene can be globally positioned based on the three-dimensional visual map, so as to accurately position the terminal device. A panoramic image of the target scene can be used to construct a three-dimensional visual map. The panoramic image has a large field of view, which avoids repeated data collection of the target scene and improves data collection efficiency. When determining the three-dimensional map points, the pinhole image can be obtained by using the projection method of the virtual camera, and the posture constraints of the virtual camera can be added. The three-dimensional map points can be determined by using the virtual camera binding optimization strategy to improve the robustness, efficiency and accuracy of the map. By using a panoramic camera to implement the SFM reconstruction process, the accuracy and robustness of the map can be improved. By adding a multi-camera binding optimization strategy to constrain the fixed connection relationship between virtual cameras, the accuracy of the map can be improved. According to the spherical projection relationship, N virtual cameras that are fixed to each other are obtained, and the fixed connection constraint is added to the reconstruction process. This constraint can improve the accuracy of multi-camera registration and reduce the proportion of erroneous registration. In addition, the idea of ​​optimizing the virtual pinhole camera's fixed connection constraint binding can be extended to a multi-camera system with hardware fixed connection. The algorithm is universal and only requires modifying the camera's fixed connection external parameters. This will not be described in detail here.

[0161] Based on the same application concept as the above method, a map construction device is proposed in the embodiment of the present application. Figure 5 FIG. 1 is a structural diagram of the map construction device, which may include:

[0162] An acquisition module 51 is used to acquire a panoramic image of a target scene;

[0163] A generating module 52 is configured to generate a first pinhole image corresponding to a first virtual camera based on the panoramic image; wherein the position of the first virtual camera is the center position of a sphere in a spherical coordinate system, and the initial posture of the first virtual camera is any posture centered at the center position of the sphere;

[0164] a determination module 53 for determining a rotation matrix between a target pose of a second virtual camera and the initial pose, and determining an extrinsic parameter matrix between the target pose and the initial pose based on the rotation matrix; wherein the position of the second virtual camera is the center position of a sphere in the viewing spherical coordinate system, and the target pose is obtained by rotating the initial pose around the coordinate axis of the viewing spherical coordinate system;

[0165] The generation module 52 is also used to determine the second pinhole image corresponding to the second virtual camera based on the rotation matrix; the determination module 53 is also used to select two-dimensional feature points corresponding to the actual position of the target scene from the first pinhole image and the second pinhole image, and determine the three-dimensional map points corresponding to the actual position based on the two-dimensional feature points and the external parameter matrix; and, construct a three-dimensional visual map of the target scene based on multiple three-dimensional map points of the target scene.

[0166] Exemplarily, when the generation module 52 generates the first pinhole image corresponding to the first virtual camera based on the panoramic image, it is specifically used to: generate the visual spherical image corresponding to the visual spherical coordinate system based on the panoramic image; and generate the first pinhole image corresponding to the first virtual camera based on the visual spherical image.

[0167] Exemplarily, when the generation module 52 generates the visual spherical image corresponding to the visual spherical coordinate system based on the panoramic image, it is specifically used to: determine the mapping relationship between the longitude and latitude coordinates in the visual spherical image and the rectangular coordinates in the panoramic image based on the width and height of the panoramic image; for each longitude and latitude coordinate in the visual spherical image, determine the rectangular coordinate corresponding to the longitude and latitude coordinate from the panoramic image based on the mapping relationship, and determine the pixel value of the longitude and latitude coordinate based on the pixel value of the rectangular coordinate; generate the visual spherical image based on the pixel value of each longitude and latitude coordinate in the visual spherical image.

[0168] Exemplarily, when the generation module 52 generates the first pinhole image corresponding to the first virtual camera based on the visual spherical image, it is specifically used to: determine the center point coordinates of the first pinhole image based on the width and height of the first pinhole image, and determine the mapping relationship between the rectangular coordinates in the first pinhole image and the longitude and latitude coordinates in the visual spherical image based on the center point coordinates and the target distance, and the target distance is the distance between the center point of the first pinhole image and the center position of the sphere of the visual spherical coordinate system; for each rectangular coordinate in the first pinhole image, determine the longitude and latitude coordinates corresponding to the rectangular coordinate from the visual spherical image based on the mapping relationship, and determine the pixel value of the rectangular coordinate based on the pixel value of the longitude and latitude coordinate; generate the first pinhole image based on the pixel value of each rectangular coordinate in the first pinhole image.

[0169] Exemplarily, when the determination module 53 determines the rotation matrix between the target posture and the initial posture of the second virtual camera, it is specifically used to: determine the first rotation angle between the target posture and the initial posture in the direction of the first coordinate axis, and determine the first sub-rotation matrix in the direction of the first coordinate axis based on the first rotation angle; determine the second rotation angle between the target posture and the initial posture in the direction of the second coordinate axis, and determine the second sub-rotation matrix in the direction of the second coordinate axis based on the second rotation angle; determine the third rotation angle between the target posture and the initial posture in the direction of the third coordinate axis, and determine the third sub-rotation matrix in the direction of the third coordinate axis based on the third rotation angle; determine the rotation matrix between the target posture and the initial posture based on the first sub-rotation matrix, the second sub-rotation matrix and the third sub-rotation matrix.

[0170] Exemplarily, when the determination module 53 determines the extrinsic parameter matrix between the target posture and the initial posture based on the rotation matrix, it is specifically used to: determine the translation matrix between the first virtual camera and the second virtual camera; and determine the extrinsic parameter matrix based on the rotation matrix and the translation matrix.

[0171] Exemplarily, when the determination module 53 determines the three-dimensional map point corresponding to the actual position based on the two-dimensional feature points and the extrinsic parameter matrix, it is specifically used to: determine the target loss value of the configured loss function; determine the projection function value between the coordinate system of the virtual camera and the coordinate system of the pinhole image based on the target loss value and the two-dimensional feature points corresponding to the actual position; and determine the three-dimensional map point corresponding to the actual position based on the extrinsic parameter matrix and the projection function value.

[0172] Exemplarily, the three-dimensional visual map includes a sample global descriptor corresponding to a sample image, a three-dimensional map point corresponding to the sample image, and a sample local descriptor corresponding to the three-dimensional map point; the sample image is a pinhole image selected from a first pinhole image and a second pinhole image; the device also includes a positioning module for obtaining a target image of the terminal device in a target scene during the global positioning process of the terminal device; based on the similarity between the target image and the multi-frame sample images corresponding to the three-dimensional visual map, a candidate sample image is selected from the multi-frame sample images; a plurality of feature points are obtained from the target image; for each feature point, a target three-dimensional map point corresponding to the feature point is determined from the three-dimensional map points corresponding to the candidate sample image; based on the multiple feature points and the target three-dimensional map points corresponding to the multiple feature points, the global positioning posture in the three-dimensional visual map corresponding to the target image is determined.

[0173] Exemplarily, the positioning module is specifically used to select candidate sample images from the multiple frames of sample images based on the similarity between the target image and the multiple frames of sample images corresponding to the three-dimensional visual map: determine the global descriptor to be tested corresponding to the target image, determine the distance between the global descriptor to be tested and the sample global descriptor corresponding to each frame of sample image corresponding to the three-dimensional visual map; based on the distance between the global descriptor to be tested and each sample global descriptor, select candidate sample images from the multiple frames of sample images; wherein the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is the minimum distance; or, the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold.

[0174] Exemplarily, when the positioning module determines the target three-dimensional map point corresponding to the feature point from the three-dimensional map points corresponding to the candidate sample image, it is specifically used to: determine the local descriptor to be tested corresponding to the feature point, the local descriptor to be tested is used to represent the feature vector of the image block where the feature point is located, and the image block is located in the target image; determine the distance between the local descriptor to be tested and the sample local descriptor corresponding to each three-dimensional map point corresponding to the candidate sample image; based on the distance between the local descriptor to be tested and each sample local descriptor, select the target three-dimensional map point from the three-dimensional map points corresponding to the candidate sample image; wherein the distance between the local descriptor to be tested and the sample local descriptor corresponding to the target three-dimensional map point is the minimum distance, and the minimum distance is less than the distance threshold.

[0175] Based on the same application concept as the above method, an embodiment of the present application proposes a map construction device, which may include: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the map construction method disclosed in the above example of the present application.

[0176] Based on the same application concept as the above method, an embodiment of the present application further provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the map construction method disclosed in the above example of the present application can be implemented.

[0177] The machine-readable storage medium may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.

[0178] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0179] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0180] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0181] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0182] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0183] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A map construction method, characterized in that: The method comprises: Acquire a panoramic image of a target scene, generate a visual spherical image corresponding to a visual spherical coordinate system based on the panoramic image; generate a first pinhole image corresponding to a first virtual camera based on the visual spherical image; wherein the position of the first virtual camera is the center position of a sphere in the visual spherical coordinate system, and the initial posture of the first virtual camera is any posture centered at the center position of the sphere; generating the first pinhole image corresponding to the first virtual camera based on the visual spherical image includes: determining the center point coordinates of the first pinhole image based on the width and height of the first pinhole image, determining a mapping relationship between rectangular coordinates in the first pinhole image and longitude and latitude coordinates in the visual spherical image based on the center point coordinates and a target distance, wherein the target distance is the distance between the center point of the first pinhole image and the center position of the sphere in the visual spherical coordinate system; for each rectangular coordinate in the first pinhole image, determine the longitude and latitude coordinates corresponding to the rectangular coordinate from the visual spherical image based on the mapping relationship, and determine the pixel value of the rectangular coordinate based on the pixel value of the longitude and latitude coordinate; generate the first pinhole image based on the pixel value of each rectangular coordinate in the first pinhole image; Determining a rotation matrix between a target pose of a second virtual camera and the initial pose, and determining an extrinsic parameter matrix between the target pose and the initial pose based on the rotation matrix; wherein the position of the second virtual camera is the center position of a sphere in the viewing spherical coordinate system, and the target pose is obtained by rotating the initial pose around the coordinate axis of the viewing spherical coordinate system; Determining a second pinhole image corresponding to a second virtual camera based on the rotation matrix, selecting two-dimensional feature points corresponding to an actual position of a target scene from the first pinhole image and the second pinhole image, and determining a three-dimensional map point corresponding to the actual position based on the two-dimensional feature points and the extrinsic parameter matrix; A three-dimensional visual map of the target scene is constructed based on a plurality of three-dimensional map points of the target scene.

2. The method according to claim 1, characterized in that Generating a visual spherical image corresponding to the visual spherical coordinate system based on the panoramic image includes: Determining, based on the width and height of the panoramic image, a mapping relationship between longitude and latitude coordinates in the visual spherical image and rectangular coordinates in the panoramic image; for each longitude and latitude coordinate in the visual spherical image, determining, from the panoramic image based on the mapping relationship, a rectangular coordinate corresponding to the longitude and latitude coordinate, and determining a pixel value of the longitude and latitude coordinate based on a pixel value of the rectangular coordinate; The visual spherical image is generated based on the pixel value of each latitude and longitude coordinate in the visual spherical image.

3. The method according to claim 1, characterized in that Determining a rotation matrix between a target pose of the second virtual camera and the initial pose includes: Determine a first rotation angle between the target posture and the initial posture in the first coordinate axis direction, and determine a first sub-rotation matrix in the first coordinate axis direction based on the first rotation angle; Determine a second rotation angle between the target posture and the initial posture in the second coordinate axis direction, and determine a second sub-rotation matrix in the second coordinate axis direction based on the second rotation angle; Determining a third rotation angle between the target posture and the initial posture in the direction of a third coordinate axis, and determining a third sub-rotation matrix in the direction of the third coordinate axis based on the third rotation angle; A rotation matrix between the target posture and the initial posture is determined based on the first sub-rotation matrix, the second sub-rotation matrix, and the third sub-rotation matrix.

4. The method according to claim 1, wherein The determining of an extrinsic parameter matrix between the target posture and the initial posture based on the rotation matrix includes: determining a translation matrix between the first virtual camera and the second virtual camera; The extrinsic parameter matrix is ​​determined based on the rotation matrix and the translation matrix.

5. The method according to claim 1, wherein The determining of the three-dimensional map point corresponding to the actual position based on the two-dimensional feature point and the extrinsic parameter matrix includes: Determine the target loss value for the configured loss function; Determining a projection function value between a coordinate system of a virtual camera and a coordinate system of a pinhole image based on the target loss value and the two-dimensional feature points corresponding to the actual position; A three-dimensional map point corresponding to the actual position is determined based on the extrinsic parameter matrix and the projection function value.

6. The method according to claim 1, characterized in that The three-dimensional visual map includes a sample global descriptor corresponding to a sample image, a three-dimensional map point corresponding to the sample image, and a sample local descriptor corresponding to the three-dimensional map point; wherein the sample image is a pinhole image selected from the first pinhole image and the second pinhole image; After constructing the three-dimensional visual map of the target scene, the method further includes: During the global positioning process of the terminal device, a target image of the terminal device in the target scene is obtained; based on the similarity between the target image and multiple frames of sample images corresponding to the three-dimensional visual map, a candidate sample image is selected from the multiple frames of sample images; Acquire multiple feature points from the target image; for each feature point, determine a target three-dimensional map point corresponding to the feature point from the three-dimensional map points corresponding to the candidate sample images; Based on the multiple feature points and the target three-dimensional map points corresponding to the multiple feature points, a global positioning pose in the three-dimensional visual map corresponding to the target image is determined.

7. The method according to claim 6, characterized in that The selecting a candidate sample image from the multiple frames of sample images based on the similarity between the target image and the multiple frames of sample images corresponding to the three-dimensional visual map includes: Determining a global descriptor to be tested corresponding to the target image, and determining a distance between the global descriptor to be tested and a sample global descriptor corresponding to each frame of the sample image corresponding to the three-dimensional visual map; Based on the distance between the global descriptor to be tested and each sample global descriptor, a candidate sample image is selected from the multiple frames of sample images; wherein the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is the minimum distance; or, the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold.

8. The method according to claim 6, characterized in that The step of determining a target three-dimensional map point corresponding to the feature point from the three-dimensional map points corresponding to the candidate sample images includes: Determine a local descriptor to be tested corresponding to the feature point, where the local descriptor to be tested is used to represent a feature vector of an image block where the feature point is located, and the image block is located in the target image; determining a distance between the local descriptor to be tested and a sample local descriptor corresponding to each three-dimensional map point corresponding to the candidate sample image; and selecting a target three-dimensional map point from the three-dimensional map points corresponding to the candidate sample image based on the distance between the local descriptor to be tested and each sample local descriptor; The distance between the local descriptor to be tested and the sample local descriptor corresponding to the target three-dimensional map point is a minimum distance, and the minimum distance is less than a distance threshold.

9. A map construction device, characterized in that: The device comprises: An acquisition module, used to acquire a panoramic image of a target scene; a generation module configured to generate a spherical image corresponding to a spherical coordinate system based on the panoramic image; and to generate a first pinhole image corresponding to a first virtual camera based on the spherical image; wherein the position of the first virtual camera is the center position of a sphere in the spherical coordinate system, and the initial posture of the first virtual camera is any posture centered at the center position of the sphere; when generating the first pinhole image corresponding to the first virtual camera based on the spherical image, the generation module is configured to: determine the center point coordinates of the first pinhole image based on the width and height of the first pinhole image, determine a mapping relationship between rectangular coordinates in the first pinhole image and longitude and latitude coordinates in the spherical image based on the center point coordinates and a target distance, wherein the target distance is the distance between the center point of the first pinhole image and the center position of the spherical coordinate system; for each rectangular coordinate in the first pinhole image, determine the longitude and latitude coordinates corresponding to the rectangular coordinate from the spherical image based on the mapping relationship, and determine the pixel value of the rectangular coordinate based on the pixel value of the longitude and latitude coordinate; and generate the first pinhole image based on the pixel value of each rectangular coordinate in the first pinhole image; a determination module, configured to determine a rotation matrix between a target pose of a second virtual camera and the initial pose, and determine an extrinsic parameter matrix between the target pose and the initial pose based on the rotation matrix; wherein the position of the second virtual camera is the center position of a sphere in the viewing spherical coordinate system, and the target pose is obtained by rotating the initial pose around the coordinate axis of the viewing spherical coordinate system; The generation module is further used to determine the second pinhole image corresponding to the second virtual camera based on the rotation matrix; the determination module is further used to select two-dimensional feature points corresponding to the actual position of the target scene from the first pinhole image and the second pinhole image, and determine the three-dimensional map points corresponding to the actual position based on the two-dimensional feature points and the extrinsic parameter matrix; and, construct a three-dimensional visual map of the target scene based on multiple three-dimensional map points of the target scene.

10. The device according to claim 9, It is characterized in that Wherein, when the generating module generates the visual spherical image corresponding to the visual spherical coordinate system based on the panoramic image, it is specifically configured to: determine, based on the width and height of the panoramic image, a mapping relationship between the longitude and latitude coordinates in the visual spherical image and the rectangular coordinates in the panoramic image; for each longitude and latitude coordinate in the visual spherical image, determine, from the panoramic image based on the mapping relationship, the rectangular coordinate corresponding to the longitude and latitude coordinate, and determine the pixel value of the longitude and latitude coordinate based on the pixel value of the rectangular coordinate; and generate the visual spherical image based on the pixel value of each longitude and latitude coordinate in the visual spherical image; Wherein, when determining the rotation matrix between the target posture and the initial posture of the second virtual camera, the determination module is specifically used to: determine a first rotation angle between the target posture and the initial posture in the direction of the first coordinate axis, and determine a first sub-rotation matrix in the direction of the first coordinate axis based on the first rotation angle; determine a second rotation angle between the target posture and the initial posture in the direction of the second coordinate axis, and determine a second sub-rotation matrix in the direction of the second coordinate axis based on the second rotation angle; determine a third rotation angle between the target posture and the initial posture in the direction of the third coordinate axis, and determine a third sub-rotation matrix in the direction of the third coordinate axis based on the third rotation angle; determine the rotation matrix between the target posture and the initial posture based on the first sub-rotation matrix, the second sub-rotation matrix and the third sub-rotation matrix; Wherein, when the determination module determines the extrinsic parameter matrix between the target posture and the initial posture based on the rotation matrix, it is specifically used to: determine the translation matrix between the first virtual camera and the second virtual camera; determine the extrinsic parameter matrix based on the rotation matrix and the translation matrix; Wherein, when determining the three-dimensional map point corresponding to the actual position based on the two-dimensional feature points and the extrinsic parameter matrix, the determination module is specifically used to: determine a target loss value of a configured loss function; determine a projection function value between a coordinate system of a virtual camera and a coordinate system of a pinhole image based on the target loss value and the two-dimensional feature points corresponding to the actual position; and determine the three-dimensional map point corresponding to the actual position based on the extrinsic parameter matrix and the projection function value; Wherein, the three-dimensional visual map includes a sample global descriptor corresponding to a sample image, a three-dimensional map point corresponding to the sample image, and a sample local descriptor corresponding to the three-dimensional map point; the sample image is a pinhole image selected from a first pinhole image and a second pinhole image; the apparatus further includes: a positioning module for obtaining a target image of the terminal device in a target scene during a global positioning process of the terminal device; selecting a candidate sample image from the multiple frames of sample images based on a similarity between the target image and the multiple frames of sample images corresponding to the three-dimensional visual map; obtaining a plurality of feature points from the target image; for each feature point, determining a target three-dimensional map point corresponding to the feature point from the three-dimensional map points corresponding to the candidate sample image; and determining a global positioning pose in the three-dimensional visual map corresponding to the target image based on the multiple feature points and the target three-dimensional map points corresponding to the multiple feature points. Wherein, the positioning module is specifically used to select candidate sample images from the multiple frames of sample images based on the similarity between the target image and the multiple frames of sample images corresponding to the three-dimensional visual map: determine the global descriptor to be tested corresponding to the target image, determine the distance between the global descriptor to be tested and the sample global descriptor corresponding to each frame of sample image corresponding to the three-dimensional visual map; based on the distance between the global descriptor to be tested and each sample global descriptor, select candidate sample images from the multiple frames of sample images; wherein the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is the minimum distance; or the distance between the global descriptor to be tested and the sample global descriptor corresponding to the candidate sample image is less than a distance threshold; Among them, when the positioning module determines the target three-dimensional map point corresponding to the feature point from the three-dimensional map points corresponding to the candidate sample image, it is specifically used to: determine the local descriptor to be tested corresponding to the feature point, the local descriptor to be tested is used to represent the feature vector of the image block where the feature point is located, and the image block is located in the target image; determine the distance between the local descriptor to be tested and the sample local descriptor corresponding to each three-dimensional map point corresponding to the candidate sample image; based on the distance between the local descriptor to be tested and each sample local descriptor, select the target three-dimensional map point from the three-dimensional map points corresponding to the candidate sample image; wherein the distance between the local descriptor to be tested and the sample local descriptor corresponding to the target three-dimensional map point is the minimum distance, and the minimum distance is less than the distance threshold.

11. A map construction device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions capable of being executed by the processor; The processor is configured to execute machine-executable instructions to implement the method steps described in any one of claims 1-8.

Citation Information

Patent Citations

  • Attitude determination, panoramic image generation and target recognition methods for intelligent machine

    CN105474033A

  • Image processing method and device

    WO2020103075A1