Three-dimensional pose generation method and device, electronic equipment and computer readable medium

Through the three-dimensional pose generation method combining depth information, the problems of low quality and long time in the prior art are solved, and higher quality and faster 3-dimensional pose generation are achieved.

CN119919494APending Publication Date: 2025-05-02LINGBAN INTELLIGENT (HANGZHOU) INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411997145.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing three-dimensional pose generation method has high positioning accuracy in areas with rich textures, but poor positioning accuracy in areas without textures or reflective lights, resulting in poor quality of 3-dimensional pose generation, and the ICP point cloud registration of the depth map positioning algorithm takes a long time.

Method used

Using a three-dimensional pose generation method combining depth information, by receiving the to-located image with image depth information, a global feature vector to-located and candidate image information group is generated, feature points and local feature matching information are extracted, a two-dimensional three-dimensional matching pair group is generated, the initial pose information is obtained, and the final pose information is generated through the point cloud data group.

Benefits of technology

Improves the quality of three-dimensional pose generation, reduces the time-consuming to optimize to the correct pose, and enhances positioning accuracy in textureless or reflective areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919494A_ABST
    Figure CN119919494A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a three-dimensional pose generation method and device, electronic equipment and a computer readable medium. A specific embodiment of the method comprises the steps of receiving a to-be-positioned image; generating a global feature vector to be positioned and a candidate image information group; generating a to-be-positioned feature point information group and a local feature matching information group set; generating a two-dimensional and three-dimensional matching pair group; generating initial pose information; generating a first local point cloud data set based on the initial pose information and a pre-stored point cloud data set; generating a second local point cloud data set based on the image depth information and pre-stored camera internal reference information; and generating final pose information based on the first local point cloud data set, the second local point cloud data set and the initial pose information. According to the embodiment, the three-dimensional pose generation quality is improved, and the time consumed for optimizing to the correct pose is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a three-dimensional pose generation method, device, electronic device, and computer-readable medium. Background Art

[0002] Three-dimensional pose is a method of determining the position and pose of a device in three-dimensional space by analyzing sensor data using spatial point cloud positioning technology. The commonly used method for obtaining three-dimensional pose is to obtain image data of the environment based on a positioning algorithm based on visual feature matching. Then, corner points with clear texture and obvious features in the space are extracted to form a point cloud map. Then, the 2D-3D feature point matching relationship is obtained through feature point matching to solve the pose. Alternatively, a depth camera is used to convert the positioning depth into a 3D point cloud, which is then ICP (Iterative Closest Point) registered with the 3D point cloud of the map to obtain the pose of the device.

[0003] However, when the above method is used to obtain the three-dimensional position and posture of the device, the following technical problems often occur:

[0004] The pure visual positioning algorithm relies solely on visual RGB information, resulting in uneven feature points in the scene. The positioning accuracy is high for areas with rich textures, but poor in areas without textures or reflective areas, resulting in poor quality of 3D pose generation. The depth map positioning algorithm uses the ICP point cloud registration algorithm to gradually iterate and optimize to the correct pose, which takes a long time.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention

[0006] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0007] Some embodiments of the present disclosure propose a method, device, electronic device and computer-readable medium for generating a three-dimensional pose in combination with depth information to solve one or more of the technical problems mentioned in the above background technology section.

[0008] In a first aspect, some embodiments of the present disclosure provide a method for generating a three-dimensional pose in combination with depth information, the method comprising: receiving an image to be positioned, wherein the image to be positioned corresponds to image depth information; generating a global feature vector to be positioned and a candidate image information group based on the image to be positioned and a preset database; generating a feature point information group to be positioned and a local feature matching information group set based on the image to be positioned and the candidate image information group; generating a two-dimensional and three-dimensional matching pair group based on the feature point information group to be positioned and the local feature matching information group set; generating initial pose information based on the two-dimensional and three-dimensional matching pair group; generating a first local point cloud data group based on the initial pose information and a pre-stored point cloud data group; generating a second local point cloud data group based on the image depth information and pre-stored camera intrinsic parameter information; generating final pose information based on the first local point cloud data group, the second local point cloud data group and the initial pose information.

[0009] In a second aspect, some embodiments of the present disclosure provide a three-dimensional pose generation device combined with depth information, the device comprising: a receiving unit, configured to receive an image to be positioned, wherein the image to be positioned corresponds to image depth information; a first generating unit, configured to generate a global feature vector to be positioned and a candidate image information group based on the image to be positioned and a preset database; a second generating unit, configured to generate a feature point information group to be positioned and a local feature matching information group set based on the image to be positioned and the candidate image information group; a third generating unit, configured to generate a feature point information group to be positioned and a local feature matching information group set based on the feature point information group to be positioned and a preset database; The above-mentioned local feature matching information set generates a two-dimensional and three-dimensional matching pair group; the fourth generation unit is configured to generate initial pose information based on the above-mentioned two-dimensional and three-dimensional matching pair group; the fifth generation unit is configured to generate a first local point cloud data group based on the above-mentioned initial pose information and a pre-stored point cloud data group; the sixth generation unit is configured to generate a second local point cloud data group based on the above-mentioned image depth information and pre-stored camera intrinsic parameter information; the seventh generation unit is configured to generate final pose information based on the above-mentioned first local point cloud data group, the above-mentioned second local point cloud data group and the above-mentioned initial pose information.

[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.

[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.

[0012] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the three-dimensional pose generation method combined with depth information of some embodiments of the present disclosure, the quality of three-dimensional pose generation is improved, and the time consumption for optimizing to the correct three-dimensional pose is reduced. Specifically, the reason for the long time consumption for optimizing to the correct three-dimensional pose is that the pure visual positioning algorithm only relies on visual RGB information, resulting in that the feature points in the scene are not very uniform, and the positioning accuracy is high for areas with rich textures, but the positioning accuracy is poor in areas without textures or reflections, resulting in poor quality of three-dimensional pose generation, and the depth map positioning algorithm uses the ICP point cloud registration algorithm to gradually iterate and optimize to the correct pose, resulting in a long time consumption. Based on this, the three-dimensional pose generation method combined with depth information of some embodiments of the present disclosure first receives the image to be positioned, wherein the above-mentioned image to be positioned corresponds to image depth information. Thus, an image with depth information can be obtained. Then, based on the above-mentioned image to be positioned and the preset database, a global feature vector to be positioned and a candidate image information group are generated. Thus, each candidate image information similar to the image to be positioned and the global feature vector of the image to be positioned can be obtained from the preset database. Then, based on the above-mentioned image to be located and the above-mentioned candidate image information group, a feature point information group to be located and a local feature matching information group set are generated. Thus, the image to be located can be subjected to feature point extraction and matching with each image in the candidate image information group, and the feature point information group to be located and the local feature matching information group set are obtained. Secondly, based on the above-mentioned feature point information group to be located and the above-mentioned local feature matching information group set, a two-dimensional three-dimensional matching pair group is generated. Thus, the three-dimensional feature information corresponding to the feature point information to be located can be obtained, and a two-dimensional three-dimensional matching pair group is obtained. Then, based on the above-mentioned two-dimensional three-dimensional matching pair group, initial posture information is generated. Thus, the initial posture information can be obtained. Then, based on the above-mentioned initial posture information and the pre-stored point cloud data group, a first local point cloud data group is generated. Thus, the first local point cloud data group corresponding to the above-mentioned initial posture can be obtained. Secondly, based on the above-mentioned image depth information and the pre-stored camera intrinsic parameter information, a second local point cloud data group is generated. Thus, the second local point cloud data group can be obtained according to the image depth information and the camera intrinsic parameter information. Finally, based on the first local point cloud data set, the second local point cloud data set and the initial pose information, the final pose information is generated. Thus, the first local point cloud data set can be matched with the second local point cloud data set to obtain the conversion relationship between the point clouds, thereby generating the final pose information. Also, because the two local point cloud data are generated using the initial pose information and the image depth information for matching processing, the uniformity of the feature points in the scene is improved, so the regional positioning accuracy is improved, and the quality of the three-dimensional pose generation is improved. Also, because the final pose information is generated using the conversion relationship between the two point cloud data and the initial pose information, the time consumed in iterative optimization to the correct pose is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0014] Figure 1 is a flowchart of some embodiments of the three-dimensional pose generation method according to the present disclosure;

[0015] Figure 2 is a schematic structural diagram of some embodiments of a three-dimensional pose generation device according to the present disclosure;

[0016] Figure 3 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0018] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0023] Figure 1 A process 100 of some embodiments of a method for generating a 3D pose in combination with depth information according to the present disclosure is shown. The method for generating a 3D pose in combination with depth information comprises the following steps:

[0024] Step 101: receiving an image to be positioned.

[0025] In some embodiments, the execution subject (e.g., a computing device) of the three-dimensional pose generation method may receive an image to be positioned. The image to be positioned may be a picture sent by a target device. The target device may be, but is not limited to, VR glasses, AR glasses, MR glasses, and a robot camera. The image to be positioned may correspond to image depth information. The image depth information may be depth information corresponding to each pixel in the image. The execution subject may be a server or computing device for storing and managing the image to be positioned.

[0026] Step 102: Generate a global feature vector to be located and a candidate image information group based on the image to be located and a preset database.

[0027] In some embodiments, the execution subject may generate a global feature vector to be located and a candidate image information group based on the image to be located and the preset database. The preset database may be a database for storing and managing all the mapping image data and feature vectors. The above-mentioned all the mapping image data and feature vectors may be image data and feature vectors pre-photographed by the target device to scan and extract the three-dimensional environment. The preset database stores various image feature vectors. The three-dimensional environment may be an environment in a real scene. Each of the above-mentioned image feature vectors corresponds to image information. The above-mentioned image information may be image data pre-photographed by the target device to scan and extract the three-dimensional environment. The above-mentioned image information includes an image and various three-dimensional coordinate information. Each of the three-dimensional coordinate information may be the coordinate of each pixel in the above-mentioned image in the corresponding three-dimensional environment. The above-mentioned image feature vector may be a feature vector pre-photographed by the target device to scan and extract the three-dimensional environment. The above-mentioned global feature vector to be located may be a feature vector used to characterize the global image to be located. Each candidate image information in the candidate image information group may be each image feature vector and each image information similar to the image to be located in the preset database.

[0028] In some optional implementations of some embodiments, the execution subject may generate a global feature vector to be located and a candidate image information group based on the image to be located and a preset database through the following steps:

[0029] The first step is to perform feature extraction on the above-mentioned image to be located to obtain a global feature vector to be located. In practice, the above-mentioned execution subject can perform feature extraction on the above-mentioned image to be located through a feature extraction algorithm to obtain a global feature vector to be located. Among them, the above-mentioned feature extraction algorithm can be an algorithm for extracting feature vectors from the image to be located. For example, the above-mentioned feature extraction algorithm can be but is not limited to: BoW (Bag of Visual Words, visual vocabulary bag model), VLAD (Vector of Locally Aggregated Descriptors, local aggregated descriptor vector), NetVLAD (Network of VLAD, networked VLAD).

[0030] In the second step, the similarity between each image feature vector included in the preset database and the global feature vector to be located is determined as each feature similarity. In practice, the execution entity may determine the similarity between the global feature vector to be located and each image feature vector as each feature similarity. Each of the feature similarities may be a distance between two image vectors. The feature similarity may be calculated by, but is not limited to, cosine similarity.

[0031] In the third step, in response to determining that each of the above-mentioned feature similarities satisfies a preset similarity condition, each feature similarity that satisfies the preset similarity condition is determined as each target similarity. The above-mentioned preset similarity condition may be that each of the above-mentioned feature similarities belongs to a preset number of similarities in a feature similarity sequence. The above-mentioned feature similarity sequence may be a sequence obtained by sorting the above-mentioned feature similarities from small to large. The above-mentioned preset number of similarities may be a pre-set value, which is not limited here. In practice, the above-mentioned execution entity may determine each of the above-mentioned feature similarities that satisfies the preset similarity condition as each target similarity.

[0032] In the fourth step, each image feature vector and each image information corresponding to each target similarity is used as each candidate image information to determine the candidate image information group. In practice, for each target similarity in each target similarity, first, the execution subject may combine the image feature vector and the image information corresponding to the target similarity into candidate image information. Then, the execution subject may determine each candidate image information as a candidate image information group.

[0033] Step 103: Based on the image to be located and the candidate image information group, generate a feature point information group to be located and a local feature matching information group set.

[0034] In some embodiments, the execution subject may generate a feature point information group to be located and a local feature matching information group set based on the image to be located and the candidate image information group. Each feature point information to be located in the feature point information group to be located may be a feature point and a descriptor used to characterize the image information to be located. Each local feature matching information in the local feature matching information group set may be a feature point used to characterize the feature point corresponding to the feature point information to be located in the candidate image information.

[0035] In some optional implementations of some embodiments, the execution subject may generate a feature point information group to be located and a local feature matching information group set based on the image to be located and the candidate image information group through the following steps:

[0036] The first step is to extract feature points from the above-mentioned image to be located to obtain an information group of feature points to be located. Among them, each feature point information to be located in the above-mentioned feature point information group to be located may include a feature point, a feature point coordinate and a descriptor of the corresponding feature point. The above-mentioned feature point coordinates may be coordinate information used to characterize the above-mentioned feature point in the above-mentioned image to be located. In practice, the above-mentioned execution subject may input the above-mentioned image to be located into a feature point extraction network to obtain an information group of feature points to be located. Among them, the above-mentioned feature point extraction network may be a neural network that takes an image as input and each feature point and each descriptor of the corresponding image as output. For example, the above-mentioned feature point extraction network may be a SuperPoint network. In practice, the above-mentioned execution subject may also use a feature point extraction algorithm to extract features from the above-mentioned image to be located to obtain an information group of feature points to be located. Among them, the above-mentioned feature point extraction algorithm may be an algorithm for performing feature detection on an image to obtain feature points and description points corresponding to the description feature points. For example, the above-mentioned feature point extraction algorithm may be a SIFT (Scale-Invariant Feature Transform) algorithm.

[0037] The second step is to extract feature points from each candidate image information in the candidate image information group to obtain each candidate feature point information group. Among them, each candidate feature point information in each candidate feature point information group in the candidate feature point information group can be a feature point and a descriptor used to characterize the corresponding candidate image information. The candidate feature point information corresponds to three-dimensional coordinate information. In practice, for each candidate image information in the candidate image information group, the execution entity can input the candidate image information into the feature point extraction network to obtain a candidate feature point information group corresponding to the candidate image information.

[0038] Step 3: for each candidate image information in the above candidate image information group, perform the following steps:

[0039] In the first sub-step, the candidate feature point information group corresponding to the candidate image information is determined as the target candidate feature point information group.

[0040] The second sub-step is to input the above-mentioned feature point information group to be located and the above-mentioned target candidate feature point information group into the feature detection network to obtain a local feature matching information group. Among them, each local feature matching information in the above-mentioned local feature matching information group includes the feature point information to be located and the candidate feature point information. The above-mentioned feature detection network can be a neural network for matching corresponding feature points in two images. The above-mentioned feature detection network can be a neural network that takes the feature points and descriptors of the two images as input and takes the matching results as output. The matching process of the above-mentioned feature detection network can be to determine the matching pair by comparing the descriptors of the feature points in the two images, and calculate the similarity between the descriptors, for example, by calculating the cosine similarity between the descriptors. The matching strategy of the above-mentioned feature detection network can be to find the best match for each feature point, ensuring that each feature point has only one corresponding matching point in the other image. For example, the above-mentioned feature detection network can be a Superglue network or a network using the NN (Nearest Neighbor, nearest neighbor matching algorithm) algorithm. In practice, the above-mentioned execution subject can input the above-mentioned feature point information group to be located and the above-mentioned target candidate feature point information group into the above-mentioned feature detection network to obtain a local feature matching information group. The fourth step is to determine the obtained local feature matching information groups as a local feature matching information group set.

[0041] Step 104: Generate two-dimensional and three-dimensional matching pairs based on the feature point information group to be located and the local feature matching information group set.

[0042] In some embodiments, the execution subject may generate a two-dimensional and three-dimensional matching pair group based on the feature point information group to be located and the local feature matching information group set. Each two-dimensional and three-dimensional matching pair in the two-dimensional and three-dimensional matching pair group may be data information used to characterize the matching relationship between the feature point information to be located and the local feature matching information. The two-dimensional and three-dimensional matching pair may include the local feature matching information and the three-dimensional coordinate information.

[0043] In some optional implementations of some embodiments, the execution subject may generate a two-dimensional or three-dimensional matching pair based on the feature point information group to be located and the local feature matching information group set by the following steps:

[0044] In the first step, for each feature point information to be located in the above feature point information group to be located, perform the following steps:

[0045] In the first sub-step, each local feature matching information corresponding to the feature point information to be located in the local feature matching information group set is determined as each feature point pair to be matched. In practice, for each local feature matching information group in the local feature matching information group set, the execution subject may determine the local feature matching information corresponding to the feature point information to be located in the local feature matching information group as a feature point pair to be matched.

[0046] In the second sub-step, the to-be-matched feature point pairs that meet the preset conditions among the above-mentioned to-be-matched feature point pairs are determined as target feature point pairs. The preset conditions may be the to-be-matched feature point pairs corresponding to the largest number of candidate feature point information among the candidate feature point information included in the above-mentioned to-be-matched feature points.

[0047] The third sub-step is to add the three-dimensional coordinate information corresponding to the candidate feature point information included in the target feature point pair to the target feature point pair to obtain a two-dimensional and three-dimensional matching pair. In practice, the execution entity may add the three-dimensional coordinate information corresponding to the candidate feature point information included in the target feature point pair and the feature point coordinates included in the feature point information to be located included in the target feature point pair to the target feature point pair to update the target feature point pair. Then, the execution entity may determine the updated target feature point pair as a two-dimensional and three-dimensional matching pair.

[0048] In the second step, each of the obtained two-dimensional and three-dimensional matching pairs is determined as a two-dimensional and three-dimensional matching pair group.

[0049] Step 105: Generate initial pose information based on the two-dimensional and three-dimensional matching pairs.

[0050] In some embodiments, the execution entity may generate initial pose information based on the two-dimensional and three-dimensional matching pairs. The initial pose information may be used to characterize the pose of the target device in three-dimensional space. In practice, first, the execution entity may obtain the internal parameter information of the target device. Then, the execution entity may obtain the initial pose information through a pose solving algorithm based on the two-dimensional and three-dimensional matching pairs and the internal parameter information. The internal parameter information may be the internal parameter of the target device. The pose solving algorithm may be a PnP (Perspective-n-Points) algorithm.

[0051] Step 106 : Generate a first local point cloud data set based on the initial pose information and the pre-stored point cloud data set.

[0052] In some embodiments, the execution subject may generate a first partial point cloud data set based on the initial posture information and the pre-stored point cloud data set. Each point cloud data in the point cloud data set may be point cloud data in a world coordinate system for representing the three-dimensional environment scanned and extracted by the target device in advance. Each first partial point cloud data in the first partial point cloud data set may be a coordinate in a camera coordinate system for representing each pixel in the image to be positioned. The camera coordinate system may be a pre-set coordinate system, which is not limited here.

[0053] In some optional implementations of some embodiments, the execution subject may generate a first local point cloud data set based on the initial pose information and a pre-stored point cloud data set by the following steps:

[0054] The first step is to generate a camera point cloud data group based on the above initial pose information and the above point cloud data group. Among them, each camera point cloud data in the above camera point cloud data group can be converted into point cloud data in the camera coordinate system. For example, the above camera point cloud data can be [x', y', z', 1]. The "1" in [x', y', z', 1] is a filling number and is not limited here. In practice, for each point cloud data in the above point cloud data group, the above execution entity can input the above point cloud data and the above initial pose information into the first preset formula to obtain the camera point cloud data. Among them, the above first preset formula can be P′=T init ·P. P′ is the camera point cloud data. init is the initial pose information. P is the point cloud data.

[0055] In the second step, each camera point cloud data satisfying the preset camera retention condition in the camera point cloud data group is determined as each point cloud data to be retained. The preset camera retention condition may be that the Z-axis coordinate in the camera point cloud data is greater than or equal to a preset value. The preset value may be a pre-set value, which is not limited here. For example, the preset threshold may be "0".

[0056] In the third step, the above-mentioned point cloud data to be retained are projected to obtain the pixel plane coordinates. Among them, the above-mentioned pixel plane coordinates can be the coordinates used to characterize the above-mentioned point cloud data to be retained on the preset projection plane. The above-mentioned preset projection plane can be a pre-set plane, which is not limited here. For example, the above-mentioned pixel plane coordinates can be (u, v). In practice, for each of the above-mentioned point cloud data to be retained, the above-mentioned execution entity can input the above-mentioned internal reference information and the above-mentioned point cloud data to be retained into the second preset formula to obtain the pixel plane coordinates. Among them, the above-mentioned second preset formula can be

[0057] u and v are pixel plane coordinates. X, Y, and Z are the coordinates of the point cloud data to be retained. f x is the focal length representing the camera in the x-axis direction in the intrinsic parameter information. f y is the focal length representing the camera in the y-axis direction in the intrinsic parameter information. c x is the optical center coordinate representing the camera in the x-axis direction in the intrinsic parameter information. c y is the optical center coordinate representing the camera in the y-axis direction in the intrinsic parameter information.

[0058] Fourth step, determine each pixel plane coordinate that meets the preset pixel retention condition among the above-mentioned pixel plane coordinates as each pixel plane coordinate to be retained. Among them, the above-mentioned preset pixel retention condition can be a preset range, which is not limited here. For example, the above-mentioned preset range can be "0 ≤ u < width and 0 ≤ v < height".

[0059] Fifth step, classify the above-mentioned pixel plane coordinates to be retained to obtain each group of pixel plane coordinates. In practice, for each pixel plane coordinate to be retained among the above-mentioned pixel plane coordinates to be retained, the execution subject can determine the corresponding point cloud data to be retained of the above-mentioned pixel plane coordinate to be retained as a group of pixel plane coordinates.

[0060] Sixth step, for each group of pixel plane coordinates among the groups of pixel plane coordinates, determine the camera point cloud data corresponding to the pixel plane coordinates to be retained that meet the preset mapping condition in the above-mentioned group of pixel plane coordinates as the first local point cloud data. Among them, the above-mentioned preset mapping condition can be the pixel plane coordinate with the smallest Z-axis coordinate in the group of pixel plane coordinates.

[0061] Seventh step, determine the determined first local point cloud data as a group of first local point cloud data.

[0062] Step 107, generate a group of second local point cloud data based on the image depth information and the pre-stored camera intrinsic parameter information.

[0063] In some embodiments, the execution subject can generate a group of second local point cloud data based on the above-mentioned image depth information and the pre-stored camera intrinsic parameter information. Among them, the above-mentioned camera intrinsic parameter information can be the intrinsic parameters of the target device. Each second local point cloud data in the above-mentioned group of second local point cloud data can be used to represent the coordinates of each pixel in the image to be located in the camera coordinate system.

[0064] In some optional implementation manners of some embodiments, the execution subject can generate a group of second local point cloud data based on the above-mentioned image depth information and the pre-stored camera intrinsic parameter information through the following steps:

[0065] The first step is to generate second local point cloud data for each pixel of the image to be located corresponding to the image depth information based on the camera intrinsic parameter information and the pixel. In practice, the execution subject can input the pixel coordinates and depth information corresponding to the pixel and the camera intrinsic parameter information into a third preset formula to obtain the second local point cloud data. The third preset formula can be in, is the second local point cloud data. (u, v) is the pixel coordinate. z=Depth(u, v) is the depth information of the pixel coordinate (u, v). x f is the focal length of the camera in the x-axis direction in the camera intrinsic information. y c is the focal length of the camera in the y-axis direction in the camera intrinsic information. x c is the optical center coordinate of the camera in the x-axis direction in the camera intrinsic information. y is the optical center coordinate of the camera in the y-axis direction in the camera intrinsic parameter information. The camera intrinsic parameter information may be the intrinsic parameter information.

[0066] In the second step, each of the generated second partial point cloud data is determined as a second partial point cloud data group.

[0067] Step 108 : generating final pose information based on the first partial point cloud data set, the second partial point cloud data set and the initial pose information.

[0068] In some embodiments, the execution subject may generate final pose information based on the first partial point cloud data set, the second partial point cloud data set and the initial pose information, wherein the final pose information may be used to characterize the pose of the target device in three-dimensional space.

[0069] In the process of adopting technical solutions to solve the above technical problems, the following problems are often accompanied:

[0070] Only the depth map positioning algorithm is used to gradually and iteratively optimize the pose information to the correct pose. Because a large number of pictures containing depth information are used, the positioning complexity is high, resulting in a long time to optimize to the correct pose.

[0071] Faced with the above technical problems, we decided to adopt the following solutions:

[0072] In some optional implementations of some embodiments, the execution subject may generate final pose information based on the first partial point cloud data set, the second partial point cloud data set and the initial pose information through the following steps:

[0073] The first step is to generate initial point cloud conversion relationship information based on the first local point cloud data group and the second local point cloud data group. The initial point cloud conversion relationship can be used to characterize the conversion relationship between two point cloud data. The initial point cloud conversion relationship can include a rotation matrix and a translation matrix. In practice, the execution entity can obtain the initial point cloud conversion relationship between the first local point cloud data group and the second local point cloud data group through a point cloud alignment algorithm. The point cloud alignment algorithm can be an ICP (Iterative Closest Point) point cloud alignment algorithm.

[0074] In the second step, based on the above initial point cloud transformation relationship information, the following loop steps are performed:

[0075] The first sub-step is to generate point cloud error information based on the first partial point cloud data set, the second partial point cloud data set and the initial point cloud conversion relationship. The point cloud error information can be a numerical value used to characterize the error between the two point clouds. In practice, the execution subject can input the first partial point cloud data set, the second partial point cloud data set and the initial point cloud conversion relationship into the fourth preset formula to obtain the point cloud error information. The fourth preset formula can be E is the point cloud error information. R is the rotation matrix in the initial point cloud transformation relationship. t is the translation matrix in the initial point cloud transformation relationship. i is the i-th first local point cloud data in the first local point cloud data group. closest is the distance p in the second local point cloud data set i The second most recent local point cloud data. |||| 2 The symbol is the square of the Euclidean norm.

[0076] In a second sub-step, in response to determining that the point cloud error information satisfies a preset error condition, the initial point cloud transformation relationship is determined as point cloud transformation relationship information. The preset error condition may be that the point cloud error information is less than a preset error threshold. The preset error threshold may be a pre-set value, which is not limited here.

[0077] The third sub-step, in response to determining that the above-mentioned point cloud error information does not meet the above-mentioned preset error condition, adjusts the parameters to regenerate the initial point cloud transformation relationship information to update the initial point cloud transformation relationship information, and uses the updated initial point cloud transformation relationship information to re-execute the above-mentioned loop steps.

[0078] The third step is to generate the final pose information based on the point cloud conversion relationship information and the initial pose information. In practice, the execution subject can input the point cloud conversion relationship information and the initial pose information into the fifth preset formula to obtain the final pose information. The fifth preset formula can be T = T icp ·T init . T is the final pose information. T icp Transform relation information for point cloud. init is the initial pose information.

[0079] The above technical solution and its related contents, as an inventive point of an embodiment of the present disclosure, solve the problem of "only using the depth map positioning algorithm to gradually iteratively optimize the pose information to the correct pose, because a large number of pictures containing depth information are used, the positioning complexity is high, resulting in a long time to optimize to the correct pose". The factors that lead to a long time to optimize to the correct pose are often as follows: only using the depth map positioning algorithm to calculate a large number of pictures containing depth information, the positioning complexity is high. If the above factors are solved, the effect of reducing the time to optimize to the correct pose can be achieved. In order to achieve this effect, the present disclosure first generates initial point cloud conversion relationship information based on the above first local point cloud data group and the above second local point cloud data group. Thus, the initial point cloud conversion relationship information between the two local point cloud data can be obtained. Then, based on the above initial point cloud conversion relationship information, the following loop steps are performed: based on the above first local point cloud data group, the above second local point cloud data group and the above initial point cloud conversion relationship, point cloud error information is generated. Thus, the error information between the two point cloud data groups can be obtained. Then, in response to determining that the above point cloud error information satisfies the preset error condition, the above initial point cloud conversion relationship is determined as the point cloud conversion relationship information. Secondly, in response to determining that the above point cloud error information does not satisfy the above preset error condition, the parameters are adjusted to regenerate the initial point cloud conversion relationship information to update the initial point cloud conversion relationship information, and the above loop steps are re-executed using the updated initial point cloud conversion relationship information. In this way, the initial point cloud conversion relationship can be optimized and adjusted. Finally, based on the above point cloud conversion relationship information and the above initial pose information, the final pose information is generated. In this way, the final pose information can be obtained. Also, because the initial point cloud conversion relationship is obtained by matching the first local point cloud data with the second local point cloud data, the point cloud conversion relationship can be obtained more quickly, so that the final pose information can be obtained, thereby reducing the time consumed in optimizing to the correct pose.

[0080] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the three-dimensional pose generation method combined with depth information of some embodiments of the present disclosure, the time consumption for optimizing to the correct three-dimensional pose is reduced. Specifically, the reason for the long time consumption for optimizing to the correct three-dimensional pose is that the pure visual positioning algorithm only relies on visual RGB information, resulting in that the feature points in the scene are not very uniform, and the positioning accuracy is high for areas with rich textures, but the positioning accuracy is poor in areas without textures or reflections, resulting in poor quality of three-dimensional pose generation, and the depth map positioning algorithm uses the ICP point cloud registration algorithm to gradually iterate and optimize to the correct pose, resulting in a long time consumption. Based on this, the three-dimensional pose generation method combined with depth information of some embodiments of the present disclosure first receives the image to be positioned, wherein the above-mentioned image to be positioned corresponds to image depth information. Thus, an image with depth information can be obtained. Then, based on the above-mentioned image to be positioned and the preset database, a global feature vector to be positioned and a candidate image information group are generated. Thus, each candidate image information similar to the image to be positioned and the global feature vector of the image to be positioned can be obtained from the preset database. Then, based on the above-mentioned image to be located and the above-mentioned candidate image information group, a feature point information group to be located and a local feature matching information group set are generated. Thus, the image to be located can be subjected to feature point extraction and matching with each image in the candidate image information group, and the feature point information group to be located and the local feature matching information group set are obtained. Secondly, based on the above-mentioned feature point information group to be located and the above-mentioned local feature matching information group set, a two-dimensional three-dimensional matching pair group is generated. Thus, the three-dimensional feature information corresponding to the feature point information to be located can be obtained, and a two-dimensional three-dimensional matching pair group is obtained. Then, based on the above-mentioned two-dimensional three-dimensional matching pair group, initial posture information is generated. Thus, the initial posture information can be obtained. Then, based on the above-mentioned initial posture information and the pre-stored point cloud data group, a first local point cloud data group is generated. Thus, the first local point cloud data group corresponding to the above-mentioned initial posture can be obtained. Secondly, based on the above-mentioned image depth information and the pre-stored camera intrinsic parameter information, a second local point cloud data group is generated. Thus, the second local point cloud data group can be obtained according to the image depth information and the camera intrinsic parameter information. Finally, based on the first local point cloud data set, the second local point cloud data set and the initial pose information, the final pose information is generated. Thus, the first local point cloud data set can be matched with the second local point cloud data set to obtain the conversion relationship between the point clouds, thereby generating the final pose information. Also, because the two local point cloud data are generated using the initial pose information and the image depth information for matching processing, the uniformity of the feature points in the scene is improved, so the regional positioning accuracy is improved, and the quality of the three-dimensional pose generation is improved. Also, because the final pose information is generated using the conversion relationship between the two point cloud data and the initial pose information, the time consumed in iterative optimization to the correct pose is reduced.

[0081] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a three-dimensional pose generation device combined with depth information, and these device embodiments are Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0082] like Figure 2 As shown, in some embodiments, the three-dimensional pose generation device 200 combined with depth information includes: a receiving unit 201, a first generating unit 202, a second generating unit 203, a third generating unit 204, a fourth generating unit 205, a fifth generating unit 206, a sixth generating unit 207 and a seventh generating unit 208. The receiving unit 201 is configured to receive an image to be positioned, wherein the image to be positioned corresponds to image depth information; the first generating unit 202 is configured to generate a global feature vector to be positioned and a candidate image information group based on the image to be positioned and a preset database; the second generating unit 203 is configured to generate a feature point information group to be positioned and a local feature matching information group set based on the image to be positioned and the candidate image information group; the third generating unit 204 is configured to generate a two-dimensional three-dimensional pose generation device 201 based on the feature point information group to be positioned and the local feature matching information group set. dimensional matching pair group; the fourth generating unit 205 is configured to generate initial pose information based on the above two-dimensional and three-dimensional matching pair group; the fifth generating unit 206 is configured to generate a first local point cloud data group based on the above initial pose information and a pre-stored point cloud data group; the sixth generating unit 207 is configured to generate a second local point cloud data group based on the above image depth information and pre-stored camera intrinsic parameter information; the seventh generating unit 208 is configured to generate final pose information based on the above first local point cloud data group, the above second local point cloud data group and the above initial pose information.

[0083] It is understood that the units described in the device 200 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 200 and the units included therein, and will not be described in detail here.

[0084] Reference below Figure 3 , which shows a structural schematic diagram of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0085] like Figure 3As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0086] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0087] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0088] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0089] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0090] The computer-readable medium may be included in the electronic device; or it may exist independently without being assembled into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: receives an image to be located, wherein the image to be located corresponds to image depth information; generates a global feature vector to be located and a candidate image information group based on the image to be located and a preset database; generates a feature point information group to be located and a local feature matching information group set based on the image to be located and the candidate image information group; generates a two-dimensional and three-dimensional matching pair group based on the feature point information group to be located and the local feature matching information group set; generates initial pose information based on the two-dimensional and three-dimensional matching pair group; generates a first local point cloud data group based on the initial pose information and a pre-stored point cloud data group; generates a second local point cloud data group based on the image depth information and the pre-stored camera intrinsic parameter information; generates final pose information based on the first local point cloud data group, the second local point cloud data group and the initial pose information.

[0091] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0092] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0093] The units described in some embodiments of the present disclosure may be implemented by software or hardware. The units described may also be set in a processor, for example, it may be described as: a processor includes a receiving unit, a first generating unit, a second generating unit, a third generating unit, a fourth generating unit, a fifth generating unit, a sixth generating unit, and a seventh generating unit. The names of these units do not constitute a limitation on the units themselves in some cases, for example, the receiving unit may also be described as a "unit for receiving an image to be located".

[0094] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0095] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. A method for generating a three-dimensional pose in combination with depth information, comprising: Receiving an image to be positioned, wherein the image to be positioned corresponds to image depth information; Based on the image to be located and a preset database, generating a global feature vector to be located and a candidate image information group; Based on the image to be located and the candidate image information group, generating a feature point information group to be located and a local feature matching information group set; Generate a two-dimensional and three-dimensional matching pair group based on the feature point information group to be located and the local feature matching information group set; Based on the two-dimensional and three-dimensional matching pairs, generating initial pose information; Based on the initial pose information and the pre-stored point cloud data set, generating a first local point cloud data set; Generate a second local point cloud data set based on the image depth information and pre-stored camera intrinsic parameter information; Final pose information is generated based on the first local point cloud data set, the second local point cloud data set and the initial pose information.

2. The method according to claim 1, wherein: The preset database stores various picture feature vectors, each of which corresponds to picture information; And the generating of the global feature vector to be located and the candidate image information group based on the image to be located and the preset database includes: Extracting features from the image to be located to obtain a global feature vector to be located; Determine the similarity between each picture feature vector included in the preset database and the global feature vector to be located as each feature similarity; In response to determining that the respective feature similarities satisfy a preset similarity condition, determining the respective feature similarities satisfying the preset similarity condition as respective target similarities; Each picture feature vector and each picture information corresponding to each target similarity is used as each candidate image information to determine a candidate image information group.

3. The method according to claim 1, wherein: The step of generating a feature point information group to be located and a local feature matching information group set based on the image to be located and the candidate image information group comprises: Extracting feature points from the image to be located to obtain an information group of feature points to be located; Extracting feature points from each candidate image information in the candidate image information group to obtain each candidate feature point information group; For each candidate image information in the candidate image information group, the following steps are performed: Determining the candidate feature point information group corresponding to the candidate image information as the target candidate feature point information group; Inputting the to-be-located feature point information group and the target candidate feature point information group into a feature detection network to obtain a local feature matching information group, wherein each local feature matching information in the local feature matching information group includes the to-be-located feature point information and the candidate feature point information; The obtained local feature matching information groups are determined as a local feature matching information group set.

4. The method according to claim 3, wherein: Each candidate feature point information in each candidate feature point information group in each candidate feature point information group corresponds to three-dimensional coordinate information; And the generating of two-dimensional and three-dimensional matching pairs based on the feature point information group to be located and the local feature matching information group set comprises: For each feature point information to be located in the feature point information group to be located, perform the following steps: Determine each local feature matching information corresponding to the feature point information to be located in the local feature matching information group as each feature point pair to be matched; Determine the to-be-matched feature point pairs that meet the preset conditions among the to-be-matched feature point pairs as target feature point pairs; Adding the three-dimensional coordinate information corresponding to the candidate feature point information included in the target feature point pair to the target feature point pair to obtain a two-dimensional and three-dimensional matching pair; The obtained two-dimensional and three-dimensional matching pairs are determined as a two-dimensional and three-dimensional matching pair group.

5. The method according to claim 1, wherein: The step of generating a first local point cloud data set based on the initial pose information and a pre-stored point cloud data set comprises: Based on the initial pose information and the point cloud data group, generating a camera point cloud data group; Determine each camera point cloud data satisfying a preset camera retention condition in the camera point cloud data group as each point cloud data to be retained; Projecting each of the point cloud data to be retained to obtain the plane coordinates of each pixel; Determining each pixel plane coordinate that satisfies a preset pixel retention condition among the pixel plane coordinates as each pixel plane coordinate to be retained; Classifying and processing the pixel plane coordinates to be retained to obtain pixel plane coordinate groups; For each pixel plane coordinate group in each pixel plane coordinate group, determining the camera point cloud data corresponding to the pixel plane coordinates to be retained that meet the preset mapping condition in the pixel plane coordinate group as the first local point cloud data; The determined first partial point cloud data are determined as a first partial point cloud data set.

6. The method according to claim 1, wherein: The generating a second local point cloud data set based on the image depth information and pre-stored camera intrinsic parameter information includes: For each pixel of the image to be located corresponding to the image depth information, generating second local point cloud data based on the camera intrinsic parameter information and the pixel; The generated pieces of second partial point cloud data are determined as a second partial point cloud data set.

7. A spatial point cloud positioning device combined with depth information, comprising: A receiving unit is configured to receive an image to be positioned, wherein the image to be positioned corresponds to image depth information; A first generating unit is configured to generate a global feature vector to be located and a candidate image information group based on the image to be located and a preset database; A second generating unit is configured to generate a feature point information group to be located and a local feature matching information group set based on the image to be located and the candidate image information group; A third generating unit is configured to generate a two-dimensional and three-dimensional matching pair group based on the feature point information group to be located and the local feature matching information group set; a fourth generating unit, configured to generate initial pose information based on the two-dimensional and three-dimensional matching pairs; a fifth generating unit, configured to generate a first local point cloud data set based on the initial pose information and a pre-stored point cloud data set; a sixth generating unit, configured to generate a second local point cloud data set based on the image depth information and pre-stored camera intrinsic parameter information; The seventh generating unit is configured to generate final pose information based on the first partial point cloud data set, the second partial point cloud data set and the initial pose information.

8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • A method and device for determining pose of visual device

    CN111325796A

  • Visual positioning method and device, chip system and storage medium

    CN114359392A