Image positioning method and device, electronic equipment, storage medium and program product
By constructing a 3D Gaussian model and removing occluded objects, the number of visible matching pairs in image localization is increased, thereby improving the pose accuracy of image localization and solving the problem of low pose accuracy in traditional image localization techniques.
Patent Information
- Application Number
- CN202511416910.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-01-20
AI Technical Summary
In traditional image localization techniques, the lack of visible matching pairs results in low pose accuracy of the query image obtained by applying the PnP algorithm.
By constructing a 3D Gaussian model of the target scene, removing occluding objects from candidate keyframe images, and re-rendering the images, the number of visible matching pairs between the image to be located and the keyframe images is increased.
It improves the pose accuracy of the image to be localized and solves the problem of low pose accuracy caused by the lack of visible matching pairs in traditional image localization technology.
Smart Images

Figure CN121366201A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, and particularly relates to an image positioning method and device, an electronic device, a storage medium and a program product, which can be applied to the field of three-dimensional reconstruction. BACKGROUND
[0002] In a traditional image positioning technology, if a camera moves, a large parallax problem occurs between an image captured by the camera and a key frame image saved in a three-dimensional modeling process of a scene, and an object in the image captured by the camera is easily blocked in the key frame. Since the blocked object cannot be reflected in the key frame, the number of visible matching pairs is insufficient, and thus a pose precision of a query image (such as the image captured by the camera) solved by applying a PnP (Perspective-n-Point) algorithm is low. SUMMARY
[0003] Embodiments of the present application provide an image positioning method, device, electronic device, storage medium and program product, which can solve the problem of low pose precision of a query image solved by applying a PnP algorithm due to insufficient number of visible matching pairs in a traditional image positioning technology.
[0004] In a first aspect, embodiments of the present application provide an image positioning method, comprising:
[0005] constructing a three-dimensional Gaussian model of a target scene according to a plurality of key frame images of the target scene;
[0006] obtaining a to-be-positioned image of the target scene, and performing similar frame retrieval on the to-be-positioned image and the plurality of key frame images to obtain a candidate key frame image for positioning;
[0007] deleting one or more occluded objects in the candidate key frame image from the three-dimensional Gaussian model, and re-rendering one or more images under a candidate key frame pose; wherein the one or more occluded objects are deleted in sequence, and one image under the candidate key frame pose is re-rendered each time an occluded object is deleted;
[0008] obtaining a target reference key frame image based on the one or more images, and performing pose calculation on feature point pairs matched between the to-be-positioned image and the target reference key frame image to determine pose information of the to-be-positioned image.
[0009] In a second aspect, embodiments of the present application provide an image positioning device, comprising:
[0010] a model construction module configured to construct a three-dimensional Gaussian model of a target scene according to a plurality of key frame images of the target scene;
[0011] The similar frame retrieval module is configured to acquire a to-be-positioned image of the target scene, and perform similar frame retrieval on the to-be-positioned image and the plurality of key frame images to obtain a candidate key frame image for positioning.
[0012] The occlusion object removal module is configured to remove one or more occlusion objects in the candidate key frame image from the three-dimensional Gaussian model, and re-render one or more images under the candidate key frame pose; wherein the one or more occlusion objects are removed in sequence, and one image under the candidate key frame pose is re-rendered each time one occlusion object is removed.
[0013] The pose calculation module is configured to acquire a target reference key frame image based on the one or more images, and perform pose calculation on feature point pairs matched between the to-be-positioned image and the target reference key frame image to determine pose information of the to-be-positioned image.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising:
[0015] one or more processors;
[0016] The processor is configured to invoke instructions to cause the electronic device to perform the method of the first aspect.
[0017] In a fourth aspect, an embodiment of the present application provides a storage medium, which stores instructions, and when the instructions run on an electronic device, the electronic device performs the method of the first aspect.
[0018] In a fifth aspect, an embodiment of the present application provides a program product, which comprises at least one of a program and instructions, and when the at least one of the program and instructions is executed by an electronic device, the steps of the method of the first aspect are implemented.
[0019] According to the technical solution of the present application, the occlusion objects in the key frame image for positioning are removed by using the 3D Gaussian model, and the occluded objects are exposed, which can increase the number of visible matching pairs between the to-be-positioned image and the key frame image, and solve the problem of low pose accuracy of the query image solved by the PnP algorithm due to the lack of the number of visible matching pairs in the traditional image positioning technology, thereby improving the pose accuracy of the to-be-positioned image.
[0020] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0021] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description, taken in conjunction with the following drawings of which:
[0022] Figure 1 A flowchart of an image positioning method according to an exemplary embodiment is shown;
[0023] Figure 2 A flowchart of an image positioning method according to an exemplary embodiment is shown;
[0024] Figure 3 A flowchart of an image positioning method according to an exemplary embodiment is shown;
[0025] Figure 4 A block diagram of an image positioning apparatus according to an exemplary embodiment is shown;
[0026] Figure 5 A block diagram of an image positioning apparatus according to an exemplary embodiment is shown;
[0027] Figure 6 A block diagram of an apparatus 600 for image positioning according to an exemplary embodiment is shown. DETAILED DESCRIPTION
[0028] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which examples of the embodiments are shown, wherein the same or similar notations are used to denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present application, and are not to be understood as limiting the present application.
[0029] The embodiments of the present application are not exhaustive, and are only schematic of some embodiments, and are not specific limitations on the scope of protection of the present application. Each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be combined arbitrarily, for example, a scheme after removing some steps in a certain embodiment can be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged, in addition, the optional implementation in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be combined arbitrarily, for example, some or all steps of different embodiments can be combined arbitrarily, a certain embodiment can be combined with optional implementations of other embodiments.
[0030] In each embodiment of the present application, the terms and / or descriptions of the embodiments are consistent and can be referred to each other if there is no special description and logical conflict, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0031] It should be noted that the acquisition, transmission, storage, use, processing and the like of data in the technical solutions of the present application comply with relevant provisions of national laws and regulations and do not violate public order and good customs.
[0032] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0033] It should be noted that in the embodiments of the present application, some software, components, models and the like may be mentioned, which are considered to be exemplary, and the purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present application, but it does not mean that the applicant has or will necessarily use the scheme.
[0034] The embodiments of the present application relate to the technical field of computer vision and three-dimensional reconstruction.
[0035] Computer vision is to replace the visual organ as an input sensitive means with various imaging systems, and to replace the brain with a computer to complete processing and interpretation. The ultimate research goal of computer vision is to enable computers to observe and understand the world through vision like people, with the ability to adapt to the environment independently.
[0036] Three-dimensional reconstruction (3D Reconstruction) refers to establishing a mathematical model suitable for computer representation and processing of a three-dimensional object, which is the basis for processing, operating and analyzing its properties in a computer environment, and is also a key technology for establishing a virtual reality in a computer to express the objective world. In computer vision, three-dimensional reconstruction refers to the process of reconstructing three-dimensional information from a single view or multiple views. Since the information of a single view is not complete, three-dimensional reconstruction needs to use experience knowledge. While multi-view three-dimensional reconstruction (similar to human binocular positioning) is relatively easy, the method is to first calibrate the camera, i.e. calculate the relationship between the image coordinate system of the camera and the world coordinate system, and then use the information in multiple two-dimensional images to reconstruct three-dimensional information.
[0037] Structure from Motion (SfM, also called from motion to structure, or structure from motion) is a technology for restoring camera parameters and three-dimensional scene structure by analyzing image sequences. This technology is suitable for the case where the camera moves in a static scene, and the core process includes feature point extraction, image matching, camera pose estimation and three-dimensional point cloud reconstruction. It has wide application in the fields of robot navigation, augmented reality, industrial detection and the like.
[0038] The image positioning method, device, electronic device, storage medium and program product of the embodiments of the present application are described below with reference to the accompanying drawings.
[0039] It should be noted that the execution subject of the image positioning method of the embodiments of the present application can be an image positioning device, which can be implemented in software and / or hardware, and can be configured in an electronic device, which can include but is not limited to a terminal, a server end, etc.
[0040] Figure 1 A flowchart of the image positioning method according to an exemplary embodiment is shown. As shown in the flowchart, the image positioning method can include but is not limited to the following steps. Figure 1
[0041] In step 101, a three-dimensional Gaussian model of a target scene is constructed according to multiple key frame images of the target scene.
[0042] In some embodiments, a sparse point cloud of the target scene can be constructed by using the multiple key frame images of the target scene through the SfM method, and a set of three-dimensional (3D) Gaussian distributions (3D GS) is initialized by using the sparse point cloud obtained by SfM to construct a 3D Gaussian model of the target scene.
[0043] 3D Gaussian Splatting can model the target scene as a set of 3D Gaussian distributions, which is an explicit representation form. Each Gaussian distribution is characterized by a covariance matrix ∑ and a center (mean) point μ, and the mean μ of the 3D Gaussian is usually initialized by the SfM sparse point cloud, and the covariance matrix ∑ can be decomposed into a rotation matrix R and a scaling matrix S for differentiable optimization, where the formula of the 3D Gaussian is as follows:
[0044]
[0045] where the 3D Gaussian has the following parameters: center position spherical harmonic (SH) coefficients representing color
[0046] k represents the degree of freedom; rotation factor (rotation quaternion) scale factor opacity
[0047] The covariance matrix ∑ describes an ellipsoid configured by a scaling matrix S = diag([sx, sy, sz]) and a rotation matrix R = q2R([rw, n, ry, rz]), where diag is a diagonal matrix function, q2R is a function for converting a Quaternion to a Rotation Matrix, rw is a component of the Quaternion representing the scalar part of the Quaternion, and n, ry, rz are components of the Quaternion, which can also be represented as rx elsewhere. The covariance matrix can then be calculated as follows:
[0048]
[0049] For rendering from a given camera view W, which involves the process of scattering Gaussians onto the image plane, this is achieved by approximating the projection of a 3D Gaussian along the depth dimension as a pixel coordinate. Given a viewing transform W (also referred to as a camera pose), the covariance matrix ∑ in camera coordinates 2D can be expressed as:
[0050]
[0051] where J is the Jacobian matrix of the affine approximation of the projective transform. And the final rendered color can be formulated as a mixture of the a of N ordered points that overlap the pixel:
[0052]
[0053] where c i , a i represent the color and opacity of the point, which are derived from the learnable SH color coefficients and per-point opacity, represents the product of the cumulative opacity of the first i-1 Gaussians.
[0054] After the 3D Gaussians are splatted onto the image plane, the color loss between the input color image and the rendered image is calculated, and the Gaussian parameters are optimized based on the color loss. The input color image refers to the actually photographed color image, which provides visual information of the target scene. In step 101, these images are used to construct a sparse point cloud model of the target scene through SfM technology, and to initialize the 3D Gaussian distribution. The rendered image refers to the image rendered by the 3D Gaussian distribution.
[0055] In step 102, the image to be positioned of the target scene is obtained, and the image to be positioned is searched for similar frames with multiple key frame images to obtain candidate key frame images for positioning.
[0056] In the embodiments of the present application, for image positioning of a target scene, a to-be-positioned image of the target scene can be acquired, and the to-be-positioned image is subjected to similar frame retrieval with a plurality of key frame images used in the process of constructing a three-dimensional Gaussian model of the target scene, and the key frame images retrieved similar to the to-be-positioned image are taken as candidate key frame images for positioning. In some embodiments, the to-be-positioned image can refer to an image of the target scene taken by a camera.
[0057] In step 103, one or more occluded objects in the candidate key frame images are deleted from the three-dimensional Gaussian model, and one or more images under the candidate key frame pose are re-rendered; wherein the one or more occluded objects are deleted in turn, and one image under the candidate key frame pose is re-rendered each time an occluded object is deleted.
[0058] In the embodiments of the present application, the occluded objects in the candidate key frame images can be deleted from the 3D Gaussian model, and the images under the perspective are re-rendered, so as to realize the occluded object removal.
[0059] In some embodiments, the occlusion relationship can be determined according to the depth information of each object in the candidate key frame image, and the occluded object removal can be performed based on the 3D Gaussian model in the order of the depth information. Optionally, when there are a plurality of occluded objects, the plurality of occluded objects can be deleted from the three-dimensional Gaussian model in the order of the depth information, and one image under the candidate key frame pose is re-rendered each time an occluded object is deleted, so as to obtain one or more rendered images under the candidate key frame pose. The one or more rendered images can be taken as a reference image set I m ={I0, I1, …, In} for determining a final target reference key frame image, facilitating image positioning. n
[0060] In some embodiments, the occlusion relationship between objects can be determined according to the semantic segmentation result and monocular depth estimation of the candidate key frame image, and the relationship between the mask obtained according to the 2D projection of the 3D Gaussian and the image segmentation model (such as the SAM (Segment Anything Model) segmentation model) is determined to obtain the 3D Gaussian point cloud set corresponding to the object to be removed, and the 3D Gaussian point cloud set corresponding to the object to be removed is deleted from the 3D Gaussian model, so as to realize the occluded object removal. Optionally, when there are a plurality of occluded objects, the plurality of occluded objects can be deleted from the three-dimensional Gaussian model in the order of the depth information, and one image under the candidate key frame pose is re-rendered each time an occluded object is deleted, so as to obtain one or more rendered images under the candidate key frame pose. The image segmentation model can be segmented according to the depth value order of the 2D hint point corresponding to the object to be removed.
[0061] In step 104, a target reference key frame image is acquired based on the one or more images, and pose calculation is performed on the matched feature point pairs between the image to be positioned and the target reference key frame image to determine the pose information of the image to be positioned.
[0062] In some embodiments, one image with the most feature matching points with the image to be positioned can be selected from the one or more images as the target reference key frame image, so as to perform pose calculation on the matched feature point pairs between the image to be positioned and the target reference key frame image. For example, PnP algorithm can be used to solve the camera pose, so as to obtain the pose information of the image to be positioned.
[0063] By implementing the embodiments of the present disclosure, the occluded objects in the key frame image for positioning are removed by using the 3D Gaussian model, and the occluded objects are exposed. The number of visible matching pairs between the image to be positioned and the key frame image can be increased, and the problem of low pose accuracy of the query image solved by the PnP algorithm due to the lack of the number of visible matching pairs in the traditional image positioning technology can be solved, so as to improve the pose accuracy of the image to be positioned.
[0064] Figure 2 A flowchart of an image positioning method according to an exemplary embodiment is shown. In some embodiments, as shown in Figure 2 as shown in Figure 1 The image positioning method can include but is not limited to the following steps.
[0065] In step 201, the candidate key frame image is subjected to semantic segmentation processing to obtain semantic labels and a two-dimensional (2D) set of hint points of each object.
[0066] In some embodiments, the candidate key frame image can be subjected to semantic segmentation processing by using a pre-trained semantic segmentation model to obtain semantic labels and a two-dimensional (2D) set of hint points of each object (e.g., represented by P 2d ={d0,d1,…,d m}. Optionally, the semantic segmentation model can be a Mask R-CNN model, or a DeepLab model, etc., but is not limited thereto, for example, can also be a DANet model. Optionally, the above-mentioned semantic labels can be pixel-level semantic labels. For example, the candidate key frame image can be subjected to semantic segmentation processing to obtain pixel-level semantic labels of each object.
[0067] In step 202, monocular depth estimation is performed on the candidate key frame image to obtain a depth map of the candidate key frame image.
[0068] In some embodiments, monocular depth estimation can be performed on the candidate key frame image by using a monocular depth estimation model to obtain a depth map of the candidate key frame image. The monocular depth estimation model can be a MonoDepth model or a ZoeDepth model, but is not limited thereto.
[0069] In step 203, an area in the candidate key frame image that is not matched with the image to be positioned is selected, and a depth value of the area in the depth map of the candidate key frame image is obtained.
[0070] In the embodiments of the present application, the candidate key frame image can be matched with the image to be positioned, an area in the candidate key frame image that is not matched with the image to be positioned is selected, and a depth value of the area in the depth map of the candidate key frame image is obtained.
[0071] In step 204, one or more occluded objects in the area are sequentially deleted from the three-dimensional Gaussian model according to an order of the depth values of the semantic labels in the area.
[0072] In the embodiments of the present application, the one or more occluded objects are sequentially deleted, and an image under the candidate key frame pose is re-rendered each time an occluded object is deleted.
[0073] It is worth noting that the occluded object to be deleted can be determined based on the depth values of the semantic labels in the area, and one semantic label corresponds to one object. In some embodiments, according to the order of the depth values of the semantic labels in the area, a currently processed semantic label is determined, a two-dimensional hint point in a two-dimensional hint point set corresponding to the currently processed semantic label is mapped to a 3D hint point in the three-dimensional Gaussian model, the 3D hint point is projected to a 2D plane of a corresponding view to obtain a 2D hint point in all views, a three-dimensional Gaussian point cloud set of an occluded object to be removed in the three-dimensional model is determined according to a relationship between the 2D hint point in all views and a mask obtained by the image segmentation model, and the three-dimensional Gaussian point cloud set of the occluded object is deleted from the three-dimensional Gaussian model.
[0074] In some embodiments, as shown in FIG. 3, the optional implementation of sequentially deleting the one or more occluded objects in the area from the three-dimensional Gaussian model according to the order of the depth values of the semantic labels in the area includes but is not limited to the following steps. Figure 3
[0075] In step 301, a two-dimensional hint point set corresponding to a currently processed semantic label is determined according to an order of depth values of semantic labels in the area.
[0076] In some embodiments, the semantic label with the smallest depth value can be determined as the first semantic label to be processed in ascending order of the depth values of the semantic labels in the region. After the corresponding object is deleted based on the 2D hint point set corresponding to the semantic label, the semantic label with the smallest depth value among the remaining unprocessed semantic labels can be determined as the next semantic label to be processed, and so on, until all the occluded objects in the region are removed.
[0077] In step 302, the 2D hint points in the 2D hint point set corresponding to the current semantic label to be processed are mapped to 3D hint points in the 3D Gaussian model, and the 3D hint points to which the 2D hint points corresponding to the current semantic label to be processed are projected into the 2D plane of the corresponding view to obtain the 2D hint points in all views.
[0078] In some embodiments, the i-th 2D hint point in the candidate key frame image can be represented as The corresponding 3D hint point is defined as:
[0079]
[0080] where d(μ) is the depth of the Gaussian center μ, P0 is the projection matrix of the candidate key frame image, and thus P0μ is the position of μ in the view of the candidate key frame image. The above 3D hint point definition formula represents: The corresponding 3D hint point is the center of a certain 3D Gaussian point cloud. In some embodiments, the center can satisfy the following requirements:
[0081] 1) It has a similar projection position as the 2D hint point , where the Manhattan distance is less than a threshold ∈;
[0082] 2) If there are multiple 3D Gaussian point cloud centers that satisfy the above requirement 1), the center of the 3D Gaussian point cloud with the smallest positive depth is selected as the 3D hint point.
[0083] For all 2D hint points in the first view (i.e., the above candidate key frame image), a set of 3D hint points can be obtained by the above method. Then, for the key frame images of other views (i.e., the key frame images other than the candidate key frame image in the above plurality of key frame images used to construct the 3D Gaussian model), the set of 3D hint points can be projected into the 2D plane of the corresponding view to generate 2D hints in the view. In this way, 2D hint points can be obtained in all key frame images, i.e., 2D hint points in all views.
[0084] In step 303, the 2D hint points in all views are segmented according to the image segmentation model to obtain the corresponding masks.
[0085] In the embodiments of the present application, after obtaining the 2D hint points in all views, the 2D hint points in all views can be segmented by using an image segmentation model, for example, a SAM segmentation model can be used to segment the 2D hint points in all views to obtain the corresponding mask (such as m j ).
[0086] In step 304, according to the mask corresponding to the 2D hint points in all views, the 3D Gaussian point cloud set of the occluded object to be removed in the 3D Gaussian model is determined.
[0087] In some embodiments, it can be judged whether the 2D projection point corresponding to the 3D Gaussian point cloud falls within the mask generated by the image segmentation model, and the 3D Gaussian point cloud set within the mask is determined as the 3D Gaussian point cloud set of the occluded object to be removed in the 3D Gaussian model.
[0088] In an optional implementation, for each 3D Gaussian point cloud, whether the 3D Gaussian point cloud is a Gaussian point cloud to be removed can be determined by the following formula:
[0089]
[0090] Wherein, G i represents the 3D Gaussian point cloud to be removed, C i represents the weighted score of all views, τ is a threshold value, G i with a value of 1 represents a Gaussian point cloud to be removed, and G i with a value of 0 represents a Gaussian point cloud not to be removed. Wherein, the formula of C i is as follows:
[0091]
[0092] Wherein, i represents the i-th 3D Gaussian point cloud, j represents the j-th view, and N is the number of multiple key frame images; X ij represents whether the projection point of the i-th 3D Gaussian point cloud when performing 2D projection falls within the mask in the j-th view, for example, the value of X ij may be:
[0093]
[0094] Wherein, μ i is the center of the i-th 3D Gaussian point cloud of the target scene, m j is the mask in the j-th view, P j is the projection matrix of the j-th view.
[0095] Thus, the 3D Gaussian point cloud set G representing the occluded object to be removed can be obtained m = {G0, G1, …, G n}.
[0096] In step 305, the 3D Gaussian point cloud set of the occluded object is deleted from the 3D Gaussian model.
[0097] In an embodiment of the present application, after obtaining the 3D Gaussian point cloud set G representing the occluded object m = {G0, G1, …, G n}, the 3D Gaussian point cloud set can be deleted from the 3D Gaussian model. For each semantic label to be processed, steps 302 to 305 are performed, so that all occluded objects to be removed can be deleted from the 3D Gaussian model. In an embodiment of the present application, after deleting one occluded object from the 3D Gaussian model, the image under the candidate key frame image pose needs to be re-rendered, and the re-rendered image is used as a key frame for positioning, instead of the original key frame, to match the to-be-positioned image for positioning. In this way, one or more re-rendered images under the candidate key frame image pose can be obtained, and these re-rendered images are used as a reference image set, so as to determine a target reference key frame image from the re-rendered image, for feature point matching with the to-be-positioned image and pose calculation, so as to obtain the pose information of the to-be-positioned image.
[0098] In the above embodiment, whether the 2D projection point corresponding to the 3D Gaussian point cloud falls within the mask generated by the image segmentation model can be used to delete the 3D Gaussian point cloud within the mask from the 3D Gaussian model, and the corresponding image is re-rendered, so as to remove the occluded object. The image after removing the occluded object is used to implement PnP algorithm solving, so as to obtain the pose information of the to-be-positioned image. The occluded object can be exposed, the number of visible matching pairs of the query image and the key frame can be increased, and the quality of pose estimation can be improved. This is crucial for autonomous navigation, obstacle avoidance and construction of accurate three-dimensional models in the process of three-dimensional reconstruction.
[0099] Figure 4 A block diagram of an image positioning device according to an exemplary embodiment is shown. As shown in Figure 4 the image positioning device can include a model construction module 410, a similar frame retrieval module 420, an occluded object removal module 430, and a pose calculation module 440.
[0100] The model construction module 410 is configured to construct a three-dimensional Gaussian model of a target scene according to multiple key frame images of the target scene.
[0101] The similar frame searching module 420 is configured to acquire a to-be-positioned image of a target scene, and perform similar frame searching on the to-be-positioned image and a plurality of key frame images to obtain a candidate key frame image used for positioning.
[0102] The occlusion object removing module 430 is configured to remove one or more occlusion objects in the candidate key frame image from the three-dimensional Gaussian model, and re-render one or more images under the candidate key frame pose. The one or more occlusion objects are removed in sequence, and one image under the candidate key frame pose is re-rendered each time one occlusion object is removed.
[0103] The pose calculation module 440 is configured to acquire a target reference key frame image based on the one or more images, and perform pose calculation on feature point pairs matched between the to-be-positioned image and the target reference key frame image to determine pose information of the to-be-positioned image.
[0104] In some embodiments, the model construction module 410 is configured to construct a sparse point cloud of the target scene by using the plurality of key frame images of the target scene through an SfM method, and initialize a set of three-dimensional Gaussian distributions by using the sparse point cloud to construct a three-dimensional Gaussian model of the target scene.
[0105] In some embodiments, the pose calculation module 440 is configured to select an image with the most feature matching points with the to-be-positioned image from the one or more images as the target reference key frame image.
[0106] Optionally, in some embodiments, as shown in FIG. 5B, on the basis of FIG. 5A, the occlusion object removing module 530 can further include a semantic segmentation unit 531, a depth estimation unit 532, an acquisition unit 533, and a deletion unit 534. Figure 5 Figure 4 The semantic segmentation unit 531 is configured to perform semantic segmentation processing on the candidate key frame image to obtain a semantic label and a two-dimensional hint point set of each object. The depth estimation unit 532 is configured to perform monocular depth estimation on the candidate key frame image to obtain a depth map of the candidate key frame image. The acquisition unit 533 is configured to select a region that is not matched between the candidate key frame image and the to-be-positioned image, and acquire a depth value of the region from the depth map of the candidate key frame image. The deletion unit 534 is configured to sequentially remove one or more occlusion objects in the region from the three-dimensional Gaussian model according to an ordering sequence of the depth values of the semantic labels in the region. In this way, the occlusion object removing module 530 can remove the one or more occlusion objects in the region from the three-dimensional Gaussian model according to the ordering sequence of the depth values of the semantic labels in the region, and re-render one or more images under the candidate key frame pose each time one occlusion object is removed. Figure 5 The model construction module 410, the similar frame searching module 420, and the pose calculation module 440 in FIG. 4 have the same functions and structures as the model construction module 410, the similar frame searching module 420, and the pose calculation module 440 in FIG. 5A. Figure 4 The model construction module 410, the similar frame searching module 420, and the pose calculation module 440 in FIG. 4 have the same functions and structures as the model construction module 410, the similar frame searching module 420, and the pose calculation module 440 in FIG. 5A.
[0107] In some embodiments, the deleting unit 534 is configured to: determine a set of two-dimensional hint points corresponding to the current semantic label to be processed according to an order of depth values of the semantic label in the region; map the two-dimensional hint points in the set of two-dimensional hint points corresponding to the current semantic label to be processed to three-dimensional hint points in the three-dimensional Gaussian model, and project the three-dimensional hint points to which the two-dimensional hint points corresponding to the current semantic label to be processed are mapped to a two-dimensional plane of the corresponding view to obtain two-dimensional hint points in all views; perform segmentation processing on the two-dimensional hint points in all views according to the image segmentation model to obtain corresponding masks; determine a set of three-dimensional Gaussian points of the occluded object to be removed in the three-dimensional Gaussian model according to the masks of the two-dimensional hint points in all views; and delete the set of three-dimensional Gaussian points of the occluded object from the three-dimensional Gaussian model.
[0108] In some embodiments, the three-dimensional hint point to which the two-dimensional hint point is mapped is the center of the three-dimensional Gaussian point cloud, wherein the center satisfies the following requirements: 1) has a similar projection position as the two-dimensional hint point, wherein the Manhattan distance is less than a threshold; 2) if there are multiple centers of three-dimensional Gaussian point clouds satisfying the above requirement 1), the center of the three-dimensional Gaussian point cloud with the smallest positive depth is selected as the three-dimensional hint point.
[0109] In some embodiments, the deleting unit 534 is configured to: determine whether the two-dimensional projection point corresponding to the three-dimensional Gaussian point cloud falls within the mask generated by the image segmentation model, and determine the set of three-dimensional Gaussian points within the mask as the set of three-dimensional Gaussian points of the occluded object to be removed in the three-dimensional Gaussian model.
[0110] It should be noted that the foregoing explanations of the embodiments of the image positioning method are also applicable to the image positioning apparatus of the embodiments, which will not be described here again.
[0111] Figure 6 is a block diagram of an apparatus 600 for image positioning according to an exemplary embodiment. For example, the apparatus 600 can be an electronic device, which can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like.
[0112] Referring to Figure 6 , the apparatus 600 can include one or more of the following components: a processing component 602, a memory 604, a power supply component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.
[0113] The processing component 602 generally controls the overall operations of the device 600, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 602 can include one or more processors 620 to execute instructions
[0114] The memory 604 is configured to store various types of data to support operations of the device 600. Examples of such data include instructions for any applications or methods operating on the device 600, contact data, phonebook data, messages, pictures, videos, and so on. The memory 604 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic or optical disk.
[0115] The power component 606 supplies electrical power for the various components of the device 600. The power component 606 can include a power supply management system, one or more power sources, and other components associated with generating, managing and distributing electrical power for the device 600.
[0116] The multimedia component 608 includes a screen providing an output interface between the device 600 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, swiping and gestures on the touch panel. The touch sensors can not only sense a boundary of a touching or swiping action, but also detect duration and pressure related to the touching or swiping action. In some embodiments, the multimedia component 608 includes a front-facing camera and / or a rear-facing camera. The front-facing camera and / or the rear-facing camera can receive external multimedia data when the device 600 is in an operation mode, such as a shooting mode or a video mode. Each of the front-facing camera and the rear-facing camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0117] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC) that is configured to receive an external audio signal when the device 600 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 also includes a speaker for outputting audio signals.
[0118] The I / O interface 612 provides an interface between the processing component 602 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0119] The sensor component 614 includes one or more sensors for providing status assessments of various aspects of the device 600. For example, the sensor component 614 can detect an open / closed position of the device 600, relative positioning of components, such as a display and a keypad of the device 600, a change of position of the device 600 or a component of the device 600, presence or absence of user contact with the device 600, changes in orientation or acceleration / deceleration
[0120] The communication component 616 is configured to facilitate wired or wireless communication between the device 600 and other devices. The device 600 can access a wireless network based on a corresponding communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an example embodiment, the communication component 616 receives broadcast signals or broadcast-related information from external broadcast management systems via a broadcast channel. In an example embodiment, the communication component 616 also includes a Near Field Communication (NFC) module to promote short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) techniques, infrared data association (IrDA) techniques, ultra-wideband (UWB) techniques, Bluetooth (BT) techniques, and other techniques.
[0121] In exemplary embodiments, the apparatus 600 can be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronic devices, to perform the above-described methods.
[0122] In exemplary embodiments, a non-transitory computer readable storage medium including instructions, such as the memory 604 including instructions, is also provided, which can be executed by the processor 620 of the apparatus 600 to complete the above-described methods. For example, the non-transitory computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.
[0123] In exemplary embodiments, a program product including at least one of a program, instructions, which is executed by an electronic device to implement the steps of the above-described methods, is also provided.
[0124] In the description of the specification, the description referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Illustrative expressions of the above terms in the specification do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples, without contradiction.
[0125] Any process or method descriptions or descriptions of the flow diagrams in the specification or elsewhere in this document, can be understood as representing the steps of a method or process that can be implemented in hardware, software, or a combination of both, unless specifically stated otherwise. The preferred embodiments of the present application include additional implementation in which the steps of the method or process are performed in an order different from the order shown or discussed, including substantially simultaneously, or in reverse order, as will be understood by those skilled in the art.
[0126] The logic and / or steps represented in the flowcharts and / or described herein, for example, can be considered as a sequence of instructions to implement logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device, such as a computer-based system, processor- based system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be a computer- readable storage medium or a computer-readable signal medium. The computer- readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include the following: an electrical connection having one or more wires (electrical connections), a portable computer diskette (a magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory.
[0127] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. As such, in some embodiments, specifically configured hardware can be used to implement at least some of the functionality described herein. For example, if implemented in hardware, the hardware can include any or a combination of the following: a discrete logic circuit having logic gates for implementing logic functions upon data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0128] Those of skill in the art would understand that information and signals can be represented using any of a variety of technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that can be referenced throughout the above description can be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0129] In addition, each of the functional units in the various embodiments of the present application can be integrated in one processing module, or each of the units can be physically present separately, or two or more units can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.
[0130] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.
Claims
1. An image positioning method characterized by, The method comprises the following steps: constructing a three-dimensional Gaussian model of a target scene according to multiple key frame images of the target scene; obtaining a to-be-positioned image of the target scene, and performing similar frame retrieval on the to-be-positioned image and the multiple key frame images to obtain a candidate key frame image for positioning; removing one or more occluded objects in the candidate key frame image from the three-dimensional Gaussian model, and re-rendering one or more images under the candidate key frame pose; wherein the one or more occluded objects are removed in sequence, and one image under the candidate key frame pose is re-rendered each time an occluded object is removed; obtaining a target reference key frame image based on the one or more images, and performing pose calculation on feature point pairs matched between the to-be-positioned image and the target reference key frame image to determine pose information of the to-be-positioned image.
2. The method of claim 1, wherein, The method of constructing a three-dimensional Gaussian model of a target scene according to multiple key frame images of the target scene comprises the following steps: constructing a sparse point cloud of the target scene by using a structure from motion (SfM) method according to the multiple key frame images of the target scene; initializing a set of three-dimensional Gaussian distributions by using the sparse point cloud to construct the three-dimensional Gaussian model of the target scene.
3. The method of claim 1, wherein, The method of removing one or more occluded objects in the candidate key frame image from the three-dimensional Gaussian model comprises the following steps: performing semantic segmentation processing on the candidate key frame image to obtain semantic labels and a two-dimensional hint point set of each object; performing monocular depth estimation on the candidate key frame image to obtain a depth map of the candidate key frame image; selecting an area that is not matched between the candidate key frame image and the to-be-positioned image, and obtaining a depth value of the area from the depth map of the candidate key frame image; sequentially removing one or more occluded objects in the area from the three-dimensional Gaussian model according to an ordering sequence of the depth values of the semantic labels in the area.
4. The method of claim 3, wherein, The method of sequentially removing one or more occluded objects in the area from the three-dimensional Gaussian model according to an ordering sequence of the depth values of the semantic labels in the area comprises the following steps: determining a two-dimensional hint point set corresponding to a current to-be-processed semantic label according to the ordering sequence of the depth values of the semantic labels in the area; mapping a two-dimensional hint point in the two-dimensional hint point set corresponding to the current to-be-processed semantic label to a three-dimensional hint point in the three-dimensional Gaussian model, and projecting the three-dimensional hint point mapped by the two-dimensional hint point corresponding to the current to-be-processed semantic label to a two-dimensional plane of a corresponding view to obtain two-dimensional hint points in all views; performing segmentation processing on the two-dimensional hint points in all views according to an image segmentation model to obtain corresponding masks; determining a three-dimensional Gaussian point cloud set of an occluded object to be removed in the three-dimensional Gaussian model according to the corresponding masks of the two-dimensional hint points in all views; removing the three-dimensional Gaussian point cloud set of the occluded object from the three-dimensional Gaussian model.
5. The method of claim 4, wherein, The three-dimensional hint point mapped by the two-dimensional hint point is the center of a three-dimensional Gaussian point cloud, wherein the center satisfies the following requirements: 1) has a projection position similar to a two-dimensional hint point, wherein the Manhattan distance is less than a threshold value; 2) if there are multiple three-dimensional Gaussian point clouds that meet the above requirement 1), the center of the three-dimensional Gaussian point cloud with the smallest positive depth is selected as the three-dimensional hint point.
6. The method of claim 4, wherein, The method further includes determining a set of three-dimensional Gaussian point clouds of occluded objects that need to be removed from the three-dimensional Gaussian model according to the masks corresponding to the two-dimensional hint points in all views, including: determining a set of three-dimensional Gaussian point clouds of occluded objects that need to be removed from the three-dimensional Gaussian model according to the masks corresponding to the two-dimensional hint points in all views, including:
7. The method of any one of claims 1 to 6, wherein, The method further includes acquiring a target reference key frame image based on the one or more images, including: selecting an image with the most matching points with the image features of the image to be positioned as the target reference key frame image from the one or more images.
8. An image positioning apparatus characterized by comprising: The method further includes: a model construction module configured to construct a three-dimensional Gaussian model of a target scene according to a plurality of key frame images of the target scene; a similar frame retrieval module configured to acquire an image to be positioned of the target scene, and perform similar frame retrieval on the image to be positioned and the plurality of key frame images to obtain a candidate key frame image for positioning; an occluded object removal module configured to remove one or more occluded objects in the candidate key frame image from the three-dimensional Gaussian model, and re-render one or more images under the candidate key frame pose; wherein the one or more occluded objects are removed in sequence, and one image under the candidate key frame pose is re-rendered each time an occluded object is removed; a pose calculation module configured to acquire a target reference key frame image based on the one or more images, and perform pose calculation on feature point pairs matched between the image to be positioned and the target reference key frame image to determine pose information of the image to be positioned.
9. An electronic device, comprising: The method further includes: one or more processors; wherein the processor is configured to invoke instructions to cause the electronic device to perform the method of any one of claims 1-7.
10. A storage medium, the storage medium storing instructions, wherein, The instructions, when executed on the electronic device, cause the electronic device to perform the method of any one of claims 1-7.
11. A program product comprising at least one of a program, instructions, characterized in that The program, instructions, or at least one of them, when executed by the electronic device, implement the steps of the method of any one of claims 1-7.