A pose determination method and device capable of eliminating the influence of interfering objects

By identifying and eliminating interferers in query images and reference images in visual positioning technology, and extracting features using semantic segmentation and attention mechanisms, the problem of interference between moving objects and textureless objects in the prior art is solved, and the accuracy of positioning is improved.

CN114708176BActive Publication Date: 2025-06-24SUN YAT SEN UNIVERSITY SHENZHEN +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210329265.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-06-24
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

Existing visual positioning techniques are susceptible to interference when dealing with moving objects and textureless objects, resulting in a decrease in positioning accuracy.

Method used

A method is proposed to reduce the impact of interfering objects on positioning by identifying and eliminating interfering objects in query images and reference images, and extracting features using semantic segmentation, bias conversion and attention mechanisms.

Benefits of technology

Effectively eliminate interferences in the image, improve positioning accuracy, and reduce the possibility of mislocalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708176B_ABST
    Figure CN114708176B_ABST
Patent Text Reader

Abstract

The present invention discloses a pose determination method, device, electronic device and computer-readable storage medium capable of eliminating the influence of interfering objects. The method includes: after obtaining a query image, searching for a plurality of reference images corresponding to the query image in a preset image library; forming an image set by combining the query image with each of the reference images, and respectively performing interference object elimination processing on the images in each image set to obtain the relative pose corresponding to each image set; and determining the target pose by combining a plurality of the relative poses. The present invention can, after screening the reference images corresponding to the query image, respectively identify the interfering objects in the screened query image and the reference image, thereby eliminating the interfering objects in the image, reducing the influence of the interfering objects on positioning, and improving the positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Visual Localization is a technology for providing location services and a key task in the field of computer vision. Its purpose is to determine the absolute pose of the camera that captured an image in the world coordinate system, that is, the position and orientation (abbreviation: pose), given only a query image. With the development and application of intelligent technologies, visual localization plays a key role in fields such as intelligent robots, autonomous driving, and augmented reality (AR).

[0003] One commonly used pose determination technology is the indirect localization technology. Its positioning method is to first use image retrieval technology to find reference images similar to the current image in an image library, then determine the reference poses of the reference images and the relative poses between the current image and the reference images respectively, and finally perform epipolar constraint estimation based on the reference poses and relative poses or use deep learning methods for relative pose estimation.

[0004] However, the commonly used pose determination methods currently have the following technical problems: Since the collected environmental images may correspond to various different scenes, and people or objects in different scenes may be in a moving state, if people or objects in a moving state are used as references for positions, it will cause serious interference to the positioning, resulting in positioning errors; moreover, the collected environmental images may also contain various textureless objects (for example, the sky or clouds), and textureless objects have no specific shape, increasing the number of non-positionable references, thus increasing the interference factors for positioning and reducing the positioning accuracy. Summary of the Invention

[0005] The present invention proposes a pose determination method and device that can eliminate the influence of interfering objects. The method can distinguish and eliminate interfering objects in the query image and the reference image to reduce the influence of interfering objects on the positioning process, thereby improving the positioning accuracy.

[0006] The first aspect of the embodiments of the present invention provides a pose determination method that can eliminate the influence of interfering objects. The method includes:

[0007] After obtaining a query image, find several reference images corresponding to the query image in a preset image library;

[0008] Form an image set by combining the query image with each reference image, and perform interference object elimination processing on the images in each image set respectively to obtain the relative pose corresponding to each image set;

[0009] Determine the target pose by combining several of the relative poses.

[0010] In a possible implementation of the first aspect, the process of respectively performing interference removal processing on the images in each of the image sets to obtain the relative pose corresponding to each of the image sets includes:

[0011] Performing semantic segmentation and bias conversion on the images in the image set respectively to obtain a fused image, where the fused image includes a fused query image corresponding to the query image and a fused reference image corresponding to the reference image;

[0012] Using an attention mechanism to extract attention features from the fused image, where the attention features include a query feature image and a reference feature image;

[0013] Associating the query feature image and the reference feature image to obtain the relative pose.

[0014] In a possible implementation of the first aspect, the process of performing semantic segmentation and bias conversion on the images in the image set respectively to obtain a fused image includes:

[0015] Performing semantic segmentation processing on the query image and the reference image respectively to obtain a segmented query image and a segmented reference image;

[0016] Invoking a preset bias network to respectively convert the segmented query image and the segmented reference image to generate a biased query image and a biased reference image;

[0017] Performing element - level corresponding addition of the biased query image and the query image to obtain a fused query image, and performing element - level corresponding addition of the biased reference image and the reference image to obtain a fused reference image.

[0018] In a possible implementation of the first aspect, the process of using an attention mechanism to extract attention features from the fused image includes:

[0019] Inputting the fused query image and the fused reference image into a preset attention feature extraction network respectively to obtain a query feature image and a reference feature image, where the preset attention feature extraction network is composed of a residual network and a CBAM channel - spatial attention module.

[0020] In a possible implementation of the first aspect, the process of associating the query feature image and the reference feature image to obtain the relative pose includes:

[0021] Performing a dot product on the features at each pixel position of the query feature image and the features at each pixel position of the reference feature image to obtain an associated image;

[0022] Inputting the associated image into a preset regression network to obtain the relative pose.

[0023] In a possible implementation of the first aspect, determining the target pose by combining several of the relative poses includes:

[0024] Pairwise combine several of the relative poses to obtain multiple combined poses;

[0025] Perform triangulation on two relative poses within each combined pose to obtain a hypothesized pose;

[0026] Calculate the inlier count value for each hypothesized pose, select the inlier count value with the largest numerical value from multiple inlier count values, and use the hypothesized pose corresponding to the largest inlier count value as the target pose.

[0027] In a possible implementation of the first aspect, searching for several reference images corresponding to the query image in a preset image library includes:

[0028] Use the DenseVLAD algorithm to edit a corresponding query descriptor for the query image, and use the DenseVLAD algorithm to edit a corresponding preset descriptor for each image stored in the preset image library;

[0029] Calculate the Euclidean distance between the query descriptor and each preset descriptor to obtain multiple Euclidean distance values;

[0030] Select several Euclidean distance values less than a preset distance value from multiple Euclidean distance values, and use the images corresponding to the Euclidean distance values less than the preset distance value as reference images.

[0031] A second aspect of the embodiments of the present invention provides a pose determination device capable of eliminating the influence of interfering objects, and the device includes:

[0032] An acquisition and search module, configured to, after acquiring a query image, search for several reference images corresponding to the query image in a preset image library;

[0033] An interference elimination module, configured to form an image set by combining the query image with each reference image, and perform interference object elimination processing on the images within each image set respectively to obtain the relative pose corresponding to each image set;

[0034] A determination module, configured to determine the target pose by combining several of the relative poses.

[0035] Compared with the prior art, a pose determination method and device capable of eliminating the influence of interfering objects provided by an embodiment of the present invention have the following beneficial effects: After screening and querying the reference image corresponding to the image, the interfering objects in the screened query image and the reference image can be respectively identified, so as to eliminate the interfering objects in the image, reduce the influence of the interfering objects on positioning, and improve the accuracy of positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 FIG. 6 is a schematic flowchart of a pose determination method capable of eliminating the influence of interfering objects provided by an embodiment of the present invention;

[0037] Figure 2 FIG. 10 is a schematic structural diagram of a bias network provided by an embodiment of the present invention;

[0038] Figure 3 FIG. 14 is a schematic structural diagram of an attention feature extraction network provided by an embodiment of the present invention;

[0039] Figure 4 FIG. 18 is a schematic diagram of the effect of image association provided by an embodiment of the present invention;

[0040] Figure 5 FIG. 22 is a schematic structural diagram of a regression network provided by an embodiment of the present invention;

[0041] Figure 6 FIG. 26 is a schematic flowchart of steps for generating a relative pose provided by an embodiment of the present invention;

[0042] Figure 7 FIG. 30 is a schematic structural diagram of a pose determination device capable of eliminating the influence of interfering objects provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. It is obvious that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] The currently commonly used pose determination methods have the following technical problems: Since the collected environmental images may correspond to various different scenes, and people or objects in different scenes may be in a moving state, if people or objects in a moving state are used as references for positions, it will cause serious interference to positioning, and thus lead to positioning errors; moreover, the collected environmental images may also contain various different textureless objects (for example, the sky or clouds), and textureless objects have no specific shape, increasing the number of references for which positioning is impossible, thereby increasing the interference factors for positioning and reducing the accuracy of positioning.

[0045] To solve the above problems, the following specific embodiments will be used to introduce and illustrate in detail a pose determination method capable of eliminating the influence of interfering objects provided by the embodiments of the present application.

[0046] Referring to Figure 1 , a flowchart of a pose determination method capable of eliminating the influence of interfering objects provided by an embodiment of the present invention is shown.

[0047] Among them, by way of example, the pose determination method capable of eliminating the influence of interfering objects may include:

[0048] S11. After obtaining the query image, find several reference images corresponding to the query image in the preset image library.

[0049] The query image may be an image that the user needs to perform visual positioning on, and the reference image may be an image similar to the query image.

[0050] The preset image library may be a reference image library generated by combining multiple environmental images collected and labeled with absolute poses by the user. Specifically, the user can collect environmental images at a certain density, and then calibrate the absolute poses of the environmental images through the SfM technology, and combine these environmental images with absolute poses into a reference image library that can be used for visual positioning.

[0051] In order to accurately position, several images can be selected from the multiple images stored in the image library that are labeled with absolute poses by the user as reference images for the query image to perform positioning reference.

[0052] Since there are multiple images in the image library, in order to accurately select the required reference images to improve the positioning accuracy, among them, by way of example, step S11 may include the following sub-steps:

[0053] Sub-step S111: Use the DenseVLAD algorithm to edit a corresponding query descriptor for the query image, and use the DenseVLAD algorithm to edit a corresponding preset descriptor for each image stored in the preset image library.

[0054] DenseVLAD is one of many algorithms for generating image descriptors and is a coding algorithm with minor differences.

[0055] The descriptor can be a compressed representation of the image and can be an image tag. By editing the descriptor corresponding to the image, the similarity between the query image and each image stored in the image library can be measured based on the descriptor, so as to perform image retrieval and improve the efficiency of image screening.

[0056] Sub-step S112: Calculate the Euclidean distance between the query descriptor and each of the preset descriptors to obtain a plurality of Euclidean distance values.

[0057] The Euclidean distance between the query descriptor and each preset descriptor can be calculated, so that a plurality of Euclidean distance values can be obtained.

[0058] Sub-step S113: Screen out several Euclidean distance values less than a preset distance value from the plurality of Euclidean distance values, and use the images corresponding to the Euclidean distance values less than the preset distance value as reference images.

[0059] In one embodiment, if there are multiple Euclidean distance values less than the preset distance value, several Euclidean distance values with the largest numerical values can be selected from the multiple Euclidean distance values according to a preset quantity, and the images corresponding to the Euclidean distance values are used as reference images.

[0060] Optionally, the preset quantity can be 3, 5, 8, 10, etc.

[0061] S12: Combine the query image with each of the reference images to form an image set, and respectively perform interference removal processing on the images in each image set to obtain the relative pose corresponding to each image set.

[0062] In one embodiment, the query image and each reference image can be used as an image set, and interference removal processing is performed on the images in the image set to eliminate each interference in the image, so as to reduce the influence of the interference on the later positioning and improve the positioning accuracy; and at the same time as removing the interference, the relative pose corresponding to this image set can be calculated, so that the final target pose can be determined by using the relative poses corresponding to multiple image sets.

[0063] In order to distinguish different types of objects in the image to eliminate the influence of interfering objects, and in order to screen out image features with significant characteristics from the objects after removing the interference, as an example, step S12 may include the following sub-steps:

[0064] Sub-step S121: Perform semantic segmentation and offset conversion on the images in the image set respectively to obtain a fused image, where the fused image includes a fused query image corresponding to the query image and a fused reference image corresponding to the reference image.

[0065] Among them, integrating semantic segmentation can obtain different semantic information in the image, so that different types of objects included in the query image and the reference image can be distinguished, so that the interfering objects and non-interfering objects can be well distinguished, and then the influence of the interfering objects can be eliminated.

[0066] The bias transformation can incorporate different semantic information in the image into the image, endowing different regions in the image with category attributes, so that the interfering objects and non-interfering objects can be quickly and intuitively distinguished.

[0067] In an optional embodiment, sub-step S121 may include the following sub-steps:

[0068] Sub-step S1211: Perform semantic segmentation processing on the query image and the reference image respectively to obtain a segmented query image and a segmented reference image.

[0069] Specifically, semantic segmentation processing can be performed on the query image and the reference image simultaneously, differentiating the different semantic information corresponding to different objects in the query image and the reference image, so as to obtain a segmented query image and a segmented reference image.

[0070] Sub-step S1212: Call a preset bias network to respectively convert the segmented query image and the segmented reference image to generate a biased query image and a biased reference image.

[0071] The segmented query image and the segmented reference image can be respectively input into the preset bias network, and the preset bias network encodes the input images, so as to obtain bias feature maps with the same size as the input images, and respectively obtain a biased query image and a biased reference image.

[0072] Refer to Figure 2 , which shows a schematic structural diagram of a bias network provided by an embodiment of the present invention.

[0073] In an embodiment, after using the bias network to incorporate different semantic information in the image into the image, different regions in the image have category attributes, and the content part in the image corresponding to the category of the interfering object will be suppressed, and different interfering objects are suppressed to different degrees. After biasing the pixel values of the image in the regions where these interfering objects are located, it is equivalent to removing these interfering objects from the image. Reducing the impact of interfering objects on subsequent processing.

[0074] In addition, semantic information can also be annotated for other content that is not an interfering object, which can improve the accuracy of matching and association in subsequent processing and contribute to improving the accuracy of positioning.

[0075] Sub-step S1213: Element-wise add the biased query image and the query image to obtain a fused query image, and element-wise add the biased reference image and the reference image to obtain a fused reference image.

[0076] After obtaining the biased processed image, the biased processed image can be superimposed and fused with its original image to obtain the corresponding fused image.

[0077] Specifically, the offset query image and the query image can be added element - by - element to obtain a fused query image. Similarly, the offset reference image and the reference image can be added element - by - element to fuse the reference image.

[0078] Sub - step S122: Use the attention mechanism to extract attention features from the fused image, where the attention features include a query feature image and a reference feature image.

[0079] In one embodiment, the screening method using the attention mechanism can screen out more significant image features from the image and filter out other non - significant image features. In this way, it is equivalent to only using significant features for localization, filtering out most non - significant features, and then reducing the possibility of being affected by interfering objects in the image during localization.

[0080] In an alternative embodiment, sub - step S122 may include the following sub - steps:

[0081] Sub - step S1221: Input the fused query image and the fused reference image into a preset attention feature extraction network respectively to obtain a query feature image and a reference feature image. Among them, the preset attention feature extraction network is composed of a residual network and a CBAM channel - spatial attention module.

[0082] Refer to Figure 3 , which shows the structural schematic diagram of the attention feature extraction network provided by an embodiment of the present invention.

[0083] In one embodiment, the attention feature extraction network may be composed of a residual network ResNet34 and a CBAM channel - spatial attention module, and its structure is as Figure 3 shown.

[0084] Among them, the CBAM channel - spatial attention module includes a channel attention module and a spatial attention module.

[0085] In actual operation, the fused query image and the fused reference image can be input into the attention feature extraction network respectively. After passing through the attention feature extraction network, a query feature image and a reference feature image can be output respectively.

[0086] Sub - step S123: Correlate the query feature image and the reference feature image to obtain the relative pose.

[0087] After obtaining the query feature image and the reference feature image, a correlation regression operation can be performed on the query feature image and the reference feature image to obtain the phase pose between the query image and the reference image.

[0088] In an optional embodiment, sub-step S123 may include the following sub-steps:

[0089] Sub-step S1231: Perform a dot product on the features at each pixel position of the query feature image and the features at each pixel position of the reference feature image to obtain an association image.

[0090] Refer to Figure 4 , which shows a schematic diagram of the effect of image association provided by an embodiment of the present invention.

[0091] During the association operation, the dot product of each feature of the query feature image with respect to each feature of the reference feature image can be calculated, so as to realize the association of the query feature image and the reference feature image, and then an association result feature map is generated based on the calculated dot product result to obtain an association image. The specific figure is shown in Figure 4.

[0092] Sub-step S1232: Input the association image into a preset regression network to obtain a relative pose.

[0093] Refer to Figure 5 , which shows a schematic diagram of the structure of the regression network provided by an embodiment of the present invention.

[0094] In one embodiment, the regression network consists of two convolutional layers and one fully connected layer.

[0095] Specifically, after inputting the association image into the regression network, a phase pose can be output. Among them, the result of the relative pose can be represented by an essential matrix. The essential matrix is a 3×3 matrix and is a representation form of the relative pose. The essential matrix can be decomposed to obtain displacement and rotation. Optionally, other forms can also be used, such as representing displacement with a three-dimensional vector, representing rotation with a four-dimensional quaternion, or representing displacement with a three-dimensional vector and representing rotation with a 3×3 rotation matrix.

[0096] Refer to Figure 6 , which shows a flow chart of the steps for generating a relative pose provided by an embodiment of the present invention.

[0097] In actual operation, when operating on each image set, the query image and the reference image in the image set can be processed simultaneously. Semantic segmentation, offset conversion, image fusion, and feature extraction can be performed on the query image and the reference image simultaneously, and finally, feature association and regression processing are performed on the two features, so as to obtain the relative pose of the query image and the reference image.

[0098] S13: Determine the target pose by combining several of the relative poses.

[0099] After performing the above operations on each image set to obtain the relative pose corresponding to each image set, several relative poses are obtained. Finally, combining several relative poses can determine the final target pose.

[0100] In order to accurately evaluate the final target pose, in one embodiment, step S13 may include the following sub-steps:

[0101] Sub-step S131: Combine the several relative poses in pairs to obtain a plurality of combined poses.

[0102] In one embodiment, assuming there are 5 relative poses, after combining the 5 relative poses in pairs, 10 combined poses can be obtained.

[0103] Sub-step S132: Perform triangulation on the two relative poses within each combined pose to obtain a hypothesized pose.

[0104] The two relative poses within each combined pose can be triangulated, so that a hypothesized pose corresponding to these two relative poses can be obtained.

[0105] The specific calculation process is as follows: As described above, the relative pose is an essential matrix. An essential matrix can be decomposed into two possible relative displacements (i.e., the two possible displacement values are opposite) and two possible relative rotations, a total of 2 times 2, a total of 4 possibilities.

[0106] Multiplying the two possible relative rotations by the absolute rotation of the reference image in the image pair that generates the essential matrix can obtain two absolute rotations. Similarly, two relative poses correspond to two essential matrices, each essential matrix can correspond to two absolute rotations, and two essential matrices can correspond to four absolute rotations. Among these four absolute rotations, the two absolute rotations with the smallest Euclidean distance are selected as the correct absolute rotations. On the basis of determining the absolute rotation, continue to use triangulation to select one of the two relative displacements as the correct relative displacement in the above manner. Finally, an absolute pose hypothesis is formed based on the determined absolute rotation and displacement.

[0107] Sub-step S133: Calculate the inlier count value of each hypothesized pose, and select the inlier count value with the largest value from the multiple inlier count values. The hypothesized pose corresponding to the largest inlier count value is used as the target pose.

[0108] In one embodiment, the inlier count calculation process is as follows: For an absolute pose hypothesis, judge one by one whether five groups of image pairs belong to inliers. Let the image pair be (I k , I q ), I k represents the reference image, and I q represents the query image.

[0109] Use the value of the inverse cosine function to determine whether it is an inlier, and the calculation is as shown in the following formula:

[0110]

[0111] In the above formula, where R represents the rotation matrix, and t represents the displacement.

[0112] In an optional embodiment, the upper limit of α can be α max = 10°. If the calculated α is less than this value, it is regarded as an inlier.

[0113] In addition, before application, each required neural network can be pre-trained. Specifically, the collected environmental images and the poses used can be input into the convolutional neural network for training. During the training process, there is no hypothesis testing step, but it ends after obtaining the relative pose (represented by the essential matrix) through regression. The loss function is:

[0114]

[0115] That is, the mean square error loss between the predicted essential matrix and the true essential matrix.

[0116] In this embodiment, the embodiment of the present invention provides a pose determination method that can eliminate the influence of interfering objects. Its beneficial effect is that: after screening the reference images corresponding to the query image, the present invention can respectively identify the interfering objects in the query image and the reference image, so as to eliminate the interfering objects in the image and reduce the influence of the interfering objects on positioning, thereby improving the positioning accuracy.

[0117] The embodiment of the present invention also provides a pose determination device that can eliminate the influence of interfering objects. Refer to Figure 7 , which shows a schematic structural diagram of a pose determination device that can eliminate the influence of interfering objects provided by an embodiment of the present invention.

[0118] Among them, by way of example, the pose determination device that can eliminate the influence of interfering objects may include:

[0119] An acquisition and search module 701, configured to, after acquiring a query image, search for a plurality of reference images corresponding to the query image in a preset image library;

[0120] An interference elimination module 702, configured to form an image set by combining the query image with each of the reference images, and respectively perform interference object elimination processing on the images in each image set to obtain the relative pose corresponding to each image set;

[0121] A determination module 703, configured to determine a target pose by combining a plurality of the relative poses.

[0122] Optionally, the interference elimination module is further configured to:

[0123] Perform semantic segmentation and bias conversion on the images in the image set respectively to obtain fused images, where the fused images include a fused query image corresponding to the query image and a fused reference image corresponding to the reference image;

[0124] Extract attention features from the fused images by using an attention mechanism, where the attention features include a query feature image and a reference feature image;

[0125] Associate the query feature image and the reference feature image to obtain a relative pose.

[0126] Optionally, the interference elimination module is further configured to:

[0127] Perform semantic segmentation processing on the query image and the reference image respectively to obtain a segmented query image and a segmented reference image;

[0128] Call a preset bias network to convert the segmented query image and the segmented reference image respectively to generate a biased query image and a biased reference image;

[0129] Perform element-wise corresponding addition of the biased query image and the query image to obtain a fused query image, and perform element-wise corresponding addition of the biased reference image and the reference image to obtain a fused reference image.

[0130] Optionally, the interference elimination module is further configured to:

[0131] Input the fused query image and the fused reference image into a preset attention feature extraction network respectively to obtain a query feature image and a reference feature image, where the preset attention feature extraction network is composed of a residual network and a CBAM channel spatial attention module.

[0132] Optionally, the interference elimination module is further configured to:

[0133] Perform a dot product on the features at each pixel position of the query feature image and the features at each pixel position of the reference feature image to obtain an associated image;

[0134] Input the associated image into a preset regression network to obtain a relative pose.

[0135] Optionally, the determination module is further configured to:

[0136] Combine the plurality of relative poses pairwise to obtain a plurality of combined poses;

[0137] Triangulate the two relative poses within each of the combined poses to obtain hypothesized poses;

[0138] Calculate the number of inlier values for each of the hypothesized poses, screen the inlier number values with the largest numerical values from the multiple inlier number values, and use the hypothesized pose corresponding to the largest inlier number value as the target pose.

[0139] Optionally, the obtaining and querying module is further configured to:

[0140] Use the DenseVLAD algorithm to edit corresponding query descriptors for the query image, and use the DenseVLAD algorithm to edit corresponding preset descriptors for each image stored in the preset image library;

[0141] Calculate the Euclidean distances between the query descriptors and each of the preset descriptors to obtain multiple Euclidean distance values;

[0142] Screen several Euclidean distance values less than a preset distance value from the multiple Euclidean distance values, and use the images corresponding to the Euclidean distance values less than the preset distance value as reference images.

[0143] Those skilled in the art can clearly understand that for the convenience of description and simplicity, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein.

[0144] Furthermore, an embodiment of the present application further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements a pose determination method capable of eliminating the influence of interference objects as described in the above embodiments.

[0145] Furthermore, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute a pose determination method capable of eliminating the influence of interference objects as described in the above embodiments.

[0146] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A pose determination method capable of eliminating the influence of interfering objects, characterized in that, The method includes: After obtaining the query image, finding several reference images corresponding to the query image in a preset image library; Combining the query image with each of the reference images to form an image set, and respectively performing interference removal processing on the images in each image set to obtain the relative pose corresponding to each image set, where the interference removal processing is to remove the interference of the images in the image set; Determining the target pose by combining several of the relative poses; The step of respectively performing interference removal processing on the images in each image set to obtain the relative pose corresponding to each image set includes: Performing semantic segmentation and bias conversion on the images in the image set respectively to obtain a fused image, where the fused image includes a fused query image corresponding to the query image and a fused reference image corresponding to the reference image; Extracting attention features from the fused image by using an attention mechanism, where the attention features include a query feature image and a reference feature image; Associating the query feature image and the reference feature image to obtain the relative pose.

2. The pose determination method capable of eliminating the influence of interfering objects according to claim 1, characterized in that The step of performing semantic segmentation and bias conversion on the images in the image set respectively to obtain a fused image includes: Performing semantic segmentation processing on the query image and the reference image respectively to obtain a segmented query image and a segmented reference image; Calling a preset bias network to respectively convert the segmented query image and the segmented reference image to generate a biased query image and a biased reference image; Performing element-wise addition of the biased query image and the query image to obtain a fused query image, and performing element-wise addition of the biased reference image and the reference image to obtain a fused reference image.

3. The pose determination method capable of eliminating the influence of interfering objects according to claim 1, characterized in that, The step of extracting attention features from the fused image by using an attention mechanism includes: Inputting the fused query image and the fused reference image into a preset attention feature extraction network respectively to obtain a query feature image and a reference feature image, where the preset attention feature extraction network is composed of a residual network and a CBAM channel-spatial attention module.

4. The pose determination method capable of eliminating the influence of interfering objects according to claim 1, characterized in that The step of associating the query feature image and the reference feature image to obtain the relative pose includes: Performing a dot product on the features at each pixel position of the query feature image and the features at each pixel position of the reference feature image to obtain an associated image; Inputting the associated image into a preset regression network to obtain the relative pose.

5. The pose determination method capable of eliminating the influence of interfering objects according to claim 1, characterized in that The step of determining the target pose by combining several of the relative poses includes: Pairwise combining several of the relative poses to obtain multiple combined poses; Performing triangulation processing on the two relative poses in each combined pose to obtain a hypothesized pose; Calculating the inlier number value of each hypothesized pose, screening the inlier number value with the largest value from multiple inlier number values, and taking the hypothesized pose corresponding to the largest inlier number value as the target pose.

6. The pose determination method capable of eliminating the influence of interfering objects according to any one of claims 1-5, characterized in that The step of finding several reference images corresponding to the query image in a preset image library includes: Edit a corresponding query descriptor for the query image using the DenseVLAD algorithm, and edit a corresponding preset descriptor for each image stored in a preset image library using the DenseVLAD algorithm; Calculate the Euclidean distance between the query descriptor and each of the preset descriptors to obtain a plurality of Euclidean distance values; Select several Euclidean distance values less than a preset distance value from the plurality of Euclidean distance values, and use the images corresponding to the Euclidean distance values less than the preset distance value as reference images.

7. A pose determination device capable of eliminating the influence of interfering objects, characterized in that, The device includes: An acquisition and search module, configured to, after acquiring a query image, search for several reference images corresponding to the query image in a preset image library; An interference elimination module, configured to form an image set by combining the query image and each of the reference images, and perform interference elimination processing on the images in each image set respectively to obtain a relative pose corresponding to each image set, where the interference elimination processing is a process of eliminating interference objects in the images in the image set; A determination module, configured to determine a target pose by combining several of the relative poses; The performing interference elimination processing on the images in each image set respectively to obtain a relative pose corresponding to each image set includes: Performing semantic segmentation and offset conversion on the images in the image set respectively to obtain fused images, where the fused images include a fused query image corresponding to the query image and a fused reference image corresponding to the reference image; Extract attention features from the fused images using an attention mechanism, where the attention features include a query feature image and a reference feature image; Associate the query feature image and the reference feature image to obtain a relative pose.

8. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the pose determination method capable of eliminating the influence of interference objects as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the pose determination method capable of eliminating the influence of interference objects as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Object pose estimation method and device, storage medium and robot

    CN111179342A

  • Method, device and system for determining pose

    CN112784174A