Method, apparatus, and storage medium for object localization under non-continuous observation
By acquiring object models and performing association and alignment in discontinuous observation scenarios, the problem of localization and reconstruction of dynamic objects after observation is interrupted is solved, achieving efficient and accurate object localization and reconstruction, which is suitable for robot operation and grasping tasks.
Patent Information
- Application Number
- CN202111349093.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-15
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2041-11-15
AI Technical Summary
Existing technologies struggle to accurately locate and reconstruct dynamic objects in discontinuous observation scenarios, especially when observations are interrupted or obstructed. They cannot effectively utilize computer-aided design models and continuous observations, resulting in significant changes in object layout and affecting the accuracy of object observation, reconstruction, and location.
By recovering the reference image after the observation is interrupted, the object model is obtained, and the model is associated with the model reconstructed before the observation was interrupted. The geometric and color features of the object are used for matching, and the point cloud registration network is used for alignment, so as to realize the localization and reconstruction of the object in the discontinuous observation scenario.
Under discontinuous observation conditions, it can efficiently and accurately associate and align objects to generate complete point cloud models, supporting robot operation and grasping tasks, improving the accuracy of object reconstruction and pose estimation, without relying on CAD models and continuous observation.
Smart Images

Figure CN113989374B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to computer vision, including object localization in computer vision. BACKGROUND
[0002] Dynamic object reconstruction and localization is a key task in the field of computer vision and robotics, which can be applied to a variety of application scenarios, including autonomous navigation, augmented reality, to robotic grasping and manipulation. In the related art, either a computer-aided design (CAD) model of the object is relied on, or continuous observation is needed to handle the reconstruction and localization of dynamic objects. However, conventional methods tend to ignore the fact that the shapes and sizes of everyday objects vary, and the computer-aided design model can be unknown or not easily available. In practice, the obtained object model can be segmented due to limited viewing angles or inter-object occlusion, and the sensor can not be able to continuously observe multiple objects. During the observation interruption / loss, the layout of the objects can have changed dramatically. This will adversely affect the object observation, reconstruction and localization. SUMMARY
[0003] This summary is provided to introduce a selection of concepts that are further described below in the section. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0004] According to some embodiments of the present disclosure, an object localization method in a discontinuous observation scene is provided, comprising the steps of: obtaining an object model based on a reference image obtained when observation is resumed after observation interruption; and correlating objects before and after observation interruption based on the obtained object model and an object reconstruction model.
[0005] According to some other embodiments of the present disclosure, an object correlation apparatus in a discontinuous observation scene is provided, comprising: a model obtaining unit configured to obtain an object model based on a reference image obtained when observation is resumed after observation interruption; and a correlation unit configured to correlate objects before and after observation interruption based on the obtained object model and an object reconstruction model.
[0006] According to some embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute a method of any of the embodiments described in the present disclosure based on instructions stored in the memory.
[0007] According to some embodiments of the present disclosure, a computer readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, causes a method of any of the embodiments described in the present disclosure to be implemented.
[0008] According to some embodiments of the present disclosure, there is provided a computer program product comprising instructions which, when executed by a processor, cause a method according to any of the embodiments described in the present disclosure to be implemented.
[0009] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0010] The preferred embodiments of the present disclosure will be described herein below with reference to the accompanying drawings. The accompanying drawings are used in order to provide a further understanding about the present disclosure and they together with the detailed description below contain part of the disclosure and form a part of the disclosure for explaining the present disclosure. It should be understood that the accompanying drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the drawings:
[0011] Figure 1 A discontinuous observation scenario is schematically shown.
[0012] Figure 2A And 2B A method of object localization in a discontinuous observation scenario according to embodiments of the present disclosure is shown, Figure 2C A schematic diagram of an object matching process according to embodiments of the present disclosure is shown.
[0013] Figure 3 An example of object association according to embodiments of the present disclosure is shown.
[0014] Figure 4 An apparatus of object localization in a discontinuous observation scenario according to embodiments of the present disclosure is shown.
[0015] Figure 5 A block diagram of some embodiments of an electronic device of the present disclosure is shown.
[0016] Figure 6 A block diagram of some other embodiments of an electronic device of the present disclosure is shown.
[0017] It should be understood that the dimensions of the various parts shown in the drawings are not necessarily to scale as the drawings are for the purpose of illustration only. Wherever possible, the same reference numbers will be used throughout the drawings to depict the same or similar parts. As such, once a part has been defined, it can not be discussed further in subsequent drawings. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present disclosure will be clearly and completely described in combination with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, and not all the embodiments. The description of the embodiments below is actually only illustrative, and is by no means any limitation on the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein.
[0019] It should be understood that each step described in the method embodiments of the present disclosure can be performed in different order and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specified, the relative arrangement of the components and steps, numerical expressions, and numerical values set forth in these embodiments should be interpreted as merely illustrative, not limiting the scope of the present disclosure.
[0020] The term "comprise" and variations thereof used in the present disclosure means an open term that includes at least the recited elements / features, but does not exclude other elements / features, i.e., "comprising but not limited to". In addition, the term "include" and variations thereof used in the present disclosure means an open term that includes at least the recited elements / features, but does not exclude other elements / features, i.e., "including but not limited to". Therefore, include and comprise are synonymous. The term "based on" means "at least partially based on".
[0021] Throughout the specification, the term "one embodiment", "some embodiments", or "embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments". Moreover, the appearance of the phrase "in one embodiment", "in some embodiments", or "in embodiments" at various places in the specification does not necessarily all refer to the same embodiment, but can refer to different embodiments.
[0022] It should be noted that the "first", "second", and the like concepts mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the functions performed by these devices, modules or units in a given order or mutual dependency. Unless otherwise specified, the "first", "second", and the like concepts are not intended to imply a given order or any other manner of given order in time, space, ranking or any other manner.
[0023] It should be noted that the modification of "one", "multiple" mentioned in the present disclosure is illustrative but not restrictive, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0024] The names of the data, messages or information exchanged in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the data, messages or information.
[0025] Dynamic object reconstruction and localization is crucial for robots to understand the surrounding environment and manipulate objects in the environment in robotic operations. On one hand, reconstruction can help to complete the modeling of partially observed or occluded objects. On the other hand, accurate pose estimation can improve the completeness and accuracy of object reconstruction.
[0026] Some current works achieve dynamic object reconstruction by introducing an additional segmentation network module in a simultaneous localization and mapping (SLAM) system to distinguish objects of interest. In these works, either the object class is assumed to be known or continuous observation is required. However, in real robotic operations, these assumptions may not be guaranteed. Current object pose estimation methods mainly rely on known computer-aided design (CAD) models or require a large amount of cost to scan objects to obtain high-quality models. In addition, these methods may need to train new weights for each object or class, which limits the generalizability and is obviously not suitable for real scenarios.
[0027] Moreover, in practice, there are scenarios where discontinuous observations occur due to observation interruption / loss, including but not limited to scene transformation, object occlusion, object entry and exit, etc. which can cause changes in the observation scene. During the observation interruption / loss, the layout of the object can have changed greatly. Figure 1 The scenario of discontinuous observation over time is shown, in which the layout of the object can have changed greatly during the observation interruption / loss. Figure 1 Figs. (a) and (b) show the object layout before and after the observation interruption / loss, respectively. Compared with (a), the object is completely disordered in the observation after the interruption is resumed in (b). This can adversely affect object observation, reconstruction and localization. How to associate objects in discontinuous observations and accurately localize objects in the new scene is a quite challenging problem in this case.
[0028] In view of this, we propose an improved scheme that can perform dynamic object localization without continuous observation and known computer-aided design models.
[0029] In practical application scenarios, such as robot operation and grasping tasks, continuous observation of objects in the scene cannot be guaranteed due to limited view or mutual occlusion between objects. In discontinuous observation, the spatial and motion continuity of objects cannot be guaranteed, but most rigid object models do not change in texture and structure. Therefore, rigid object models are essential for many applications and have become indispensable for association between different observations. Thus, the scheme of the present disclosure obtains the models of objects before and after the observation interruption to locate objects in the discontinuous observation scene. In particular, the scheme of the present disclosure can obtain the object model after the resumption of observation in the case of discontinuous observation, and realize the association between objects before and after the observation interruption based on the obtained object model and the object reconstruction model obtained according to the information before the observation interruption, thereby further accurately locating the objects.
[0030] In addition, after the object association in the scheme of the present disclosure, object pose estimation can be further performed to align the objects to facilitate subsequent processing of the objects. For example, dynamic object reconstruction and pose estimation tasks can be robustly processed without CAD models and continuous observation to generate explicit point cloud models suitable for robot object grasping.
[0031] Embodiments according to the present disclosure will be described in detail below with reference to the accompanying drawings, in particular, object association and alignment in the discontinuous observation scene.
[0032] Figure 2A An object localization method in a discontinuous observation scene according to some embodiments of the present disclosure is shown. In method 200, in step S201, an object model is obtained based on a reference image obtained when the observation is resumed after the observation interruption; and in step S202, the obtained object model is associated with an object reconstruction model to realize the association between objects before and after the observation interruption.
[0033] According to some embodiments of the present disclosure, the object model obtained based on the reference image is a model of the object in the observation scene after the observation recovery / scene change. In some embodiments, the object model obtained based on the reference image can be a model of various appropriate forms, which can contain / indicate / describe various attribute information of the object in the observation scene, including texture, structure, pose, color, etc. In some embodiments, the object model is an object point cloud model, which can be in any appropriate form.
[0034] In some embodiments, the reference image used to generate the object model can be a predetermined number of images observed after the observation recovery. Preferably, to achieve object localization after observation recovery as soon as possible, to correlate the object before and after the observation interruption as soon as possible, a predetermined number of consecutive images after observation recovery can be used to generate the object model, which should be as small as possible in order to achieve object correlation quickly and efficiently. In particular, the reference image is a monocular image obtained at the time of observation recovery, such as the starting image obtained at the time of observation recovery, for example, the first frame image. In some examples, the reference image can be referred to as a query image.
[0035] In some embodiments, a 2.5D instance point cloud can be obtained from the reference image as the object point cloud model, in particular, a 2.5D instance point cloud is obtained from the starting image. As an example, the 2.5D instance point cloud is obtained from the depth image back projection under the guidance of the instance segmentation network. In particular, in embodiments of the present disclosure, due to possible object arrangement, occlusion, etc., the reference image observed after the observation recovery, for example, the monocular image, can not fully reflect all objects in the scene, or even only reflect part of the object, so that the obtained object point cloud model is essentially a point cloud model of an incomplete object or a partial object, for example, it can belong to a partial point cloud observed under monocular observation, which is an incomplete point cloud model. Therefore, in this paper, the object point cloud model obtained from the reference image is also referred to as a "partial object point cloud model" or an "incomplete object point cloud model", which are synonymous in the context of the present disclosure.
[0036] According to some embodiments of the present disclosure, the reconstructed model of the object can refer to the model of the object in the scene before the observation interruption / scene change, which can be used for collaborative processing with the object model obtained after the observation interruption / scene change to achieve the correlation between the objects before and after the observation interruption / scene change. In some embodiments, the reconstructed model of the object is obtained by model reconstruction of the object based on the continuous images of the object. In some embodiments, the continuous images of the object are a predetermined number of consecutive observation images before the observation interruption.
[0037] In some embodiments, the reconstructed model of the object can be predetermined and stored in a suitable storage device. In particular, the model reconstruction is performed as the object is continuously observed and stored in a suitable storage device. For example, the model reconstruction can be performed periodically, e.g. periodically in the continuous observation of the object. In addition, the model reconstruction can be performed continuously, e.g. the model reconstruction is performed every time a predetermined number of images are continuously observed. In this way, the pre-stored reconstructed object model can be directly invoked when the object localization / correlation operation starts in the non-continuous observation scenario. In other embodiments, when the object localization / correlation operation starts in the non-continuous observation scenario, the object reconstruction can be performed first, e.g. based on a predetermined number of continuously observed images before the observation is interrupted, whereby the object localization is performed based on the reconstructed model.
[0038] In some embodiments, the reconstructed model of the object is any suitable type of object model, in particular, can be a surfel-based model, a point cloud model, or other suitable object model. Similarly, it can also contain / indicate / describe various attribute information of the object in the observation scenario, including texture, structure, pose, color, etc. In some embodiments, the continuous images include continuous RGB images and depth images of the object, also known as RGB-D images, and preferably, the model reconstructed based on the continuous images is a surfel-based model. Compared with a single-view point cloud model, the surfel-based model can obtain more comprehensive attribute information of the object, and construct a more complete model, thereby reducing or even eliminating the geometric and density errors of the object.
[0039] In the scheme of the present disclosure, the model reconstruction of the object can be performed in various ways. As an example, the object-level model construction given an RGB-D image as input can be realized by introducing a learning-based instance segmentation method in SLAM to obtain the reconstructed model of the object, e.g. SLAM++, Fusion++, Co-Fusion and MaskFusion, MID-Fusion, etc. in the related art. In some embodiments, preferably, based on MaskFusion, a surfel is used to represent the object model when constructing, so that the surfel object model obtained can more accurately and comprehensively reflect the object features than the point cloud model, making the scheme of the present disclosure more efficient when operating than when applying the point cloud model.
[0040] According to some embodiments of the present disclosure, the association between objects before and after the observation interruption can refer to achieving a match between the objects before and after the observation interruption, i.e., finding a correspondence, especially a one-to-one correspondence, between the objects before and after the observation interruption, so as to be able to associate the objects before the observation interruption with the objects re-observed after the interruption is resumed, to facilitate subsequent operations. In different tasks, object association across different frames is studied, usually in multiple object tracking (MOT). MOT focuses on tracking dynamic objects across consecutive frames. Most MOT methods rely on continuity assumptions (such as GIoU or Bayesian filtering) to perform data association, but fail when observations are discontinuous. In the embodiments of the present disclosure, the object model before and after the observation interruption / loss in the discontinuous observation scenario is utilized to achieve association, so as to efficiently achieve object association.
[0041] In some embodiments, in step S202, the association between objects before and after the observation interruption based on the obtained object model and the object reconstruction model further comprises: in step S2021, determining the similarity between the object point cloud model and the object reconstruction model based on the object information, and in step S2022, associating between objects before and after the observation interruption based on the similarity. As shown in Figure 2B
[0042] In some embodiments, the object information can be various attribute information characterizing the object, for example, various attribute information that can be obtained from the object observation result, which can include at least one of object geometric features, object texture features, object color features, for example. As an example, the object information can be extracted from the observed object image, the object model obtained from the object image, etc.
[0043] In some embodiments, the object information includes both geometric features and color features, and the similarity between the object point cloud model and the object reconstruction model is determined based on both geometric features and color features. In particular, in visual perception, similar objects tend to be similar in structure and texture. Since some objects have similar shapes or textures, it is difficult to distinguish them only by geometric information or only by color information. Therefore, the present disclosure proposes to utilize both geometric features and color features of the object as object information to determine model similarity. As an example, both geometric features and color features of the object can be extracted from the colored point cloud of the object obtained from the object observation image.
[0044] In some embodiments, the associating objects before and after the observation interruption based on the similarity further comprises determining a one-to-one correspondence between the objects before and after the observation interruption based on the similarity so as to associate the objects before and after the observation interruption. In some embodiments, the one-to-one correspondence between the objects before and after the observation interruption is determined based on a maximum sum similarity between the acquired object model and the object reconstruction model. Both the acquired object model and the object reconstruction model can contain parameter information of multiple objects, and the object parameters corresponding to the maximum similarity / maximum matching condition can indicate the correspondence between the objects. As an example, various appropriate algorithms can be employed to determine the maximum matching condition between the acquired object model and the object reconstruction model to determine the correspondence between the objects.
[0045] Figure 2C An object matching process diagram according to an embodiment of the present disclosure is shown. First, from the object reconstruction model before the observation interruption and the acquired object model after the observation recovery, their mixed features composed of geometry and color respectively are extracted. Then the similarity of the two set models is estimated based on these features. Finally, we use a suitable algorithm, a matching algorithm such as Sinkhorn algorithm, to find a one-to-one correspondence between the two sets with the maximum total similarity.
[0046] In some embodiments, the method 200 further comprises a step S203 of aligning the associated objects before and after the observation interruption. After the model and 2.5D instance point cloud association, we obtain the rough position information of the object model in the new scene, that is, the respective position information of each observed object in the scene after the observation recovery. However, the pose of the object relative to the camera can have changed greatly. Therefore, it is necessary to align the object with the new scene in order to better estimate the object pose. The purpose of object pose estimation is to estimate the orientation and transformation of the object, which is crucial for robot operation. Object pose estimation can be implemented in various appropriate ways, for example, a framework LatentFusion for 6D pose estimation of invisible objects, which proposes to solve the 6D pose estimation of invisible objects by reconstructing the latent 3D representation of the object using sparse reference views. Of course, object pose estimation can also be implemented in other ways, which will not be described in detail here.
[0047] In some embodiments, after obtaining the object association before and after the observation interruption, the poses between the associated / corresponding objects before and after the observation interruption are aligned. The object alignment can be achieved by various appropriate methods. In some embodiments, the poses of the objects after the observation interruption are aligned with the objects before the observation interruption by a spatial transformation. As an example, the object alignment can be achieved by aligning the obtained object model, such as an object point cloud model, with the object reconstruction model, such as an object mesh model, by a certain transformation. The object alignment can be achieved by various appropriate methods / algorithms, which will not be described in detail here. Thus, by the object association and alignment according to the present disclosure, the object point cloud model obtained from a single or a small number of reference images can be further corrected based on the object reconstruction model after the observation recovery, so that the object model obtained after the observation recovery can be more complete and better reflect the observed object condition, for example, compared with a single view, the scheme according to the present disclosure can better obtain the model of the object that is partially or completely occluded due to scene transformation, observation interruption, etc.
[0048] It should be pointed out that the alignment operation is not essential for the scheme of the present disclosure, that is, even if the alignment operation is not performed, the scheme of the present disclosure can still accurately and efficiently determine the association between the objects before and after the observation interruption by using the object model, to efficiently realize the object positioning in the discontinuous observation scene.
[0049] The following will describe an example of object positioning in a discontinuous observation scene according to an embodiment of the present disclosure in combination with Figure 3
[0050] The present disclosure aims to reconstruct and position dynamic objects without known computer-aided design models and continuous observation, and proposes to use object models, in particular object reconstruction models and object models obtained after observation interruption recovery, to solve this problem. Figure 3 The three components of the exemplary example of the present disclosure are schematically shown: object model reconstruction, object association, and object alignment. In the object model reconstruction, continuous frames can be input, and the object model is reconstructed based on the SLAM system, in the object association, object-level data association between discontinuous frames, i.e., previous observation frames and new observation frames, is performed; in the object alignment, the object model is aligned with the new observation value by a point cloud registration network. The implementation of each component in the scheme of the present disclosure will be described in detail below.
[0051] For object model reconstruction, in the present example, a video segment V t is input to reconstruct the object model M m , m = 0...N. Specifically, to reconstruct the object model, we use the implementation of MaskFusion to achieve camera tracking and object-level construction during the video clip. MaskFusion represents the object model in surfels and performs camera tracking by aligning the given RGB-D image with the projection of the reconstructed model. To achieve the reconstruction of each object, it uses Mask (Mask) R-CNN to obtain instance masks and fuse each instance mask into the object-level construction. In addition, to better cope with object reconstruction, the present disclosure further trains a class-agnostic segmentation network that can be combined with MaskFusion to enable dynamic object reconstruction of a wider class of objects.
[0052] When previous observations are lost and new observations come, we need to find the correspondence between the reconstructed objects and the new scene. However, directly aligning the object model to the new scene is time-consuming and can lead to ambiguous matching. Therefore, a coarse-to-fine multi-object alignment process is proposed. In the coarse matching, we introduce a correlation module to estimate the similarity between the object model and the 2.5D instance point cloud in the reference image (query image), and then find the matching between them. These matches provide the approximate location of each object in the new scene. Under the guidance of the instance segmentation network, the 2.5D instance point cloud is obtained by back-projection from the depth image.
[0053] The correlation operation of the present example takes two sets of color point clouds as input: the object reconstruction model M m and the 2.5D instance point cloud P n extracted from the reference image by back-projection, and extracts respective geometric and color features from both.
[0054] 1) Geometric features: The geometric feature extraction part is implemented through a PointNet++ network. Specifically, we use the function which processes unordered point clouds and encodes them into fixed-length vectors, where N = 1024.
[0055] 2) Color features: Analyzing color distribution, the statistical histogram helps to distinguish different objects. Specifically, a three-dimensional histogram of size (32, 32, 32) is used to calculate the RGB distribution. The three channels of the histogram represent red, green, and blue in the same way as the RGB image. The (256, 256, 256) color space in the image can be scaled to a smaller (32, 32, 32) color space to improve efficiency and robustness to lighting changes in different scenes.
[0056]
[0057]
[0058] where i, j, k can represent the three-channel element coordinates in the three-dimensional histogram, and the corresponding histogram element value is x when the object model has x R, G, B color elements.
[0059] Feature extraction can be performed using various appropriate methods, such as various methods known in the art, which will not be described in detail here. As an example, a multi-scale grouping classification version of PointNet++ can be used in the feature extractor.
[0060] Then, based on the obtained geometric features and color features, one-to-one matching is found between the m object models M m and the 2.5D instance point cloud P n in the reference image. We formulate the one-to-one matching problem as an optimal transport problem and determine the one-to-one matching by solving the problem.
[0061] Specifically, the similarity between the two sets is evaluated using the weighted sum S geo and S rgb . Let be two different point clouds, for example, corresponding to the object reconstruction model and the object model obtained after observation recovery, respectively.
[0062]
[0063] where φ is an L2 normalization function. β is a flattening function that can convert a three-dimensional histogram of color features into a vector. λ can be any appropriate value, such as 0.1.
[0064] The goal of the one-to-one matching problem is to find the corresponding relationship between the two sets with the maximum total similarity. Specifically, the maximum weight matching between the m object models and the n 2.5D instance point clouds needs to be found. In this example, it is formulated as an optimal transport theory model. In addition, n and m can not be equal when some objects disappear or new objects appear. To handle these cases, a relaxation variable is introduced in the formula to find objects without a corresponding relationship. Taking the case where m < n as an example, the n x n distance matrix D is defined as follows:
[0065]
[0066] The transport matrix is T, where T ij is the matching probability of M i and P j . The matching problem can be formulated as:
[0067]
[0068]
[0069]
[0070] Finally, Sinkhorn algorithm is used to solve (7) to get one-to-one correspondence. When T ij >0.5 and M i and P j are not the relaxation variables, it can be considered that (i, j) is a good match, otherwise this match will be abandoned.
[0071] After the model and the 2.5D instance point cloud are associated, the rough position information of the object model in the new scene is obtained. However, considering that the pose of the object relative to the camera may have changed greatly, it is necessary to align the object with the new scene in order to obtain the accurate 6-DOF (degree of freedom) pose. In the present disclosure, the object is represented by a patch in the reconstructed model, and the patch is similar to the point cloud in geometry and color, so the object pose can be aligned with the new scene by point cloud registration. Point cloud registration refers to the problem of finding a rigid transformation to align two given point clouds, such as the reconstructed model and the object model after observation recovery. Formally, given two point clouds X and Y, the goal is to find a transformation T e SE3 that aligns the two point clouds. In implementation, RPMNet, a point cloud registration network, can be used for each pre-matched point cloud set, which achieves the best performance in partial, noisy and invisible point cloud registration tasks.
[0072] According to embodiments of the present disclosure, further processing can be performed before the point cloud is input into the network to optimize the alignment process. As an example, filtering is performed using a filter with a radius of 0.005 and at least 16 neighboring points. Then the point cloud is down-sampled with a voxel size of 0.01. Finally, the point cloud is scaled to the unit sphere and translated to the origin. Before the reference model is passed into RPMNet, we generate several hypotheses for each axis with an initial angle of 90 degrees, 180 degrees, and 270 degrees, which further reduces the sensitivity to the initial angle between the two point clouds. These operations can be performed based on the Open3D library, of course, any other appropriate library can also be used.
[0073] Figure 4 An object association apparatus in a discontinuous observation scene according to some embodiments of the present disclosure is shown. In the apparatus 400, a model acquisition unit 401 is configured to acquire an object model based on reference images obtained when observation is resumed after observation interruption; and an association unit 402 is configured to associate the object based on the acquired object model and a reconstructed model of the object to realize object association before and after observation interruption.
[0074] In some embodiments, the association unit 402 further comprises a similarity determining unit 4021 configured to determine a similarity between the object point cloud model and the object reconstruction model based on the object information, and the association unit further associates the objects before and after the observation interruption based on the similarity.
[0075] In some embodiments, the object information comprises both geometric features and color features, and the similarity determining unit is configured to determine the similarity between the object point cloud model and the object reconstruction model based on both the geometric features and the color features.
[0076] In some embodiments, the association unit 402 is further configured to determine a one-to-one correspondence between the objects before and after the observation interruption based on the similarity to associate the objects before and after the observation interruption. In some embodiments, the one-to-one correspondence between the objects before and after the observation interruption is determined based on a maximum sum similarity between the plurality of object information and the plurality of object models. In particular, determining the one-to-one correspondence between the objects can be implemented by a matching unit, i.e. the association unit comprises a matching unit configured to determine the one-to-one correspondence between the objects before and after the observation interruption based on the similarity. Of course in implementation, the operations of the similarity determining unit and the matching unit can both be implemented by the association unit itself.
[0077] In some embodiments, the association apparatus 400 can further comprise an alignment unit 403 configured to align the associated objects before and after the observation interruption. In some embodiments, the alignment unit is configured to align the object after the observation interruption with the pose of the object before the observation interruption by a spatial transformation.
[0078] In some embodiments, the association apparatus can further comprise a model reconstruction unit 404 configured to reconstruct a model of the object based on the consecutive images of the object. In some embodiments, the consecutive images of the object are a predetermined number of consecutive observation images before the observation interruption. In some embodiments, the reconstruction model of the object is selected from any one of a group comprising a mesh-based model, a point cloud model. In some embodiments, the consecutive images comprise consecutive RGB images and depth images of the object, and the model reconstructed based on the consecutive images is a mesh-based model. It should be noted that the model reconstruction unit can not be included in the association apparatus, and can be invoked by the association apparatus to perform model reconstruction when in operation.
[0079] It should be noted that the model reconstruction unit 404 is shown in dashed line to indicate that the model reconstruction unit 404 can also be located outside the model training apparatus 400, for example in this case, the apparatus 400 is still able to achieve the advantageous effects of the present disclosure as described before.
[0080] It should be noted that each unit described above is a logical module according to the specific function it implements, and is not intended to limit the specific implementation manner, for example, it can be implemented in software, hardware or a combination of software and hardware. In actual implementation, each unit described above can be implemented as an independent physical entity, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, each unit described above is indicated by a dashed line in the drawing, indicating that these units can not actually exist, and the operations / functions implemented by them can be implemented by the processing circuit itself.
[0081] In addition, although not shown, the device can also include a memory, which can store various information generated by the device, each unit included in the device in operation, programs and data for operation, data to be transmitted by the communication unit, etc. The memory can be a volatile memory and / or a non-volatile memory. For example, the memory can include, but is not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), flash memory. Of course, the memory can also be located outside the device. Alternatively, although not shown, the device can also include a communication unit, which can be used for communication with other devices. In one example, the communication unit can be implemented in a suitable manner known in the art, for example, including communication components such as antenna array and / or radio frequency link, various types of interfaces, communication units, etc. Here will not be described in detail. In addition, the device can also include other components not shown, such as radio frequency link, baseband processing unit, network interface, processor, controller, etc. Here will not be described in detail.
[0082] The scheme according to the present disclosure can form an object observation system alone or in combination with any existing object observation scheme, which can be used to observe objects, such as continuous observation, and when the observation is resumed after interruption / loss, the object localization in the discontinuous observation case according to the present disclosure is performed. Specifically, when the object observation system according to the disclosure is started, the object model in the scene can be reconstructed with continuous RGB-D video frames as input along with the object observation, which can be performed by a dynamic object reconstruction module. When the observation is interrupted / changed, the association module obtains the 2.5D instance point cloud of the new observation, and then evaluates the similarity between the reconstructed object model and the 2.5D instance point cloud in the new observation to find the one-to-one correspondence relationship before and after the observation interruption / scene change. These correspondences provide the approximate position of each object in the new scene. Then, in the alignment module, the object alignment is performed, for example, using a deep learning-based rigid point cloud registration method which is less sensitive to initialization and more robust to align two point cloud instance clusters. Thus, a novel dynamic object observation system is realized for the reconstruction, association and alignment of multiple invisible objects without additional training of new objects between different scenes. And experiments show that our system is a general and state-of-the-art system that can support various tasks such as model-free object pose estimation, single-view object completion, and real robot grasping.
[0083] The effectiveness of the scheme according to the present disclosure will be further demonstrated below in combination with experimental examples, including 6-DOF object pose estimation and robot grasping.
[0084] By evaluating on the public YCBVideo and Model-Free Object Pose Estimation Dataset (MOPED) datasets, it can be shown that the performance of the system according to the present disclosure in 6-DOF pose estimation is excellent, compared with zero-shot methods such as ZePHyR and model-based methods such as CosyPose, Pix2Pose and EPOS, etc. The method according to the present disclosure can obtain more optimal and accurate pose estimation.
[0085] The system according to the present disclosure can also be well applied to various application tasks, especially robot grasping tasks. A UR5 robot arm with a robotiq-2f-85 gripper and a wrist Realsense D435i RGB-D camera is used as the hardware platform. The system according to the present disclosure can help the robot grasping task in the following steps: a) the robot arm scans the cluttered objects placed on the table, and during the scanning process, the unknown object model is reconstructed by the system according to the present disclosure. b) the robot arm obtains a query image from a single view, and aligns the reconstructed object model with the object in the query image using the system according to the present disclosure, and obtains the aligned object point cloud. c) the aligned object point cloud is fed into the ready-made grasp pose generation model to generate candidate grasps. Compared with the single-view point cloud, the aligned object point cloud output by the system according to the present disclosure is more complete, for example, the objects that are completely or partially occluded can still be properly positioned, so that the grasp pose generation module can generate grasps for the occluded parts of the object or the occluded object.
[0086] Therefore, the present disclosure mainly considers a brand new task, i.e. dynamic object reconstruction and localization without continuous observation and known CAD model prior, and proposes a new system to perform dynamic object-level reconstruction, multi-object association and alignment. The system according to the present disclosure is universal and state-of-the-art in various tasks such as model-free object pose estimation, single-view object completion based on model alignment, and dynamic multi-object robot grasping.
[0087] Some embodiments of the present disclosure also provide an electronic device which can be operated to implement the operations / functions of the aforementioned model pre-training device and / or model training device. Figure 5 A block diagram of some embodiments of the electronic device of the present disclosure is shown. For example, in some embodiments, the electronic device 5 can be various types of devices, for example, can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet PCs), PMPs (Portable Multimedia Players), vehicle terminals (such as car navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. For example, the electronic device 5 can include a display panel for displaying data and / or execution results utilized in the scheme according to the present disclosure. For example, the display panel can be various shapes, such as a rectangular panel, an oval panel, or a polygonal panel, and the like. In addition, the display panel can not only be a flat panel, but also a curved panel, or even a spherical panel.
[0088] As Figure 5 shown, the electronic device 5 of this embodiment includes a memory 51 and a processor 52 coupled to the memory 51. It should be noted that Figure 5The components of the electronic device 50 shown are merely exemplary and not limiting, and the electronic device 50 can have other components according to actual application requirements. The processor 52 can control other components in the electronic device 50 to perform desired functions.
[0089] In some embodiments, the memory 51 is configured to store one or more computer readable instructions. When the processor 52 is configured to execute the computer readable instructions, the computer readable instructions are executed by the processor 52 to implement the method according to any of the above embodiments. For specific implementation of each step of the method and related explanations, please refer to the above embodiments, and the repeated parts will not be described here.
[0090] For example, the processor 52 and the memory 51 can communicate with each other directly or indirectly. For example, the processor 52 and the memory 51 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 52 and the memory 51 can also communicate with each other through a system bus, and the present disclosure does not limit the processor 52 and the memory 51.
[0091] For example, the processor 52 can be embodied as various appropriate processors, processing devices, etc., such as a central processing unit (CPU), a graphics processing unit (GPU), a network processing unit (NP), etc.; and can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The central processing unit (CPU) can be X86 or ARM architecture, etc. For example, the memory 51 can include any combination of various forms of computer readable storage media, such as volatile memory and / or non-volatile memory. The memory 51 may, for example, include system memory, which may, for example, store an operating system, application programs, a boot loader, a database, and other programs, etc. Various application programs and various data, etc. can also be stored in the storage medium.
[0092] In addition, according to some embodiments of the present disclosure, various operations / processes according to the present disclosure, when implemented by software and / or firmware, can be loaded from a storage medium or a network to a computer system with a dedicated hardware structure, such as Figure 6 The computer system 600 shown is installed with programs constituting the software, and the computer system, when installed with various programs, can perform various functions, including functions such as those described above, etc. Figure 6 is a block diagram showing an example structure of a computer system that can be employed in the embodiments according to the present disclosure.
[0093] In Figure 6In the central processing unit (CPU) 601, various processes are executed according to a program stored in a read only memory (ROM) 602 or a program loaded from the storage section 608 to a random access memory (RAM) 603. In the RAM 603, data required when the CPU 601 executes various processes and the like is also stored as necessary. The central processing unit is merely exemplary, and can be other types of processors, such as the various processors described above. The ROM 602, the RAM 603, and the storage section 608 can be various forms of computer readable storage media, as described below. Note that, although Figure 6 The ROM 602, the RAM 603, and the storage section 608 are shown separately in the central processing unit, but one or more of them can be combined or located in the same or different memory or storage module.
[0094] The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output interface 605 is also connected to the bus 604.
[0095] The following components are connected to the input / output interface 605: an input section 606, such as a touch panel, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, and the like; an output section 607, including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage section 608, including a hard disk, a magnetic tape, and the like; and a communication section 609, including a network interface card, such as a LAN card, a modem, and the like. The communication section 609 allows communication processing to be performed via a network, such as the Internet. It is readily understood that, although Figure 6 The various devices or modules in the electronic device 600 are shown in the central processing unit as communicating via the bus 604, but they can also communicate through a network or other means, where the network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0096] A drive 610 is also connected to the input / output interface 605 as necessary. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, is mounted on the drive 610 as necessary, so that a computer program read therefrom is installed in the storage section 608 as necessary.
[0097] In the case where the above series of processes are implemented by software, the program constituting the software can be installed from a network, such as the Internet, or a storage medium, such as the removable medium 611.
[0098] According to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods according to embodiments of the present disclosure. In such embodiments, the computer program can be downloaded and installed from a network by the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the CPU 601, the above-described functions defined in the methods of the embodiments of the present disclosure are executed.
[0099] It should be noted that, in the context of the present disclosure, a computer readable medium can be a tangible medium which can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer readable medium can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), or the like, or any suitable combination of the above.
[0100] The above computer readable medium can be included in the above electronic device; or can exist separately, without being assembled into the electronic device.
[0101] In some embodiments, a computer program including instructions which, when executed by a processor, causes the processor to carry out the method of any of the above embodiments is also provided. For example, the instructions can be embodied in a computer program code.
[0102] In an embodiment of the disclosure, computer program code to carry out operations of the disclosure can be written in one or more programming languages, or combinations thereof, including object oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0103] The flow diagrams and the block diagrams in the drawings are illustrations of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0104] The modules, components or units described in the embodiments of the present disclosure can be implemented by software or by hardware. In some cases, the name of the module, component or unit does not constitute a limitation on the module, component or unit itself.
[0105] The functionality described herein above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.
[0106] According to some embodiments of the present disclosure, a method for object localization in a discontinuous observation scenario is proposed, comprising the steps of: obtaining an object model based on a reference image obtained when observation resumes after observation interruption; and correlating objects before and after observation interruption based on the obtained object model and a reconstructed model of the object.
[0107] In some embodiments, the reconstructed model of the object is obtained by model reconstruction of the object based on consecutive images of the object. In some embodiments, the consecutive images of the object are a predetermined number of consecutive observation images before the observation interruption.
[0108] In some embodiments, the reconstructed model of the object is selected from any one of the group comprising a mesh-based model, a point cloud model.
[0109] In some embodiments, the consecutive images comprise consecutive RGB images and depth images of the object, and the model reconstructed based on the consecutive images is a mesh-based model.
[0110] In some embodiments, the object model obtained based on the reference image is a point cloud model of the object. In some embodiments, the reference image is a starting image obtained when observation resumes, and a 2.5D instance point cloud is obtained from the starting image as the point cloud model of the object.
[0111] In some embodiments, correlating objects before and after observation interruption based on the obtained object model and a reconstructed model of the object further comprises: determining similarity between the point cloud model of the object and the reconstructed model of the object based on object information, and correlating objects before and after observation interruption based on the similarity.
[0112] In some embodiments, the object information comprises at least one of object geometric features, object texture features, object color features. In some embodiments, the object information comprises both geometric features and color features, and the similarity between the point cloud model of the object and the reconstructed model of the object is determined based on both geometric features and color features.
[0113] In some embodiments, the associating the objects before and after the observation interruption based on the similarity further comprises determining a one-to-one correspondence between the objects before and after the observation interruption based on the similarity such that the objects before and after the observation interruption are associated. In some embodiments, the one-to-one correspondence between the objects before and after the observation interruption is determined based on a maximum sum similarity between the plurality of object information and the plurality of object models.
[0114] In some embodiments, the method further comprises aligning the associated objects before and after the observation interruption. In some embodiments, the objects after the observation interruption are aligned with the pose of the objects before the observation interruption by a spatial transformation.
[0115] According to some embodiments of the present disclosure, there is provided an apparatus for associating objects in a discontinuous observation scene, comprising: a model obtaining unit configured to obtain object models based on reference images obtained when observation is resumed after an observation interruption; and an associating unit configured to associate objects before and after the observation interruption based on the obtained object models and object reconstruction models.
[0116] In some embodiments, the associating unit further comprises a similarity determining unit configured to determine a similarity between the object point cloud models and the object reconstruction models based on the object information, and the associating unit further associates the objects before and after the observation interruption based on the similarity.
[0117] In some embodiments, the object information comprises both geometric features and color features, and the similarity determining unit is configured to determine the similarity between the object point cloud models and the object reconstruction models based on both the geometric features and the color features.
[0118] In some embodiments, the associating unit is further configured to determine a one-to-one correspondence between the objects before and after the observation interruption based on the similarity such that the objects before and after the observation interruption are associated. In some embodiments, the one-to-one correspondence between the objects before and after the observation interruption is determined based on a maximum sum similarity between the plurality of object information and the plurality of object models.
[0119] In some embodiments, the apparatus for associating objects can further comprise an aligning unit configured to align the associated objects before and after the observation interruption. In some embodiments, the aligning unit is configured to align the objects after the observation interruption with the pose of the objects before the observation interruption by a spatial transformation.
[0120] In some embodiments, the association apparatus can further comprise a model reconstruction unit configured to perform model reconstruction of the object based on the consecutive images of the object. In some embodiments, the consecutive images of the object are a predetermined number of consecutive observation images before the observation interruption. In some embodiments, the reconstructed model of the object is any one selected from a group comprising a mesh-based model, a point cloud model. In some embodiments, the consecutive images comprise consecutive RGB images and depth images of the object, and the model reconstructed based on the consecutive images is a mesh-based model.
[0121] According to yet some embodiments of the present disclosure, there is provided an electronic device comprising: a memory; and a processor coupled to the memory, the memory having stored therein instructions that, when executed by the processor, cause the electronic device to perform the method of any embodiment described in the present disclosure.
[0122] According to yet some embodiments of the present disclosure, there is provided a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the method of any embodiment described in the present disclosure.
[0123] According to yet some embodiments of the present disclosure, there is provided a computer program comprising: instructions which, when executed by a processor, cause the processor to perform the method of any embodiment described in the present disclosure.
[0124] According to some embodiments of the present disclosure, there is provided a computer program product comprising instructions which, when executed by a processor, implement the method of any embodiment described in the present disclosure.
[0125] The above description is merely exemplary of some embodiments of the present disclosure and of the principles thereof. It is understood that the scope of the disclosure is not limited to the specific combinations of technical features disclosed above, and that other technical solutions formed by any combination of the above technical features or equivalent features thereof without departing from the above disclosure should be included in the scope of the disclosure. For example, technical solutions formed by mutual replacement of the above features with technical features disclosed in the present disclosure (but not limited to) having similar functions should be included in the scope of the disclosure.
[0126] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure the understanding of this description.
[0127] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor imposing a sequential order of executed operations. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, while several specific implementation details are contained in the above discussion, these should not be taken as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented together in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0128] While certain aspects of the disclosure have been described with reference to particular embodiments, those skilled in the art will understand that the disclosure is illu- stative only and is not intended to be limiting. Many modifications and variations are possible in light of the above teachings. It is contemplated that, where any component or module is implemented in the embodiments described herein, there are a number of alternative ways to implement the same. The elements and acts of the various embodiments described above can be combined to provide further implementations. It is also contemplated that the components can be implemented using software, hardware, firmware, or combinations thereof that can be implemented in software and / or firmware. Additionally, one or more components of each of the individual embodiments can constitute separate devices, applications, or programs. Other specific arrangements and orderings will be apparent to those of ordinary skill in the art upon reading the above description. Additionally, the various embodiments described above can be implemented in a computer program product tangibly embodied in a machine-readable storage medium (e.g., magnetic disk, optical disk, memory, etc.) including a machine-readable program of instructions (e.g., program code) implementing the various processes, operations, and / or methods described herein. The program of instructions can be executed by one or more processors to cause the functions described herein to be performed.
Claims
1. A method for object localization in a discontinuous observation scene, comprising the following steps: An object model is obtained based on a reference image acquired during the resumption of observations after an interruption. The object model is a model of objects in the scene after observation resumption. The reference image is a single-view image acquired during observation resumption, and the object model obtained based on the reference image is an object point cloud model. The acquired object model and object reconstruction model are used to establish the association between objects before and after the observation interruption. The object reconstruction model is a model of objects in the scene before the observation interruption. This object reconstruction model is a surface-based model obtained by reconstructing the object model from continuous images of the object. The method of establishing object association before and after observation interruption based on the acquired object model and object reconstruction model further includes: The similarity between the object point cloud model and the object reconstruction model is determined based on object information, and The similarity is used to establish a correlation between objects before and after the observation interruption.
2. The method according to claim 1, wherein, The object reconstruction model is selected from any of the groups including surface-based models and point cloud models.
3. The method according to claim 1, wherein, Continuous images include continuous RGB images and depth images of an object.
4. The method according to any one of claims 1-3, wherein, A series of images of an object is a predetermined number of consecutive observation images before the observation is interrupted.
5. The method according to claim 1, wherein, The reference image is the initial image obtained during the recovery observation, and a 2.5D instance point cloud is obtained from the initial image as the object point cloud model.
6. The method according to claim 1, wherein, The object information includes at least one of the object's geometric features, object's texture features, and object's color features.
7. The method according to claim 1, wherein the object information includes both geometric features and color features, and the similarity between the object point cloud model and the object reconstruction model is determined based on both geometric features and color features.
8. The method according to claim 1, wherein, The correlation between objects before and after the observation interruption based on similarity further includes: Similarity is used to determine the one-to-one correspondence between objects before and after the observation interruption, so that the objects before and after the observation interruption can be associated.
9. The method according to claim 8, wherein, The one-to-one correspondence between objects before and after the observation interruption is determined based on the maximum sum similarity between multiple object information and multiple object models.
10. The method according to claim 1, wherein the method further comprises: Align objects before and after the associated observation is interrupted.
11. The method according to claim 10, wherein, Spatial transformation is used to align the pose of an object after the observation is interrupted with that of an object before the observation was interrupted.
12. An object association device in a discontinuous observation scene, comprising: The model acquisition unit is configured to acquire an object model based on a reference image obtained when observation is resumed after an observation interruption, wherein the reference image is an image obtained from a single viewpoint when observation is resumed, and the object model acquired based on the reference image is an object point cloud model. as well as The association unit is configured to associate objects before and after observation interruption based on the acquired object model and object reconstruction model, wherein the object reconstruction model is a surface-based model obtained by reconstructing the object from continuous images of the object, and the association unit further includes: A similarity determination unit is configured to determine the similarity between the object point cloud model and the object reconstruction model based on object information, and The association unit further associates objects before and after the observation interruption based on the similarity.
13. The apparatus of claim 12, further comprising: Alignment units are configured to align objects before and after an associated observation interruption.
14. An electronic device comprising: Memory; and A processor coupled to the memory, the memory storing instructions that, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1-11.
15. A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method according to any one of claims 1-11.
16. A computer program product comprising instructions that, when executed by a processor, cause implementation of the method according to any one of claims 1-11.
Citation Information
Patent Citations
Scene recovery method and device based on low-quality GRB-D data
CN105469103A
Sequence ISAR image scattering center multi-hypothesis tracking track correlation method
CN112782696A