Method and system for creating annotated object models for new real-world objects

The system interactively predicts and adjusts object characteristics using operator inputs, reducing effort and errors in annotating robot environments, thereby improving robotic systems' understanding.

JP7794866B2Active Publication Date: 2026-01-06HONDA MOTOR CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024028378
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-03-31
Filing Date
2024-02-28
Publication Date
2026-01-06
Estimated Expiration
2044-02-28

AI Technical Summary

Technical Problem

Current systems require significant manual effort to define and annotate the properties of objects in a robot's environment for behavioral planning, making it time-consuming and prone to errors.

Method used

A system that predicts object characteristics interactively, allowing operators to correct and adjust these properties using gestures, voice, or gaze, and stores the annotated models for robotic systems.

Benefits of technology

Reduces annotation effort and minimizes errors by providing a quick and intuitive method for creating annotated object models, enhancing robotic systems' understanding of their environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794866000001
    Figure 0007794866000001
  • Figure 0007794866000002
    Figure 0007794866000002
  • Figure 0007794866000003
    Figure 0007794866000003
Patent Text Reader

Abstract

To provide a method and a corresponding system for creating an annotated object model of a real-world object.SOLUTION: The method comprises the steps of: providing an initial object model for an object for which an annotated object model shall be created; predicting properties of the object; visualizing a representation of the object based on the initial object model wherein the predicted object properties are displayed associated with the representation; and obtaining selection information based on at least one of a user gesture, user pointing operation, user speech input, and user gaze perceived by a user perception device. The method determines a portion of the object corresponding to the selection information, receives property information from a user input, and associates the input property information with the corresponding portion of the object.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of providing knowledge about real-world objects for robotic systems, and in particular to creating annotated object models that can be used for behavior planning for autonomous robots such as domestic robots. [Background technology]

[0002] In recent years, assistance systems, including autonomous devices such as domestic robots, have become increasingly popular. Such systems require a reasonable understanding of the environment in which they are intended to operate to facilitate the behavioral planning required to accomplish a desired task. Unfortunately, this requires providing a definition of the objects that may potentially be present in the robot's environment, because the robot itself is not yet capable of automatically learning every detail of the robot's environment that may be important when performing a particular task. A deep "understanding" of the environment is essential for the robot and must be established before any task can be accomplished by the robot without being taught every single step required. Therefore, defining the characteristics of all potential objects in the robot's environment that are required for behavioral planning to accomplish the robot's desired task requires a significant effort. Currently, this information is largely input by the system operator, who recognizes the need to add new objects to the robot's knowledge base, allowing the robot to then operate autonomously in the environment, including these newly added objects. While adding objects can be easily done, adding all the properties that each object has, including its relations or even possible relations to other objects, is more difficult and time-consuming if this has to be started from scratch. However, these properties are necessary to improve the robot's understanding of its environment, so that the planning module of the robot (or robot system) can act on or use such objects. Summary of the Invention [Problem to be solved by the invention]

[0003] It would therefore be desirable if a system used to input properties about objects or parts of objects could automatically generate at least some of the information needed to be entered to annotate an object model about a new object. Such annotations (properties associated with the object or particular parts thereof) generated by the system should then preferably be corrected and adjusted by an operator of the system. Such correction and adjustment of annotations is much faster than having an operator directly input all the details about the properties an object may have. [Means for solving the problem]

[0004] The method and system according to the present invention are suitable for supporting the interactive creation of an annotated object model of a real-world object. The system predicts characteristics of an unannotated object model of a new object that can be corrected, confirmed, or rejected by an operator. The operator can also add additional characteristics. First, an initial object model is provided, which is the basis for creating an annotated object model of a new, as yet unknown real-world object. Therefore, the characteristics of the new object must be added to the initial object model, which defines the geometry of the new object. Based on this initial object model, a prediction of the object characteristics is provided by the system. This prediction of the object characteristics is then visualized in a representation of the object based on the initial object model.

[0005] Object characteristics are visualized along with the object's representation, making the association between the characteristics and the object or specific parts thereof clear to the operator. To add, remove, or adjust the object's predicted (proposed) characteristics, the operator can then input selection information, which is obtained by the system using a perceptual device that receives a user gesture, a user pointing operation, a user's voice input, or a user's gaze. A "gesture" can be understood as any movement of the operator's body part, including, for example, pointing specifically at something using an index finger. Such gestures can be observed by a camera used as a perceptual device. Alternatively, or as a supplementary input, the user's gaze can be identified and analyzed. Furthermore, a pointing operation can be obtained from a computer mouse as a perceptual device. For voice input, a microphone can be the perceptual device. Based on this input selection information, the part of the object corresponding to the selection information is identified. Note that various methods of inputting selection information can be combined.

[0006] The system then receives characteristic information about the identified area of ​​the object. The characteristic information received by the system can be provided by an operator, for example, using voice input, and can include a category or any type of characteristic of the object or part thereof. The voice input can include deletion of the originally predicted characteristic as well as further information to be annotated to the identified area of ​​the object. After the object characteristics have been corrected, deleted, or added, the resulting annotated object model is stored in a database containing world knowledge, for example, for a robotic system. [Brief explanation of the drawings]

[0007] [Figure 1]FIG. 1 shows a system overview including components according to a preferred embodiment of the present invention. [Figure 2] 1 is a simplified flowchart describing assisted annotation for creating a new annotated object model. [Figure 3] FIG. 1 shows an example of a visualization of the representation of the objects to be annotated and the process for refining the initial object model. [Figure 4] FIG. 1 illustrates an exemplary method for predicting annotations for an object from which an annotated object model will be created. [Figure 5] Figure 1 shows an example for predicting properties on an initial object model. [Figure 6] FIG. 10 illustrates exemplary details of a process for adapting predicted object properties. DETAILED DESCRIPTION OF THE INVENTION

[0008] In the following, embodiments of the present invention will be described in more detail with reference to the accompanying drawings. However, before going into the details and structural elements for realizing the present invention, a general understanding of the present invention will be improved by summarizing the method and system and explaining the respective advantages that the present invention has when a new object model needs to be annotated.

[0009] A key advantage of the present invention is that it provides an interactive method for annotating object models, thereby reducing the required effort and workload compared to commonly known procedures. According to the present invention, the system predicts the characteristics to be presented to the operator. Therefore, the operator is provided with a "provisionally" annotated object model, allowing the operator to quickly adapt characteristics only to the extent necessary. Interactive annotation allows intuitive selection of portions or areas of an object where annotations to the object model will be created, deleted, or adapted. This intuitive selection of relevant portions or regions of an object significantly reduces the subjective effort required for the annotation process. In particular, when the intuitive selection of a portion or region of an object is determined from the operator's gestures or general movements and actions, e.g., the operator's hands or gaze, and the adaptation of annotations (characteristics) associated with this portion or region of the object is performed using voice input from the operator, annotations for new objects can be created fairly quickly and with a reduced probability of making mistakes.

[0010] The procedures and systems described herein for annotating new objects can be combined with commonly known systems that use input of information, for example, via a keyboard or mouse. Such auxiliary methods of inputting information into the system are particularly advantageous in cases where voice recognition, the preferred method of inputting information for adjusting annotations on objects, introduces errors due to misinterpretation of voice input by the operator.

[0011] The display of the object, based on the initial object model and displayed with predicted object properties, preferably uses an augmented reality or virtual reality display device. While other types of displays can be used to present the display to the operator, allowing the operator to input selection information and make annotation adjustments, an augmented reality (AR) or virtual reality (VR) device is preferred. By using such an AR or VR device, the operator can easily recognize the association between parts or portions of the overall object represented on the AR or VR display device and their respective predicted properties. For example, the display can be visualized with the properties, where the properties are displayed in close spatial relationship to the parts to which they belong. The relationship can even be indicated by using arrows or connecting lines. Properties directly related to the appearance of the object can be displayed directly in the display of the object itself. For example, a color reflecting the property "color" can be used to render the display of the object. Again, this allows an easy and intuitive way for the operator to adapt the associated properties. Specifically, corrections to such properties can be displayed in real time. Because if the operator selects a surface for the display of a new object and inputs new color information, the display using the predicted color can be replaced by a display using the newly input color information. Thus, a change in color as a property of a particular area, or even of the entire object, becomes immediately visible, and the operator can therefore directly recognize that the correct property is now associated with the object or part thereof. It is clear that color is only one example, and the same is valid for any modification of a property that can be visualized in a display.

[0012] When a new object for which an annotated object model is to be created exists in the system's environment, a camera can be used to capture an image of the new object or perform a 3D scan of the object, so that an overlay of object characteristics can be displayed using augmented reality. The captured image for the 3D scan can also be used to automatically generate an initial object model. However, it is not absolutely necessary that the new object actually exists, or that an image can be captured or scanned, or that an image of the new object can be captured by a camera. Rather, only the (unanotated) object model of the new object can be known, and a display can be provided based on that object model using virtual reality.

[0013] As described above, the selection of a part, region, or area of ​​a new object is determined from the perception of the operator, particularly the gestures the operator makes with his or her hands. However, the perception can also include tracking the operator's eye movements, thereby determining the location on the object's surface at which the operator is focused. Therefore, in addition to analyzing gestures to obtain selection information, the analysis can also be supplemented with information received from an eye tracker.

[0014] According to one alternative embodiment, the information received from the eye tracker can even substitute for the perception of a gesture made by the operator's hand to identify at least a portion of the object selected by the operator. In any case, the perceived information can be supplemented by a command entered by the operator. For example, based on manipulation perception, selection information defining a location on the surface of the object is first required. If the operator touches a specific point on an object (virtual or real-world object), for example, with a fingertip, the touch location is interpreted as selection information. In the case of augmented reality, such contact can be a contact between the operator's fingertip and the real-world object; however, if a virtual reality display is used, the contact point between the fingertip and the representation of the object can also be identified by calculating a collision between the operator's hand and the virtual object. Once the contact location is identified, a region within which subsequent annotations or corrections are valid must be defined. Without any further instruction from the operator, the region can be considered as a region surrounding the identified location with a predefined size. The predefined size can be used to define a default area identified around the contact location. However, further input commands can be used to enhance the selection, for example by adjusting the size of the area or region.

[0015] Furthermore, depending on the identified touch location, the default for identifying a specific part of the object based on the selection information can be different; for example, in cases where an object segmentation is overlaid on a new object display as a predicted characteristic, touching a frame showing the outer boundary of the identified segment, which is used to visualize the object segment, can be interpreted as selecting the entire object segment and therefore a part of the object. This will become clearer when an example of annotation adaptation is described with reference to the drawings. However, as a general information, the specific location specified by the operator can serve as a trigger for using one of several default settings used for identifying the area corresponding to the selection information, regardless of how this location is specified (hand gesture, eye movement, ...).

[0016] The selection information entered by the operator may even include multiple different pieces of information; for example, a property of one part of an object may define a possible interaction between that part and another part of the object. For example, a bottle cap may be attached onto the open end of the bottle, whereas a cork may be inserted into the open end of the bottle. Such properties that define a potential relationship between one part and another part may also be defined using the present invention. In such cases, the selection information may include at least first information and second information. The first information defines the first part, and the second information defines the second part, whereby the relationship between these two parts may then be defined, for example, as a property associated with the first part identified based on the first selection information. Such input of selection information may even be split. For example, a first selection information may be entered into the system, followed by a property to be associated ("may be placed on"), after which a second selection information is added to complete the annotated relationship between these two parts. Obviously, the selection information may even include more than two parts.

[0017] To assist the operator in inputting selection information, when virtual reality is used to provide the representation of the object, it is advantageous if the system provides feedback about contact or collision between the operator's hand and the representation of the object. This can be achieved, for example, using a wristband or glove with an actuator that provides a stimulus indicating that the operator has "touched" the representation of the object. Such feedback would further improve the intuitive nature of inputting selection information using natural gestures by the operator. Contact can be identified from a calculated collision between the object and the operator's hand. Such calculation of collisions using, for example, a point cloud to describe an object model and the operator's hand is known in the art.

[0018] As described above, the perception of the operator's actions and movements is used to identify the location of the object indicated by the operator in order to annotate the respective parts or areas of the new object. However, the perception of the operator's gestures can even be used to improve the starting point of the annotation procedure, i.e., the initial object model. For example, the initial object model used to represent the object to which the annotations will be applied can be refined using gestures. The mesh (triangle mesh, polygonal mesh), point cloud, etc. used to form the object model and create the representation can be corrected using perceived gestures from the operator. In this way, the initial object model can be adjusted and refined to more closely resemble the desired new object. In cases where a camera is used to capture images of the new object, an overlay of the representation of the object with the mesh, point cloud, etc. can be presented to the operator, who can then directly identify specific parts of the mesh (points of the point cloud, etc.) and expand or remove specific areas to more closely reflect the true shape of the object. Such a fit of the initial object model can be immediately visualized, whereby the process of adding, correcting, or deleting annotations can be initiated by the user once the user is satisfied with the similarity between the initial object model and the real-world model.

[0019] In principle, prediction of object properties can be performed in several different ways. One possibility is to start with an initial object model, which may have been refined to correspond more closely to the new object, and then apply an algorithm to the model that analyzes the object modeling data to identify properties related to the whole object or its parts. Such an algorithm, for example a so-called part detector, can be run directly on the initial model itself, or it can take into account information retrieved from a database to identify known properties of other objects, in which case similarities between the new object and already annotated objects can be identified. Such similarities may relate to the whole object, or to parts of the object that have been identified as segments or parts of the whole object.

[0020] Alternatively, the predicted properties of an object or part thereof can be the result of a morphing process. Morphing involves morphing an already annotated template object model onto an initial object model that defines the geometry or shape of the new object. During the morphing process, annotations associated with specific nodes or points of the model are preserved and thus transferred to the morphed result, which represents the new object. Thus, the result of the morphing process is an initial object model that is already annotated with predicted properties that are transferred from the template object model by the morphing process. These annotations are then used as predicted properties that can be supplemented, deleted, or adjusted by an operator as described above.

[0021] FIG. 1 presents a system overview of a system 1 for creating a new annotated object model. The system 1 includes a processor 2 connected to an output device, preferably an augmented reality or virtual reality display output device 3. The processor 2 generates signals that are provided to the output device 3 to visualize a representation of the object to be annotated as well as predicted properties of the object. The representation is based on an initial object model that can be retrieved from a data storage 4, e.g., a database stored in an internal or external memory. After the new object model to be annotated has been annotated, the processor 2 can be configured to store the newly created annotated object model in the database in the data storage. Thus, the new object model, including its associated annotations, will be available for future annotation processes and may serve as, for example, an improved starting point.

[0022] The software used to perform the described methods may have a modular structure, whereby, for example, part detectors, manipulations, morphing, etc. are realized in multiple different software modules, each of which executes on processor 2. However, processor 2 is given only as an example, and the entire computation may be performed by multiple processors performing the computations collaboratively. This may even include cloud computing.

[0023] The processor 2 is further connected to a camera 5. The camera 5 is configured to capture an image of an object or provide a scan of the object, allowing for identifying the three-dimensional shape of the object from which an annotated object model is to be created. Image data of the image captured by the camera 5 is provided to the processor 2, whereby an analysis algorithm can be performed on the image data, e.g., to automatically create an initial object model. The image data can then be processed to prepare a signal that can be provided to an output device 3 to generate a representation of the object to be annotated.

[0024] The processor 2 is further connected to a perception device 6 that allows observation of the operator, in particular the gestures the operator makes or the operator's gaze. The perception device 6 can therefore include a camera 7 for capturing images of the hand, thereby enabling touching gestures, sliding movements, etc., performed by the operator using his or her hand to be identified. The camera 7 captures images from the operator's hand and provides the respective image data to the processor 2. The processor 2 is configured to analyze the provided data and thus calculates the pose and movement of the operator's hand. This allows the system to recognize whether the operator has touched or manipulated an object, identify the location to which the annotation will be applied if a real-world object is used by the operator, or calculate the corresponding location if a virtual object presented by the virtual reality output device 3 is to be annotated. In such cases, collisions between a point cloud or mesh modeling the operator's hand and a model representing the object to be annotated can be calculated. Note that the operator's hand is provided only as an intuitive example.

[0025] Additionally, the perception device may include an eye tracker 8. Perception of the operator's eye movements makes it possible to conclude which part of the object to be annotated the operator is looking at. This information is determined by processor 2 from data provided to processor 2 by perception device 6. Therefore, it is possible to identify areas of interest to the operator based on the information obtained by eye tracker 8. This information can then be used to supplement the location determined from the gesture determination based on input received from camera 7. However, it is also possible to use information received from eye tracker 8 to identify which area of ​​the object to be annotated by the operator. The information received from eye tracker 8, as well as the information received from camera 7 observing the operator's hand gestures and movements, is selection information because it contains information that can be analyzed to determine the location of the object corresponding to the operator's (virtual) touch or the location where the operator is looking. Based on this selection information, processor 2 identifies the part of the object that corresponds to the selection information by calculating the location of the collision between the hand and the object or the location where the operator is looking. It should be noted that camera 5 and camera 7 can be the same component, and a distinction is made between camera 5 and camera 7 only for the purposes of the present description.

[0026] The area corresponding to the selection information is then visualized using the output device 3. This can be achieved, for example, by highlighting the collision location or the location looked by the operator in a representation of the object to be annotated by the output device 3, as shown in the left part of Figure 6.

[0027] The perception device 6 may further include a microphone 9. Using the microphone 9, voice commands from the operator are obtained and respective signals are submitted to the processor 2. Voice commands can be used by the operator to instruct the system 1. These instructions can include adding annotations, deleting annotations, or adjusting annotations, but can also include augmentation information used to improve the system's understanding of gestures used to define new locations on the object. One example may be that the region of the object that corresponds to a specific location identified by the processor 2 from the operator's perceived gesture can be adjusted in terms of its dimensions. Starting from a default setting that identifies a region around the calculated contact point between the operator's finger and the object to be annotated, the operator can use voice commands to increase or decrease the size of the identified area surrounding the identified contact point.

[0028] FIG. 2 shows a simplified flowchart illustrating method steps for annotating an object model for a new, as-yet-unknown object. In step S1, the new object is identified. Identification of the new object can begin with an image captured by camera 5, based on which an appropriate object model is selected, for example, from a database or generated using a 3D scan of the new object. Selection can be performed by processor 2 searching for known object models in the database and comparing these object models with the shape and geometry of the new object. The model with the highest similarity to the new object can then be selected, thereby providing the initial object model in step S2. Alternatively, object model identification can be performed based on operator input, using their own personal understanding and considerations to identify object models with a high degree of similarity to the object to be annotated.

[0029] The term "object model" can be understood as data capable of describing the shape of an object. The object model can be a point cloud, a triangle mesh, or a polygonal mesh. Once the initial object model is provided, it can be adjusted and refined in step S3 to more closely correspond to the object for which the annotated object model will be created. This adjustment of the initial object model can be performed using operator input. For example, in visualizing a display based on the initial object model, an operator can select nodes or points of the object model and shift or delete them, so that the resulting display based on the adjusted initial object model better reflects the shape of the new object. When such adjustment and refinement of the initial object model is performed, the following description refers to the adjusted / refined initial object model.

[0030] After the initial object model is satisfactorily adjusted for the new object, prediction of object characteristics is performed in step S4. This prediction can be performed directly based on the representation using an algorithm that uses a part detector to calculate specific characteristics, for example, various segments, from the object model. Algorithms for performing such characteristic prediction exist in the art, and one skilled in the art would easily select a suitable algorithm for predicting specific object characteristics. The part detector can be trained on annotated object models from an existing database or can be the result of past interactive object annotation. To learn from the operator's response (accepting the suggested segmentation or rejecting the suggestion), the part detector can be adapted to the operator's input. Alternatively, another object model similar to the initial object model and for which annotations are already available can be selected from the database and used as a template object model. This template object model is then morphed to convert the template model into the initial object model. This conversion is used to predict characteristics for the initial object model by transferring annotation information from the template object model to the initial object model. This can be achieved by transferring segmentation or other annotations, for example, to the nearest model vertices or faces. The resulting annotations of the initial object model are then used as a starting point for adapting the annotations based on operator input.

[0031] Examples of object properties include part designations, defining different parts of an object by segmenting the whole object, affordances such as graspable, placeable, detachable, relationships between object parts, appearance properties such as color, and material properties such as wood, steel, glass, etc. It should be noted that relationships between object parts are not limited to the same object but can also include relationships to multiple objects. Examples of such relationships are "can be placed in", "can be placed on", "fits in", ...

[0032] Then, in step S5, the predicted characteristics (annotations) are visualized along with a representation of the object itself. The visualization is performed so that the location of an object where a particular characteristic is displayed corresponds to the portion or area of ​​that portion for which the respective characteristic is valid, or to the entire object. For example, if the object is a particular type of bottle that includes a cap, the characteristics associated with the cap are displayed in close spatial relationship, allowing the operator to directly recognize this association. In cases where the visualization provides the operator with a large number of characteristics that may obscure the spatial relationship, connecting lines or other aids, such as the use of color coding, can be added. For example, the color coding can use the same color to identify a particular portion (e.g., the bottle cap) and the characteristics listed in addition to the bottle cap. A different color can be used for the bottle body and the characteristics listed in addition to the bottle body.

[0033] Based on the visualization of the initial object model display and the predicted properties, the operator then begins selecting a specific portion or location of the display, and thus the object being represented. Based on such input of selection information used by the operator to select a portion of the object, the portion corresponding to the selection information is identified by processor 2. Once the portion corresponding to the selection information entered by the operator is identified, the operator can begin adding, deleting, or adjusting properties related to this identified portion of the object. Note that the portion may be an entire portion of the object to be annotated. In relation to the example using the cap and body that generally form a bottle, the location touched by the operator may correspond to a displayed frame used to indicate a specific portion of the object. The segment may be, for example, the cap, and if the location touched by the operator is identified as a point in the frame indicating the segment corresponding to the cap, the system understands that properties related to the entire portion will be adapted. Identification of the portion corresponding to the selection information entered by the user is performed in step S7.

[0034] In the case where a virtual reality output device 3 is used, feedback informing the operator of the identified contact with the virtual object is provided in step S8, for example using gloves or a wristband with multiple actuators. Both gloves and wristbands with multiple actuators allow stimulation of the operator's hand or wrist, making it intuitively recognizable that the operator has now collided with the object to be annotated and submitted selection information to the processor 2.

[0035] Once the area of ​​the object or of the object model representing the object selected by the operator through input of selection information has been identified, the operator adapts the annotations (object properties) as desired. To this end, system 1 acquires information about the desired adaptation in step S9. Preferably, microphone 9 is used to receive audio information from the operator, and based on such audio input, processor 2 identifies the adaptation intended to be performed by the operator. Such adaptation can include adding annotations / properties, deleting predicted annotations / properties, and even adjustments. Any audio input that allows identification of a portion of the object is considered to be related to the adaptation of properties associated with this particular portion or part of the object. The process of adapting the annotations / properties of the identified area of ​​the object is terminated only if a keyword is used to trigger a different function. Such a keyword could be, for example, "end annotation." After the operator has terminated the annotation adaptation using such a keyword, it can be concluded that no further adaptation of annotations / properties is intended, and the created annotated object model is then stored in the database as described above in step S12. Therefore, in step S10, it is determined whether further input from the operator is expected. The decision may also be based on a timeout, for example, when no further voice input can be recognized for a certain period of time.

[0036] If further input is made by the operator, the procedure proceeds again to step S5 based on the updated visualization using the properties already adapted. Thus, at any given time, the operator is provided with all the information available for the object model at that time.

[0037] FIG. 3 illustrates the process of interactively identifying object models that can be used as starting points for predicting properties for the model to be annotated. An image of a new object is captured by a camera, and object models that may be suitable starting points are searched for, for example, in a database based on similarity. The system attempts to identify the object type and makes respective suggestions. In the example shown, the system suggests that the new object is a vase. Using voice input, the operator corrects the suggestion by informing the system that the object is a bottle. In the example shown, the system suggests a point cloud, which can be corrected by the user, as shown in FIG. 3a. The correction can include removing outliers in the point cloud on edge surfaces. This correction is preferably performed during the scanning phase, in which a camera 5 captures images of the object to identify and suggest the appropriate category to the operator.

[0038] The frame surrounding the object shown in Figure 3b indicates that a single segment of it has been identified by System 1 and presented as an overlay on the representation of the bottle.

[0039] Figure 4 illustrates the process of feature prediction using morphing of known objects, as described above, that are in categories suggested by the system and confirmed by the operator, or from categories directly identified by the operator. Starting with a known object model (template object model) that has already been annotated, the point cloud of the template object model is morphed onto the point cloud of an initial object model that has not yet been annotated. Morphing begins with a rough alignment of the two point clouds. If the alignment results in significant errors, such as the template object model being upside down relative to the initial object model, the operator can rotate the template object model to find a better match with the initial object model.

[0040] Annotations made for a known object and thus included in the template object model are preserved during the morphing process, thereby linking these annotations to the local geometry. Therefore, at the end of the morphing process, the initial object model of the new object inherits the corresponding annotations / properties from the corresponding local geometry. These inherited properties are then considered to represent predicted properties for the new object. As explained above, these predicted properties can be added, deleted, or adjusted to ultimately create an annotated object model for the new object. The principle of starting from a known object and its annotated object model, morphing the template model onto the initial object model of a new, unknown object, thereby inheriting the annotations because they are associated with the local geometry, is illustrated in Figure 4.

[0041] Figure 5 shows further details of the interactive annotation process. As described with respect to Figure 3, the operator confirms or corrects the category of the new object, and the system can then suggest different objects that belong to this category, if available. In the embodiment shown, two different object models for a bottle are presented that can serve as starting points for predicting annotations.

[0042] The system first suggests Example 1. This suggestion is visualized directly using the augmented reality output device 3. In case the operator would rather choose a second object model of the same category, the operator uses a voice command to force the system to switch to Example 2 as the starting point for further annotation. Therefore, in this case, the operator enters "Choose Example 2." As can be seen on the right side of Figure 5, the system immediately switches to the second example of the known object model and displays a representation based on this second object model being morphed onto the initial object model as described above, whereby the annotations are transferred to the initial object model as predicted properties.

[0043] FIG. 6 illustrates how user gestures, perceived by camera 7 and analyzed by processor 2, are used in combination with voice input to revise the predicted properties of a new object. In the lower left side of FIG. 6, it can be seen that the operator touches an area of ​​the object that, based on the predicted properties of the selected model, was identified as belonging to the body of the bottle. The predicted parts "body" and "cap" are displayed in a close spatial relationship to the corresponding parts of the bottle representation. However, the operator wants to change this area to belong to the "cap" part. Therefore, to enter selection information, the operator touches the area assigned to the incorrectly identified body and corrects this using a voice command ("That's the cap"). The system then automatically adapts the segmentation to identify the two parts of the bottle, namely, "body" and "cap," according to the identified areas identified from the selection information entered by the operator. The segments are identified by the system as shown in the upper part of FIG. 6. The locations identified from the entered selection information are indicated by open circles. As can be seen in the center of FIG. 6, this can still result in an incorrect segmentation. An additional input, shown in the lower right portion of Figure 6, allows for correction of the system's misinterpretation. A sliding gesture made by the operator is used to indicate that the entire area traversed by the sliding gesture will be considered for the annotation to be entered. This gesture can even be reinforced by using voice input. In this case, the voice command "expand cap" makes it clear that the area defined by the sliding motion is now all part of the cap.

[0044] Alternatively, the frames showing the segments shown in the upper part of FIG. 6 can be directly shifted by the operator to adjust the proposed segmentation.

[0045] In the illustrated embodiment, only a single viewpoint is shown. However, in that viewpoint, some parts of the object that are important and require annotation may be occluded. Therefore, the operator can manipulate the view of the object to make other parts visible. This manipulation can include zooming the view as well as changing the viewpoint.

[0046] The above-described method for creating annotated object models of new real-world objects does not need to be performed before operating the robotic system. This method is also particularly advantageous for application in telerobotic systems. This allows the operator of a telerobotic system to teach the system new objects (including their properties) while the system is in use. To ensure appropriate selection information in such situations, the robot can use a laser pointer to allow the teleoperator to control the location of the new object to which the teleoperator wishes to point in order to provide the desired annotation (including deleting the annotation). Furthermore, the robot's arm can be used in place of the operator's hand, as described above for inputting selection information. [Explanation of symbols]

[0047] 1 System 2 processors 3 Augmented reality or virtual reality display output devices 4. Data Storage 5. Camera 6 Perceptual Devices 7. Camera 8. Eye Tracker 9 Microphone

Claims

1. 1. A method for creating an annotated object model of a real-world object, comprising: a method step of providing an initial object model for an object for which an annotated object model is to be created; a method step of predicting properties of the object by generating predicted properties of the object with respect to the initial object model by morphing a template model onto the initial object model and carrying over known properties of the template model to the initial object model as the predicted properties; - a method step of visualizing a representation of said object based on said initial object model, wherein said predicted properties of said object are displayed in association with said representation; obtaining selection information based on at least one of a user gesture, a user pointing action, a user voice input, and a user gaze sensed by a user perception device; a method step of identifying a portion of the object corresponding to the selection information; the method steps of receiving characteristic information from a user input; and a method step of associating said input characteristic information with said corresponding portion of said object.

2. The display is visualized using an Augmented Reality (AR) or Virtual Reality (VR) display; The method of claim 1.

3. the selection information is augmented by commands entered by an operator; The method of claim 1.

4. the selection information includes at least first selection information and second selection information, and the characteristic information defines a relationship between the portions of the object corresponding to the at least first and second selection information, or a characteristic common to the portions; The method of claim 1.

5. the identification of the portion of the object is dependent on at least one of a type of gesture, a location of a collision between the representation of the object and a perceived operator's hand, and a pointing location on the displayed representation. The method of claim 1.

6. providing feedback to an operator if a collision between the representation of the object and the perceived operator's hand is identified; The method of claim 5.

7. the predicted characteristics include at least a definition of a segment defining a portion of the object; The method of claim 1.

8. the predicted characteristics include at least a definition of a segment, and the segment of the object is visualized as an overlay on the representation of the object. The method of claim 1.

9. the associated representation of the predicted properties of the object displays each predicted property in spatial relationship to a respective portion of the object for which the property is predicted, and the representation of the object, along with the associated predicted properties of the object, is manipulated according to operational input received from an operator. The method of claim 1.

10. based on the perceived operator input, the initial object model is adapted and the displayed representation is updated accordingly; The method of claim 1.

11. the initial object model is analyzed by an object part detector for automated segmentation of the object based on at least one of database information and information previously received from an operator during the creation of the annotated object model process; The method of claim 1.

12. 1. A system for creating an annotated object model of a real-world object, the system comprising: a processor; an output device; and an operator perception device; wherein the processor is configured to predict properties of the object by generating predicted properties of the object for an initial object model by morphing a template model onto the initial object model and inheriting known properties of the template model as the predicted properties into the initial object model, so as to provide an initial object model for the object for which an annotated object model is to be created; control the output device to visualize a representation of the object based on the initial object model, the predicted properties of the object being displayed in association with the representation; identify a portion of the object corresponding to selection information obtained by the operator perception device based on at least one of a user gesture, a user pointing operation, a user voice input, and a user gaze sensed by the operator perception device, and associate property information input by the operator with the corresponding portion of the object.

13. The system of claim 12 , wherein the output device comprises an Augmented Reality (AR) or Virtual Reality (VR) display.

14. The system of claim 12 , wherein the operator perception device includes at least a camera for perceiving operator movements.

15. The system of claim 12 , wherein the operator perception device includes at least a microphone.

16. The system of claim 12 , wherein the system includes a feedback device for informing the operator of an identified contact with the visualized representation of the object.

17. 13. The system of claim 12, wherein the processor is coupled to a database, the processor being configured to store the annotated object model in the database.

Citation Information

Patent Citations

  • Change detection method using ar overlay and system therefor

    JP2021131853A

  • Information processing device, information processing method, and information processing program

    WO2021235316A1