Model making device and model making method
By processing the image of the registered object and the feature information of the reference model in the model production device, the reference model is corrected to create a model that reflects the shape of the registered object, and the problems of insufficient identification performance and large data processing volume in the prior art are solved, and efficient model production and recognition performance improvement is achieved.
Patent Information
- Application Number
- CN202080059275.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-28
- Filing Date
- 2020-11-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-11-17
AI Technical Summary
The prior art is difficult to effectively evaluate the recognition impact of local areas on newly produced 3D models, resulting in insufficient recognition performance and a large amount of data and processing are required during model production.
A model production device is adopted to process the image and reference model of the registered object through a processor and memory, and to correct the reference model based on feature information to create a model reflecting the shape of the registered object.
A local information model for registered objects that reflect the impact of identification performance is produced with a smaller amount of data and processing volume, which improves the efficiency of model production and identification performance.
Smart Images

Figure CN114303173B_ABST
Abstract
Description
[0001] Reference-based Citation
[0002] This application claims the priority of Japanese Patent Application No. 2019-215673 filed on November 28, 2019, the content of which is incorporated herein by reference. Technical Field
[0003] The present invention relates to a model making apparatus and a model making method. Background Art
[0004] As background art in this technical field, there is Japanese Patent Laid-Open No. 8-233556 (Patent Document 1). In this publication, it is described that: "There is provided: a photographing unit 1; a first image storage unit 3 that stores a subject image of a subject from a specified viewpoint position photographed by the photographing unit 1; a three-dimensional shape model storage unit 2 that generates an object image from a viewpoint position closest to the photographed subject image based on a standard three-dimensional shape model; a second image storage unit 4 that stores the generated object image; a difference extraction unit 5 that extracts the difference between the subject image and the object image stored in each image storage unit; and a shape model trimming unit that trims the standard three-dimensional shape model based on the extracted difference. By trimming the standard three-dimensional shape model, which is a representative shape model of the subject, based on the difference between the subject image and the object image, the shape model of the subject is restored." (Refer to the abstract of the specification).
[0005] Prior Art Documents
[0006] Patent Documents
[0007] Patent Document 1: Japanese Patent Laid-Open No. 8-233556 Summary of the Invention
[0008] Problems to be Solved by the Invention
[0009] In the technology described in Patent Document 1, it is difficult to estimate the degree of influence that a local area has on the recognition of a newly created 3D model, so it is difficult to evaluate to what extent the local area should be correctly reflected in the 3D model. That is, in the technology described in Patent Document 1, there is a possibility that the recognition performance of the new 3D model may be insufficient due to insufficient evaluation of the local area. In addition, in the technology described in Patent Document 1, since changes (noise) in local areas that hardly affect the recognition of the 3D model of the object image are also reflected in the new 3D model, a large amount of data and processing may be required when creating the new 3D model.
[0010] In addition, in the technology described in Patent Document 1, in order to determine to what extent a local area should be correctly reflected in a 3D model, a large amount of data and processing are required. Therefore, an object of the present invention is to create a model of a registration object that reflects local information of the registration object that affects recognition performance with a smaller amount of data and processing.
[0011] Means for Solving the Problem
[0012] To solve the above problems, one aspect of the present invention adopts the following configuration. A model production device that produces a model representing the shape of a registration object, comprising a processor and a memory, the memory holding: images of one or more poses of the registration object; and a reference model representing the shape of a reference object, the processor obtains information representing the characteristics of a first pose of the registration object, and when it is determined based on a specified first condition that the shape of the first pose represented by the reference model is not similar, the reference model is corrected based on the information representing the characteristics to produce a model representing the shape of the registration object.
[0013] Advantages of the Invention
[0014] According to one aspect of the present invention, it is possible to create a model of a registration object that reflects local information of the registration object that affects recognition performance with a smaller amount of data and processing.
[0015] The problems, configurations, and effects other than the above will become clearer through the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a block diagram showing a functional configuration example of the model production device according to Embodiment 1.
[0017] Figure 2 It is a block diagram showing a hardware configuration example of the model production device according to Embodiment 1.
[0018] Figure 3 It is an example of an imaging system that captures an image of the registration object 20 provided to the model production device according to Embodiment 1.
[0019] Figure 4 It is a flowchart showing an example of a model production process for producing a 3D model of a registration object according to Embodiment 1.
[0020] Figure 5 It is a flowchart showing an example of a 3D model correction process according to Embodiment 1.
[0021] Figure 6 It is an explanatory diagram showing a specific example of a process for determining whether there is a 3D model correction according to Embodiment 1.
[0022] Figure 7 It is an explanatory diagram showing a detailed example of the 3D model correction process of Example 1.
[0023] Figure 8 It is an explanatory diagram showing a specific example of the process for determining the presence or absence of 3D model correction in Example 2.
[0024] Figure 9 It is an explanatory diagram showing a specific example of the process for determining the presence or absence of 3D model correction in Example 2.
[0025] Figure 10 It is an explanatory diagram showing an example of the 3D model selection process in Example 3.
[0026] Figure 11 It is an explanatory diagram showing an example of the 3D model selection process in Example 3.
[0027] Figure 12 It is a flowchart showing an example of the model production process in Example 4.
[0028] Figure 13 It is an explanatory diagram showing an example of the correction process of the feature extractor in Example 4.
[0029] Figure 14 It is an explanatory diagram showing an example of the correction process of the feature extractor in Example 4. Detailed implementation mode
[0030] Hereinafter, embodiments of the present invention will be described in detail based on the drawings. In this embodiment, the same reference numerals are generally given to the same components, and repeated descriptions are omitted. In addition, it should be noted that this embodiment is merely an example for implementing the present invention and does not limit the technical scope of the present invention.
[0031] Example 1
[0032] Figure 1 It is a block diagram showing a functional configuration example of a model production device. The model production device 100 produces a model representing the shape of a newly registered object to be registered using a model representing the shape of a registered reference object. A 3D (three-dimensional) model capable of representing the shape of an object using vertices and meshes (faces) is an example of such a model. In this example, an example of using a 3D model to represent the shape of an object is mainly described, but other models such as 2D models can also be used. In addition, the model can represent not only the shape of an object but also patterns, viewpoints, etc.
[0033] The model production device 100 has, for example, an image acquisition unit 111, an identification unit 112, an identification result comparison unit 113, a model correction unit 114, and an output unit 115. The image acquisition unit 111 acquires an image of the object to be registered. The identification unit 112 outputs the pose of the object by inputting the image of the object to a feature extractor described later.
[0034] The identification result comparison unit 113 determines whether the pose obtained by inputting the image of the object to be registered to the feature extractor is the correct pose. The model correction unit 114 corrects the 3D model of the reference object to produce the 3D model of the object to be registered. The output unit 115 outputs information related to the images of the reference object and the object to be registered, information related to the pose output by the feature extractor, and information related to the produced 3D model, etc.
[0035] In addition, the model production device 100 holds image data 131 and model data 132. The image data 131 is data in which images of one or more poses of one or more reference objects and images of one or more poses of a newly registered object acquired by the image acquisition unit 111 are associated with the poses. Images of one or more poses of the reference object are pre-included in the image data 131.
[0036] The model data 132 includes a 3D model representing the shape of the reference object and a 3D model representing the shape of the registered object produced by the model production device 100. The 3D model representing the shape of the reference object is pre-included in the model data 132 before the model production process is executed. In addition, in the model data 132, the object corresponding to each 3D model and the type to which the object belongs are defined.
[0037] In addition, the model data 132 has a feature extractor corresponding to each reference object. If an image of an object is input to the feature extractor, the features of the image are extracted, the pose of the object in the image is inferred based on the extracted features, and the inferred pose is output. In addition, the feature extractor may output the extracted features. The feature extractor corresponding to each reference object is produced by learning the images of the reference object. The model data 132 may include, in addition to the feature extractors corresponding to each reference object, a feature extractor that can be commonly used for all reference objects, and this feature extractor may be used instead of the feature extractors corresponding to each reference object.
[0038] In addition, the feature extractor that can be commonly used for all reference objects may be such that, if images of one or more poses of an object are input, it can extract the features of the image and output a result indicating which reference object the object in the image corresponds to (furthermore, it can also output a result that does not correspond to any reference object).
[0039] In addition, as a posture recognition method for a feature extractor corresponding to a certain reference object, for example, there is the following method: Images of one or more postures of a registration target object and images of one or more postures of a reference object are respectively input into an autoencoder, and the features of each posture of the registration target object obtained are compared with the features of each posture of the reference object, and the posture with the closest features is returned as the recognition result. The model data 132 is not limited to a feature extractor using such a posture recognition method, and may also have any feature extractor that can output a posture if an image is input, which is made from learning data obtained by learning an image of a reference object.
[0040] In addition, in the above example, the feature extractor extracts the features of the image if an image is input, and estimates the posture based on the extracted features. However, it may also be separated into a feature extractor that only extracts the features of the image if an image is input, and a posture estimator that estimates the posture by inputting the features from the feature extractor.
[0041] Figure 2 FIG. is a block diagram showing a hardware configuration example of the model production device 100. The model production device 100 is configured by a computer having a processor 110, a memory 120, an auxiliary storage device 130, an input device 140, an output device 150, and a communication IF (Interface) 160, and connecting them through an internal communication line 170 such as a bus.
[0042] The processor 110 executes a program stored in the memory 120. The memory 120 includes a ROM (Read Only Memory) as a non-volatile storage element and a RAM (Random Access Memory) as a volatile storage element. The ROM stores programs that do not change (for example, BIOS (Basic Input / Output System)) and the like. The RAM is a high-speed and volatile storage element such as a DRAM (Dynamic Random Access Memory), and temporarily stores the program executed by the processor 110 and the data used during the execution of the program.
[0043] The auxiliary storage device 130 is a large-capacity and non-volatile storage device such as a magnetic storage device (HDD (Hard Disk Drive)) or a flash memory (SSD (Solid State Drive)), and stores the program executed by the processor 110 and the data used during the execution of the program. That is, the program is read from the auxiliary storage device 130 and loaded into the memory 120, and is executed by the processor 110.
[0044] The input device 140 is a device such as a keyboard or a mouse that accepts input from an operator. The output device 150 is a device such as a display device or a printer that outputs the execution result of a program in a form recognizable by the operator. The communication IF 160 is a network interface device that controls communication with other devices according to a specified protocol.
[0045] The program executed by the processor 110 is provided to the model production device 100 via a removable medium (such as a CD-ROM or a flash memory) or a network and is stored in the non-volatile auxiliary storage device 130 as a non-temporary storage medium. Therefore, the model production device 100 can have an interface for reading data from a removable medium.
[0046] The model production device 100 is a computer system physically constituted on one computer or on a plurality of computers logically or physically constituted, and can operate either in a single thread on the same computer or on a virtual computer constructed on a plurality of physical computer resources. For example, the model production device 100 may not be a single computer but may be divided into: a teaching object registration device, which is a computer for registering a teaching object and an identification method for object identification; and a determination device, which is a computer for determining whether a certain object is a teaching object using the set identification method.
[0047] The processor 110 includes, for example, an image acquisition unit 111, an identification unit 112, an identification result comparison unit 113, a model correction unit 114, and an output unit 115, each of which serves as the above-mentioned functional unit.
[0048] For example, the processor 110 functions as the image acquisition unit 111 by operating in accordance with the image acquisition program loaded into the memory 120, and functions as the identification unit 112 by operating in accordance with the identification program loaded into the memory 120. The relationship between the program and the functional unit is the same for the other functional units included in the processor 110.
[0049] In addition, part or all of the functions of the functional units included in the processor 110 can also be implemented by hardware such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array).
[0050] The auxiliary storage device 130 holds, for example, the above-mentioned image data 131 and model data 132. In addition, part or all of the information stored in the auxiliary storage device 130 can be stored in the memory 120 or in an external database or the like connected to the model production device 100.
[0051] In addition, in this embodiment, the information used by the model manufacturing apparatus 100 is not dependent on the data structure, and it can be represented in any data structure. In this embodiment, the information is represented in a table form, but for example, it can be information stored in a data structure appropriately selected from a list, a database, or a queue.
[0052] Figure 3 is an example of an imaging system that captures an image of the registration target object 20 provided to the model manufacturing apparatus 100. The imaging system includes, for example, a camera 10, a rotating table 30, and a terminal 200. The camera 10 captures the registration target object 20. For example, an arm 11 is attached to the camera 10, and by operating the arm 11, the camera 10 can capture images from various positions and angles. The pose of an object represents the angle of the object as seen from the camera 10 and is determined by the relative positional relationship between the object and the camera.
[0053] The registration target object 20 is mounted on the rotating table 30. By rotating the rotating table 30 or operating the arm 11, the camera 10 can capture the registration target object 20 in various poses. The terminal 200 is a computer connected to the camera 10. The terminal 200 controls the shooting of the camera 10 and the operation of the arm 11. In addition, the terminal 200 acquires the image of the registration target object 20 captured by the camera 10. In addition, by controlling the operation of the rotating table 30 through the terminal 200, the camera 10 can capture images of the registration target object 20 in multiple poses.
[0054] In addition, although not shown in Figure 3 , the terminal 200 is connected to the model manufacturing apparatus 100 and sends the acquired image of the registration target object 20 to the model manufacturing apparatus 100. The image acquisition unit 111 of the model manufacturing apparatus 100 saves the received image in the image data 131. In addition, the terminal 200 can also control the camera 10, the arm 11, and the rotating table 30 according to an instruction from the image acquisition unit 111 of the model manufacturing apparatus 100.
[0055] In addition, the model manufacturing apparatus 100 and the terminal 200 can be integrated. In addition, the camera 10 can be built into the model manufacturing apparatus 100. In this case, shooting is performed according to an instruction from the image acquisition unit 111.
[0056] In addition, different from the example of Figure 3 , for example, multiple cameras 10 arranged on a spherical surface (or a hemispherical surface, etc.) centered on the registration target object 20 can capture images of the registration target object 20 in multiple poses. In addition, instead of the arm 11, a camera 10 fixed to a robot hand or the like can capture images of the registration target object 20 in multiple poses by operating the robot hand or the like.
[0057] Figure 4This is a flowchart showing an example of a model production process for creating a 3D model of the object 20 to be registered. The image acquisition unit 111 acquires images of one or more poses of the object 20 to be registered and information on the poses (S41). The model production device 100 performs the processes of steps S43 to S45 for the images of each pose (S42).
[0058] The recognition unit 112 obtains a feature extractor for recognizing the pose of the reference object from the model data 132, inputs the image of the pose of the object to be registered into the feature extractor, and outputs the pose, thereby recognizing the pose (S43). In addition, in step S43, either a feature extractor selected by the user can be used, or a feature extractor corresponding to the reference object that is closest in features to the object to be registered (for example, the reference object with the smallest squared distance between feature amounts) can be used. Among them, the feature extractor used in multiple executions of step S43 is the same. In addition, when the model data 132 includes a feature extractor that can be commonly corresponded to all reference objects, this feature extractor can also be used in step S43. The recognition result comparison unit 113 determines whether the pose of the object to be registered is the same as the pose recognized in step S43 (recognition success or failure) (S44).
[0059] When the recognition result comparison unit 113 determines that the pose of the object to be registered is the same as the pose recognized in step S43 (S44: Yes), it returns to step S42 and performs the processes of steps S43 to S45 for the next pose. Among them, when the processing for all poses is completed, the model production process is ended.
[0060] When the recognition result comparison unit 113 determines that the pose of the object to be registered is different from the pose recognized in step S43 (S44: No), the model correction unit 114 obtains the 3D model of a certain reference object from the model data 132, and creates the 3D model of the object to be registered by correcting the obtained 3D model (S45). The details of step S45 will be described later.
[0061] Figure 5 This is a flowchart showing an example of the 3D model correction process in step S45. The model correction unit 114 determines whether the 3D model correction process for creating the 3D model of the object to be registered is the first 3D model correction process (that is, whether it is the first time for step S45 for the object to be registered) (S51). When the model correction unit 114 determines that this 3D model correction process is the 3D model correction process after the second time (S51: No), it transfers to step S54 described later.
[0062] When it is determined that the 3D model correction process is the first 3D model correction process (S51: Yes), the model correction unit 114 acquires the 3D model from the model data 132 (S52). Specifically, for example, the model correction unit 114 acquires the 3D model of a reference object selected by the user of the model production device 100 from the model data 132. In addition, it may be that, for example, when the type to which the reference object belongs is given, the model correction unit 114 acquires the 3D models of all reference objects belonging to that type from the model data 132, and uses the average model of the acquired models as the 3D model acquired in step S52.
[0063] The model correction unit 114 registers a copy of the 3D model acquired in step S52 as the 3D model of the object to be registered in the model data 132 (S53). The model correction unit 114 corrects the 3D model of the object to be registered based on the image of the pose of the object to be registered (S54). Details of the method for correcting the 3D model will be described later.
[0064] The model correction unit 114 overwrites and registers the corrected 3D model as the 3D model of the object to be registered in the model data 132 (S55), and ends the 3D model correction process.
[0065] Figure 6 It is an explanatory diagram showing a specific example of the process for determining whether there is 3D model correction. In Figure 6 In the example of (a), if an image of the pose θ1 of the reference object A is input to the feature extractor A created by learning the image of the reference object A, the pose θ1 is output. In addition, if an image of the pose θ2 of the reference object A is input to the feature extractor A, the pose θ2 is output.
[0066] In Figure 6 In the example of (b), if an image of the pose θ1 of the object B to be registered is input to the feature extractor A, the pose θ1 is output, and if an image of the pose θ2 of the object B to be registered is input to the feature extractor A, the pose θ3 is output. That is, for the pose θ1 of the object B to be registered, the 3D model correction process in step S45 is not required, and for the pose θ2 of the object B to be registered, since a different pose θ3 is output, the 3D model correction process in step S45 is required.
[0067] Figure 7 It is an explanatory diagram showing a detailed example of the 3D model correction process in step S35. Hereinafter, an example in which the image of the object to be registered is RGB will be described. In Figure 7In the example, the model correction unit 114 determines that the local region 71 of the image of the posture θ1 of the object to be registered and the local region 72 of the 3D model corresponding to the local region 71 are not similar (for example, the similarity of the feature amounts of the local region 71 and the local region 72 is below a specified value (for example, the distance is above a specified value)).
[0068] When comparing the local region 71 with the local region 72, the local region 71 is composed of two faces, while the local region 72 is composed of one face. Therefore, the model correction unit 114 adds a vertex 73 to the local region 72 of the 3D model to increase the number of faces. The model correction unit 114 moves the added vertex 73 to make the local region 72 similar to or identical to the local region 71.
[0069] In this way, in Figure 7 the example, the model correction unit 114 refines the meshes of different regions in the 3D model and corrects the different regions to similar or identical regions. In addition, the model correction unit 114 can either move other vertices after deleting the vertices of the local region 72 according to the difference between the local region 72 and the local region 71, or just move a certain vertex of the local region 72.
[0070] In addition, when the model correction unit 114 refines the mesh of the 3D model in this way, for example, the number of vertices or the topology of the mesh is automatically changed by using a neural network to generate the mesh.
[0071] Furthermore, for example, it may also be that when the 3D model obtained in step S52 is the 3D model itself of a certain reference object, the image acquisition unit 111 acquires an image (for example, an image with higher resolution or a magnified image) in which the vicinity of the local region 72 of the reference object is photographed in more detail, and the model correction unit 114 also uses the acquired image to correct the 3D model in step S55 and then performs the above-mentioned mesh refinement.
[0072] In addition, when the model correction unit 114 obtains the average model of the 3D models of the same type of reference object in step S52, it can also refine the mesh of the average model and correct the average model in the same way as the above method. In addition, the model correction unit 114 can also acquire an image of the same type of reference object from the image data 131 in step S52, construct a 3D model based on the average image that is the average of the acquired images, and use it as the average model.
[0073] In addition, when the 3D model obtained in step S52 is the 3D model itself of a certain reference object, the model correction unit 114 can also create the 3D model of the registration target object by reconstructing the 3D model using an image group in which the images of the poses that failed to be recognized in step S44 in the images of each pose of the reference object are replaced with the images of the registration target object.
[0074] In addition, when the image of the registration target object is an RGB-Depth image, the model correction unit 114 corrects the 3D model by integrating the grid obtained by meshing the camera point group obtained from the image with the 3D model obtained in step S52. Further, if the image of the reference object is also an RGB-Depth image, the model correction unit 114 can also correct the 3D model by replacing the camera point group obtained from the image of the pose of the reference object corresponding to the 3D model with the camera point group obtained from the image of the registration target object.
[0075] In addition, in the present embodiment and the embodiments described later, when the 2D model of the reference object is stored in the model data 132, the model production device 100 can also correct the 2D model of the reference object to create the 2D model of the registration target object.
[0076] For example, when the 2D model obtained in step S52 and copied in step S53 by the model correction unit 114 is a 2D model composed of the images of the reference object, the 2D model is corrected by replacing the image of the pose (viewpoint) of the 2D model with the image of the registration target object. In addition, when the 2D model is a 2D model composed of one image of the reference object, the 2D model is corrected by replacing the image with the image of the registration target object.
[0077] In addition, for example, when the 2D model obtained in step S52 and copied in step S53 by the model correction unit 114 is a 2D model created based on local features such as edges and SIFT (Scale Invariant Feature Transform) in the images of the reference object, the local features are obtained from the image of the pose (viewpoint) of the 2D model, and the 2D model is corrected by replacing the local features of the 2D model with the obtained local features. In addition, when the 2D model is a 2D model composed of one image of the reference object, the 2D model is corrected by replacing the local features of the image with the local features of the registration target object.
[0078] In addition, in the case where noise is included in the image of the object to be registered, the model correction unit 114, for example, infers the contour of the object to be registered from the image, and corrects the 2D model by one of the above methods.
[0079] Through the above processing, the model production device 100 of the present embodiment produces a 3D model of the object to be registered by correcting only the part that affects the recognition performance of the feature extractor for the 3D model of the reference object. Therefore, a 3D model that reflects the local information of the object to be registered that affects the recognition performance can be produced with a small amount of data and processing volume.
[0080] Embodiment 2
[0081] In this embodiment, another example of the details of the model correction process will be described. In the following embodiments, the differences from Embodiment 1 will be described, and the descriptions repeated in Embodiment 1 will be omitted. Figure 8 It is an explanatory diagram showing a specific example of the process for determining whether to correct the 3D model.
[0082] Similar to Figure 6 the example of (b), if an image of the pose θ1 of the object to be registered B is input to the feature extractor A, the pose θ1 is output, and if an image of the pose θ2 of the object to be registered B is input to the feature extractor A, the pose θ3 is output. That is, for the pose θ1 of the object to be registered B, the correction process of the 3D model in step S45 is not required, while for the pose θ2 of the object to be registered B, since a different pose θ3 is output, the correction process of the 3D model in step S45 is required.
[0083] In addition, it is assumed that the recognition unit 112 determines that the local region 81 of the reference object obtained by the feature extractor and the local region 82 of the reference object are not similar (for example, the similarity of feature amounts is below a specified value).
[0084] At this time, the model correction unit 114 instructs the image acquisition unit 111 to acquire an image that more detailedly captures the vicinity of the local region 82 of the model of the object to be registered that is determined to require correction (for example, an image with a higher resolution or a magnified image). For example, the image acquisition unit 111 instructs the terminal 200 to capture the image, and acquires the image from the terminal 200. The model correction unit 114 uses the acquired image information to perform the model correction in step S54.
[0085] In Figure 8 this process, the model correction unit 114 corrects the 3D model based on the image near the local region (difference region) of the object to be registered that is not similar to the reference object. Therefore, a 3D model that reflects the details of the difference region of the registration reference object can be produced.
[0086] Figure 9 This is an explanatory diagram showing a specific example of the process for determining whether to correct a 3D model. Similar to the example of Figure 8 , if an image of the pose θ1 of the object B to be registered is input to the feature extractor A, the pose θ1 is output, and if an image of the pose θ2 of the object B to be registered is input to the feature extractor A, the pose θ3 is output. That is, for the pose θ1 of the object B to be registered, the correction process of the 3D model in step S45 is not required, while for the pose θ2 of the object B to be registered, since a different pose θ3 is output, the correction process of the 3D model in step S45 is required.
[0087] In addition, it is assumed that the recognition unit 112 determines that the local region 81 of the reference object obtained by the feature extractor and the local region 82 of the reference object are not similar (for example, the similarity of feature amounts is below a specified value).
[0088] At this time, the output unit 115 outputs a local region designation screen 90 to the output device 150. The local region designation screen 90 includes, for example, an object image display region 91, a local region change button 92, a save button 93, and a cancel button 94.
[0089] The local region designation screen 90 displays an image of the pose θ2 of the object B to be registered (i.e., the input image when an incorrect pose is output) and a display indicating the local region (the dotted ellipse in the figure). In addition, in the local region designation screen 90, for example, at the instruction of the user, instead of or in addition to the image of the pose θ2 of the object B to be registered, an image of the pose θ2 of the reference object (i.e., the image of the pose that should be correctly output for the reference object) can be displayed so that the user can easily grasp the dissimilar region.
[0090] The local region change button 92 is a button for changing the range of the local region. For example, if the local region change button 92 is selected, the display indicating the local region in the local region designation screen 90 becomes a state that can be changed by the user's input. The save button 93 is a button for saving the changed local region. If the save button 93 is selected, the model correction unit 114 uses the image information of the changed local region to perform the model correction in step S54.
[0091] The cancel button 94 is a button for ending without changing the local region. If the cancel button 94 is selected, the model correction unit 114 uses the image information of the local region before the change to perform the model correction in step S54.
[0092] The model correction unit 114 instructs the image acquisition unit 111 to acquire an image (for example, an image with a higher resolution or a magnified image) that more detailedly captures the vicinity of the partial region determined by the partial region designation screen 90, which is the posture of the model of the object to be registered and is determined to require correction. For example, the image acquisition unit 111 instructs the terminal 200 to capture this image and acquires this image from the terminal 200. The model correction unit 114 uses the acquired image information to perform the model correction in step S54.
[0093] In Figure 9 the process, the model correction unit 114 corrects the 3D model based on the image near the partial region (difference region) selected by the user, so a 3D model that reflects the details of the registration reference object, especially the difference region that is difficult to be recognized by the feature extractor, can be created.
[0094] Embodiment 3
[0095] This embodiment shows another example of the 3D model selection process in step S52. Figure 10 It is an explanatory diagram showing an example of the 3D model selection process in step S52. The model correction unit 114 acquires images of the object to be registered and multiple reference objects (for example, multiple reference objects selected by the user or all reference objects) from the image data 131, and inputs the acquired images to the feature extractors corresponding to the multiple reference objects respectively.
[0096] In addition, the model correction unit 114 may acquire images of a certain posture (one or more identical postures) of the object to be registered and multiple reference objects and input them to the feature extractor, or may acquire images of all postures of the object to be registered and multiple reference objects and input them to the feature extractor.
[0097] The model correction unit 114 calculates the similarity with the object to be registered for each of the multiple reference objects based on the features extracted by the feature extractor. The cosine similarity and the squared distance between feature quantities are both examples of the similarity calculated by the model correction unit 114. The model correction unit 114 determines the reference object with the highest calculated similarity as the similar object and acquires the 3D model of the similar object from the model data 132.
[0098] In Figure 10 the example, the similarity between the object to be registered B and the reference object A is 0.6, and the similarity between the object to be registered B and the reference object X is 0.4. Therefore, the model correction unit 114 uses the reference object A as the similar object and acquires the 3D model of the reference object A from the model data 132.
[0099] In Figure 10In the processing, the model correction unit 114 selects the 3D model of the reference object with a high similarity to the object to be registered. Therefore, an appropriate 3D model can be selected as the correction object, and the processing amount required for correcting the 3D model is likely to be reduced.
[0100] Figure 11 FIG. is an explanatory diagram showing an example of the 3D model selection process in step S52. Similar to the example of Figure 10 The model correction unit 114 calculates the similarity between the object to be registered and each of the multiple reference objects. When the model correction unit 114 determines that all the calculated similarities are below a specified threshold, in step S52, it does not select a model but aborts the model correction process and newly creates a 3D model of the object to be registered.
[0101] In Figure 11 In the example, the similarity threshold is 0.5. The similarity between the object to be registered B and the reference object A is 0.4, which is below the threshold, and the similarity between the object to be registered B and the reference object X is 0.3, which is also below the threshold. Therefore, the model correction unit 114 does not select the 3D model of the reference object but newly creates a 3D model of the object to be registered B.
[0102] In Figure 11 In the processing, when there is no reference object with a high similarity to the object to be registered, the model correction unit 114 newly creates a 3D model of the object to be registered. Therefore, an inappropriate 3D model as the correction object will not be selected. In addition, if the model correction unit 114 selects the 3D model of the reference object with a high similarity to the object to be registered and corrects this 3D model to create a 3D model of the object to be registered, the processing amount may instead increase or the recognition performance may become insufficient. By performing the processing of Figure 11 the model correction unit 114 can suppress the occurrence of such a situation.
[0103] Embodiment 4
[0104] This embodiment shows another example of the model creation process. The model creation device 100 of this embodiment corrects the feature extractor according to the recognition result of the object to be registered. Figure 12 FIG. is a flowchart showing an example of the model creation process of this embodiment.
[0105] When the recognition result comparison unit 113 determines that the posture of the object to be registered is the same as the posture recognized in step S43 (S44: Yes), or after the model correction process in step S45 ends, the recognition unit 112 corrects the feature extractor based on the image of the object to be registered (S46). Hereinafter, a specific example of the correction process of the feature extractor will be described.
[0106] Figure 13 This is an explanatory diagram showing an example of the correction process of the feature extractor. Similar to Figure 8 the example of, if an image of the pose θ1 of the registration target object B is input to the feature extractor A, the pose θ1 is output, and if an image of the pose θ2 of the registration target object B is input to the feature extractor A, the pose θ3 is output.
[0107] At this time, the recognition unit 112 obtains an image of the pose θ2 of the registration target object (that is, an image of the registration target object of the pose that should be correctly output from the feature extractor) from the image data 131, associates the obtained image with the pose θ2, and causes the feature extractor A to perform additional learning, thereby overwriting the feature extractor A in the model data 132. Thus, the recognition unit 112 can quickly learn an image of a pose with low recognition accuracy in the feature extractor of the registration target object.
[0108] In addition, when the feature extractor and the pose estimator are separated, the recognition unit 112 causes the pose estimator to perform the above additional learning, and further causes the feature extractor to perform additional learning on an image of the pose θ2 of the registration target object (that is, an image of the registration target object of the pose that should be correctly output from the pose estimator), and overwrites the feature extractor in the model data 132.
[0109] Furthermore, in the generation of the 3D model of the next registration target object, the recognition unit 112 uses the overwritten feature extractor A to perform the process of outputting the pose of the registration target object in step S52. Thus, pose estimation using the feature extractor A that reflects the features of the previous registration target object is performed, so the processing amount for the model production process of the registration target object having features similar to those of the previous registration target object is reduced.
[0110] In addition, when there are not enough images of the pose θ2 of the registration target object in the image data 131 (for example, only less than a specified number of images), the image acquisition unit 111 is instructed to acquire a specified number of images of the pose θ2 of the registration target object. For example, the image acquisition unit 111 instructs the terminal 200 to capture the specified number of images of the registration target object, and acquires the specified number of images of the registration target object from the terminal 200.
[0111] Figure 14 This is an explanatory diagram showing an example of the correction process of the feature extractor. Similar to Figure 8 the example of, if an image of the pose θ1 of the registration target object B is input to the feature extractor A, the pose θ1 is output, and if an image of the pose θ2 of the registration target object B is input to the feature extractor A, the pose θ3 is output.
[0112] At this time, the recognition unit 112 obtains an image of the posture θ3 of the object to be registered from the image data 131 (i.e., an image of the object to be registered with the posture erroneously output by the feature extractor), associates the obtained image with the posture θ3, and causes the feature extractor A to perform additional learning, and overwrites the feature extractor A in the model data 132. Thus, the recognition unit 112 can quickly learn an image of a posture with low recognition accuracy in the feature extractor of the object to be registered.
[0113] In addition, when the feature extractor and the posture estimator are separated, the recognition unit 112 causes the posture estimator to perform the above-described additional learning, and further causes the feature extractor to perform additional learning on an image of the posture θ3 of the object to be registered (i.e., an image of the object to be registered with the posture erroneously output by the feature extractor), and overwrites the feature extractor in the model data 132.
[0114] Furthermore, in the generation of the 3D model of the object to be registered next time, the recognition unit 112 uses the overwritten feature extractor A to perform the process of outputting the posture of the object to be registered in step S52. Thus, posture estimation using the feature extractor A reflecting the features of the previous object to be registered is performed, so that the processing amount for the model production process of the object to be registered having features close to those of the previous object to be registered is reduced.
[0115] In addition, when there are not enough images of the posture θ3 of the object to be registered in the image data 131 (for example, only a number of images equal to or less than a specified number), the image acquisition unit 111 is instructed to acquire a specified number of images of the posture θ3 of the object to be registered. For example, the image acquisition unit 111 instructs the terminal 200 to capture the specified number of images of the object to be registered, and acquires the specified number of images of the object to be registered from the terminal 200.
[0116] In addition, for example, the recognition unit 112 may also cause the feature extractor to perform additional learning on both an image of the posture θ2 of the object to be registered (i.e., an image of the object to be registered with the posture that should be correctly output by the feature extractor) and an image of the posture θ3 of the object to be registered (i.e., an image of the object to be registered with the posture erroneously output by the feature extractor).
[0117] In addition, the present invention is not limited to the above-described embodiments, and includes various modification examples. For example, the above-described embodiments are described in detail for easy understanding of the present invention, and are not limited to necessarily having all the configurations described. In addition, a part of the configuration of a certain embodiment may be replaced with the configuration of another embodiment, and in addition, the configuration of another embodiment may be added to the configuration of a certain embodiment. In addition, for a part of the configuration of each embodiment, addition, deletion, and replacement of other configurations can be performed.
[0118] In addition, a part or all of the above-described components, functions, processing units, processing elements, etc. can also be implemented by hardware, for example, through integrated circuit design. In addition, the above-described components, functions, etc. can also be implemented by software by a processor interpreting and executing programs that implement respective functions. Information such as programs, tables, and files that implement respective functions can be placed in a recording device such as a memory, a hard disk, an SSD (Solid State Drive), or a recording medium such as an IC card, an SD card, or a DVD.
[0119] In addition, regarding control lines and information lines, only parts considered necessary for explanation are shown, and not all control lines and information lines are necessarily shown on the product. In fact, it can also be considered that almost all components are interconnected.
Claims
1. A model making device that makes a model representing the shape of a registration target object, wherein, it includes a processor and a memory, the above-mentioned memory holds: images of one or more poses of the above-mentioned registration target object; and a reference model representing the shape of a reference object, the above-mentioned processor obtains information representing the characteristics of the first pose of the above-mentioned registration target object, when the above-mentioned processor determines, based on a specified first condition, that the shape of the first pose represented by the above-mentioned reference model is not similar, the above-mentioned reference model is corrected based on the information representing the above-mentioned characteristics to make a model representing the shape of the above-mentioned registration target object, the above-mentioned memory holds a feature extractor that is made by learning the image of the above-mentioned reference object and outputs a pose when an image is input, when the above-mentioned processor inputs the first image of the first pose of the above-mentioned registration target object to the above-mentioned feature extractor and outputs a second pose different from the first pose, the above-mentioned reference model is corrected based on the information representing the above-mentioned characteristics to make a model representing the shape of the above-mentioned registration target object.
2. The model making device according to claim 1, wherein, the above-mentioned memory holds: reference models representing the shapes of a plurality of the above-mentioned reference objects; and images of the above-mentioned one or more poses of the above-mentioned plurality of reference objects, when the above-mentioned processor inputs the first image and outputs the second pose to the above-mentioned feature extractor, the first image of the above-mentioned registration target object and the images of the first poses of the above-mentioned plurality of reference objects are input to the above-mentioned feature extractor, and the similarity between the above-mentioned registration target object and each of the above-mentioned plurality of reference objects is calculated, based on the information representing the characteristics of the first pose of the above-mentioned registration target object, the reference model of the reference object with the highest calculated similarity is corrected to make a model representing the shape of the above-mentioned registration target object.
3. The model making device according to claim 1, wherein, the above-mentioned memory holds: reference models representing the shapes of a plurality of the above-mentioned reference objects; and images of the above-mentioned one or more poses of the above-mentioned plurality of reference objects, the above-mentioned feature extractor is made by learning the images of the above-mentioned plurality of reference objects, when the first image is input to the above-mentioned feature extractor and the second pose is output, the first image is compared with the images of the first poses of the above-mentioned plurality of reference objects, and the similarity between the above-mentioned registration target object and each of the above-mentioned plurality of reference objects is calculated, when all of the calculated similarities are below a specified threshold value, the above-mentioned reference model is not corrected, and a new model representing the shape of the above-mentioned registration target object is made.
4. The model making device according to claim 1, wherein, the above-mentioned memory holds a second image different from the first image of the first pose of the above-mentioned registration target object, when the above-mentioned processor inputs the first image and outputs the second pose to the above-mentioned feature extractor, Cause the above-mentioned feature extractor to learn the above-mentioned second image, and save the learned above-mentioned feature extractor to the above-mentioned memory.
5. The model manufacturing device according to claim 4, wherein, the above-mentioned feature extractor includes a feature extraction unit that extracts features of an image, and a pose estimation unit that outputs a pose based on the features extracted by the above-mentioned extraction unit, when the above-mentioned first image is input to the above-mentioned feature extractor and the above-mentioned second pose is output, cause the above-mentioned pose estimation unit to learn the above-mentioned second image.
6. The model manufacturing device according to claim 1, wherein, the above-mentioned memory holds a third image of the above-mentioned second pose of the above-mentioned registration target object, when the above-mentioned first image is input to the above-mentioned feature extractor and the above-mentioned second pose is output by the above-mentioned processor, cause the above-mentioned feature extractor to learn the above-mentioned third image, and save the learned above-mentioned feature extractor to the above-mentioned memory.
7. The model manufacturing device according to claim 6, wherein, the above-mentioned feature extractor includes a feature extraction unit that extracts features of an image, and a pose estimation unit that outputs a pose based on the features extracted by the above-mentioned extraction unit, when the above-mentioned first image is input to the above-mentioned feature extractor and the above-mentioned second pose is output, cause the above-mentioned pose estimation unit to learn the above-mentioned third image.
8. The model manufacturing device according to claim 1, wherein, the above-mentioned memory holds information representing features of a local area of an image of the above-mentioned reference object, when the above-mentioned processor determines, based on the above-mentioned first condition, that the shape of the above-mentioned first pose represented by the above-mentioned reference model is not similar, in the above-mentioned registration target object and the above-mentioned reference object, determine a local area where the information representing features is not similar based on a specified second condition, acquire a detailed image of the determined above-mentioned local area of the above-mentioned registration target object, acquire information representing features of the above-mentioned detailed image, based on the information representing features of the above-mentioned detailed image, correct the above-mentioned reference model to manufacture a model representing the shape of the above-mentioned registration target object.
9. The model manufacturing device according to claim 1, wherein, it is provided with a display device, the above-mentioned memory holds information representing features of a local area of an image of the above-mentioned reference object, when the above-mentioned processor determines, based on the above-mentioned first condition, that the shape of the above-mentioned first pose represented by the above-mentioned reference model is not similar, display an image of the above-mentioned first pose of the above-mentioned registration target object on the above-mentioned display device, accept the designation of a local area, acquire information representing features of the designated above-mentioned local area, based on the information representing features of the designated above-mentioned local area, correct the above-mentioned reference model to manufacture a model representing the shape of the above-mentioned registration target object.
10. The model manufacturing device according to claim 1, wherein, the above-mentioned reference model is a three-dimensional model that defines the shape of the above-mentioned reference object by a mesh and vertices, When it is determined based on the above first condition that the shape of the first pose represented by the above reference model is not similar, the above processor increases or decreases the vertices in the above reference model based on the shape represented by the image of the above registration target object in the above first pose, moves the increased or decreased above vertices, and thereby corrects the above reference model.
11. The model production device according to claim 1, wherein, the above memory holds: images of one or more poses of each of the above plurality of reference objects; and type information indicating the types to which the above registration target object and the above plurality of reference objects belong, the above processor refers to the above type information and determines a reference object belonging to the same type as the above registration target object, the above processor produces an average model that represents the shape of an image obtained by averaging the images of the determined above reference objects, when it is determined based on the above first condition that the shape of the first pose represented by the above reference model is not similar, the above processor corrects the above average model based on the information representing the above feature, and produces a model representing the shape of the above registration target object.
12. A model production method, which is a method for a model production device to produce a model representing the shape of a registration target object, wherein, the above model production device holds: images of one or more poses of the above registration target object; and a reference model representing the shape of a reference object, in the above method, the above model production device obtains information representing the feature of the first pose of the above registration target object, when it is determined based on a prescribed first condition that the shape of the first pose represented by the above reference model is not similar, the above model production device corrects the above reference model based on the information representing the above feature, and produces a model representing the shape of the above registration target object, the above model production device holds a feature extractor that is produced by learning the images of the above reference objects and outputs a pose when an image is input, in the above method, when the above model production device inputs the first image of the first pose of the above registration target object to the above feature extractor and outputs a second pose different from the above first pose, the above model production device corrects the above reference model based on the information representing the above feature, and produces a model representing the shape of the above registration target object.
Citation Information
Patent Citations
Picked-up image processor and picked-up image processing method
JP1996233556A
Image forming apparatus, its control method and program
JP2019215673A
Apparatus for modeling three dimensional information
US5819016A