Object attitude estimation device and attitude estimation method

By generating depth and normal images through convolutional neural networks, performing super-resolution processing and feature extraction, the accuracy problem of object pose estimation under low-resolution or long-distance shooting is solved, and high-precision object pose and position estimation is achieved.

CN120997290APending Publication Date: 2025-11-21TOYOTA JIDOSHA KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510624215.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-20
Filing Date
2025-05-15
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately estimate the pose of objects when they are photographed at low resolution or far from the camera, especially for small objects, where insufficient feature data or incorrect estimation of the object's shape may occur.

Method used

Convolutional neural networks are used to parse image data, generate depth and normal images, improve object resolution through super-resolution processing, extract features, generate corrected images and estimate object state and position, and fit the actual data using the Gauss-Newton method to achieve high-precision attitude estimation.

Benefits of technology

In low-resolution or long-distance shooting situations, it can accurately estimate the pose and position of objects, thereby improving the learning effect of the learning model and the accuracy of the output results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997290A_ABST
    Figure CN120997290A_ABST
Patent Text Reader

Abstract

The invention provides an object attitude estimation device and an attitude estimation method, which can estimate the attitude of an object with high precision even if the object included in an input image is small and the resolution of the object is low. A posture estimation device (1) estimates the posture of an object on the basis of image data, acquires a depth image, acquires a normal image on the basis of the depth image, generates a super-resolution image of the object cut from the depth image, extracts feature quantities of the object included in the image data, the super-resolution image, and the normal image, and estimates the posture of the object on the basis of the extracted feature quantities. A correction image in which the shape of the object is restored on the basis of feature quantities extracted from the super-resolution image and the normal image is generated, a parameter relating to the state of the object is estimated on the basis of the correction image, the position of the object is estimated from the image data, and the attitude of the object is estimated on the basis of the parameter relating to the state of the object and the position of the object.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a device and a method for estimating a posture by analyzing an object included in image data and video data. BACKGROUND

[0002] In recent years, practicalization of a technology for efficiently performing learning, estimation, recognition, determination, and the like using advanced technologies such as artificial intelligence (AI), information and communication technology (ICT), and the like is continuously progressing. Among them, as a method of learning performed by AI, there is machine learning. With machine learning, a machine (computer) learns by itself using a large amount of data provided, and based on a learning result (learned model) thereof, performs optimization of output data with respect to input data, and performs estimation, prediction, and the like based on the output data. In such machine learning, there is a processing technology such as a convolutional neural network (CNN) that performs convolution processing based on input data. Machine learning is applied, for example, to various fields such as image recognition, speech recognition, natural language processing, and machine translation.

[0003] In Patent Literature 1, a technology for estimating a posture of an object using the above-described machine learning is described. In Patent Literature 1, a posture estimation and collation system of an object is disclosed, and the purpose thereof is to perform estimation and collation of a posture with high precision with respect to an image of an object regardless of conditions such as a posture of the object, brightness of the object, and the like. In the system of Patent Literature 1, a plurality of posture candidates are generated based on an input image. Based on the plurality of posture candidates, a plurality of comparison images in which conditions such as illumination are close to the input image are generated while a three-dimensional object model is projected onto a two-dimensional image. A feature amount (sharpness feature amount) that reflects sharpness of each of the generated comparison images is extracted. Then, a difference degree (weighted difference degree) in which the difference degree of the input image and the plurality of comparison images is weighted by the sharpness feature amount is calculated for each comparison image. Then, based on the plurality of weighted difference degrees, a comparison image closest to the input image is selected from the plurality of comparison images, and a best (closest) posture is estimated from the selected comparison image. In the system of Patent Literature 1, the sharpness feature amount that reflects the sharpness of the comparison image is considered, so a large difference degree is easily generated in a case where the posture and the object do not match between the three-dimensional object model and the input image, and thus the precision of the posture estimation is improved.

[0004] PRIOR ART DOCUMENTS

[0005] PATENT LITERATURE

[0006] Patent Literature 1: Japanese Patent No. 4692773 SUMMARY

[0007] PROBLEMS TO BE SOLVED BY THE INVENTION

[0008] In the system of Patent Literature 1, the pose is estimated based on an input image and a plurality of comparison images are generated, and the pose of an object is estimated from a feature amount of the comparison image, a difference degree between the input image and the comparison image, and the like. However, in a case where a camera is provided at a position far from an object, in a case where the object as a target is a small object, and the like, the object included in the input image is small or the resolution of the object is low. In such a case, it is possible that the feature amount obtained from the input image is less, or the shape of the object such as a concave-convex is erroneously estimated. Therefore, in the system of Patent Literature 1, even if the resolution is considered, it is possible that the pose of the object cannot be estimated with a desired accuracy according to the size, the resolution, and the like of the object represented in the input image.

[0009] The present application has been made in view of the above-described technical problems, and aims to provide a pose estimation device and a pose estimation method capable of estimating the pose of an object with high accuracy even in a case where the object included in an input image is small or the resolution of the object is low.

[0010] Means for solving the problems

[0011] To achieve the above object, the present application provides a pose estimation device of an object, which analyzes image data obtained by photographing an object using a convolutional neural network, and estimates the pose of the object included in the image data, characterized by comprising: a depth data acquisition section which acquires a depth image including depth data which is data related to the depth of the object included in the image data; a normal data acquisition section which acquires a normal image which is an image obtained by finding a normal vector of a surface of the object from the depth image; a super-resolution image generation section which generates a super-resolution image which is an image obtained by improving the resolution of the object cropped from the depth image; a feature amount extraction section which extracts a feature amount which quantitatively represents a feature of the object included in the image data, the super-resolution image, and the normal image; a corrected image generation section which generates a corrected image which is an image obtained by restoring the shape of the object based on the feature amount extracted from the super-resolution image and the normal image; a state estimation section which estimates a parameter related to the state of the object based on the corrected image; a position estimation section which estimates the position of the object from the image data; and a pose estimation section which estimates the pose of the object based on the parameter related to the state of the object estimated by the state estimation section and the position of the object estimated by the position estimation section.

[0012] Further, in the posture estimation device of the present application, the optimization section can be further provided, which performs fitting of output data of the position of the object estimated by the position estimation section and the posture of the object estimated by the posture estimation section and actual data related to the position and posture of the object by the Gauss-Newton method.

[0013] Further, in the posture estimation device of the present application, the correction image generation section can be configured to generate a synthetic feature quantity in which the feature quantity of the object extracted from the depth image and the feature quantity of the object extracted from the normal image are fused in a channel direction, reproduce the object based on the synthetic feature quantity by performing up-sampling processing using a copy value of the periphery to interpolate a missing portion, and generate the correction image.

[0014] On the other hand, the present application is a posture estimation method of an object, which analyzes image data obtained by photographing an object using a convolutional neural network, and estimates a posture of the object included in the image data, characterized by acquiring a depth image including depth data which is data related to a depth of the object included in the image data, acquiring a normal image which is an image obtained by calculating a normal vector of a surface of the object from the depth image, generating a super-resolution image which is an image obtained by increasing a resolution of the object cropped from the depth image, extracting a feature quantity which quantitatively represents a feature of the object included in the image data, the super-resolution image, and the normal image, generating a correction image which is an image obtained by reproducing a shape of the object based on the feature quantity extracted from the super-resolution image and the normal image, estimating a parameter related to a state of the object based on the correction image, estimating a position of the object from the image data, and estimating a posture of the object based on the estimated parameter related to the state of the object and the estimated position of the object.

[0015] Further, in the posture estimation method of the present application, the output data related to the estimated position of the object and the estimated posture of the object and the actual data related to the position and posture of the object can be fitted by the Gauss-Newton method.

[0016] Further, in the posture estimation method of the present application, a synthetic feature quantity in which the feature quantity of the object extracted from the depth image and the feature quantity of the object extracted from the normal image are fused in a channel direction can be generated, the object can be reproduced based on the synthetic feature quantity by performing up-sampling processing using a copy value of the periphery to interpolate a missing portion, and the correction image can be generated.

[0017] Further, a posture estimation system of an object, which analyzes image data obtained by photographing an object using a convolutional neural network, estimates a posture of the object included in the image data, characterized by comprising: a camera that acquires a depth image including depth data that is data related to a depth of the object; and a posture estimation device that estimates the posture of the object, the posture estimation device having: a normal data acquisition section that acquires a normal image that is an image obtained by finding a normal vector of a surface of the object from the depth image; a super-resolution image generation section that generates a super-resolution image that is an image obtained by increasing a resolution of the object cropped from the depth image; a feature quantity extraction section that extracts a feature quantity that quantitatively indicates a feature of the object included in the image data, the super-resolution image, and the normal image; a corrected image generation section that generates a corrected image that is an image obtained by restoring a shape of the object based on the feature quantity extracted from the super-resolution image and the normal image; a state estimation section that estimates a parameter related to a state of the object based on the corrected image; a position estimation section that estimates a position of the object from the image data; and a posture estimation section that estimates the posture of the object based on the parameter related to the state of the object estimated by the state estimation section and the position of the object estimated by the position estimation section.

[0018] Effects of Invention

[0019] In the object pose estimation apparatus and method according to embodiments of the present invention, a depth image and a normal vector image (normal image) of the object are generated based on image data obtained by taking an image of the object in a top-down view. Based on the generated depth image, a super-resolution image is generated, which performs super-resolution processing to improve the resolution of the object contained in the image data. Then, feature quantities of the object are extracted from the super-resolution image and the normal image, respectively. A corrected image is generated, in which the shape of the object is restored based on the extracted feature quantities. Then, a result is output, which estimates data related to the object's shape, such as its size, label (classification), rotation amount (rotation angle), and thickness, based on the object's feature quantities extracted from the corrected image. Additionally, data related to the object's position is estimated based on the image data taken of the object in a top-down view and input. Then, the object's pose is estimated based on the data related to the estimated shape of the object and the data related to the object's position. In other words, by creating a super-resolution image from the input image data (depth image) and generating a normal image of the object based on that super-resolution image, the shape of the object can be restored with high precision, thus enabling the estimation of the object's shape and the like with high accuracy. Furthermore, since the object's pose is estimated along with the position of the object obtained from image data obtained by taking an image of the object from a bird's-eye view, the object's pose can be estimated with high accuracy.

[0020] That is, for example, when an object must be positioned far from the camera due to the location of the camera, the object in the captured image data becomes smaller. In this case, the object's resolution is low and the number of object features is reduced, or the object's shape may be perceived as planar due to the increased distance between the locations where features are extracted, potentially resulting in insufficient feature data for estimating the object's pose based on the image data. Even in such cases, the pose estimation apparatus or method in the embodiments of the present invention can estimate the object's pose with high accuracy.

[0021] Furthermore, since the estimated shape, position, and other data of the object are highly accurate as described above, the learning model can be effectively trained by learning based on this data, and the accuracy of the output result, namely the estimated pose of the object, can be further improved. Attached Figure Description

[0022] Figure 1 This is an explanatory diagram illustrating the state of obtaining dynamic image data parsed by the attitude estimation device in the embodiments of the present invention.

[0023] Figure 2 This is an explanatory diagram illustrating the overall structure of the attitude estimation device in an embodiment of the present invention.

[0024] Figure 3 is a block diagram for explaining functional configuration of a posture estimation device in an embodiment of the present application.

[0025] Figure 4 is a graph showing an image of an object output by processing of image data by the functional configuration of the posture estimation device shown in Figure 3 (a) is a graph showing an image of a segment image obtained by detecting or extracting an object from the image data, (b) is a graph showing an image of a region of the object detected from the segment image converted into a rectangle, and (c) is a graph showing an image of a mask image obtained by masking other parts of the object than the segment region.

[0026] Figure 5 is a graph showing an image of an object output by processing of image data by the functional configuration of the posture estimation device shown in Figure 3 (a) is a graph showing an image of a cropped image in which only the object included in the mask image is cropped, (b) is a graph showing an image of a super-resolution image in which the object included in the cropped image is super-resolved, and (c) is a graph showing an image of a normal image in which an orientation of a surface of the object included in the cropped image is represented.

[0027] Figure 6 is a graph showing an image of an object output by processing of image data by the functional configuration of the posture estimation device shown in Figure 3 (a) is a graph showing an image of an image of the object input to a convolutional neural network in order to extract a feature quantity of the object included in the super-resolution image, (b) is a graph showing an image of an image of the object input to the convolutional neural network in order to extract a feature quantity of the object included in the normal image, and (c) is a graph showing an image of a correction image generated on the basis of the extracted feature quantity.

[0028] Figure 7 is a flowchart (procedure chart) showing an example of contents and procedures of an arithmetic process (control) used by the posture estimation device and the posture estimation method in the embodiment of the present application.

[0029] Figure 8 is a flowchart showing an example of contents and procedures of an arithmetic process used by the posture estimation device and the posture estimation method in the embodiment of the present application, and is a flowchart for explaining the processing and procedures of the flowchart shown in Figure 7

[0030] Figure 9 is a flowchart showing an example of contents and procedures of an arithmetic process used by the posture estimation device and the posture estimation method in the embodiment of the present application, and is a flowchart for explaining the processing and procedures of the flowchart shown in Figure 7 ​and Figure 8 the flowchart of the processes of the flowchart.

[0031] BRIEF DESCRIPTION OF DRAWINGS

[0032] 1 posture estimation device

[0033] 2 camera

[0034] 3 work site

[0035] 4 object

[0036] 5 first processing section

[0037] 6 second processing section

[0038] 7 image acquisition section

[0039] 8 image information preprocessing section

[0040] 9 super-resolution processing section

[0041] 10 normal image generation section

[0042] 11 corrected image generation section

[0043] 12 object information estimation section

[0044] 13 object position estimation section

[0045] 14 object posture estimation section

[0046] 15 output section DETAILED DESCRIPTION

[0047] Next, the present application will be described based on the illustrated embodiments. Note that the following described embodiments are merely one example of a case where the present application is embodied, and do not limit the present application.

[0048] In the object posture estimation device and the object posture estimation method in the embodiments of the present application, a convolutional neural network is used, and a convolution operation is performed based on a plurality of pieces of information (object data) collected, and a feature (a feature amount, a feature point) of the information source is extracted. For example, a feature amount is extracted from image data (object data) obtained by imaging an object such as a workpiece or a transported object, and the position and the posture of the object are estimated based on the feature amount. In the object posture estimation device and the object posture estimation method in the embodiments of the present application, the posture of an object (a workpiece) included in image data imaged by an imaging device is estimated by using a convolutional neural network. For example, the posture of an object such as a tool placed in a work site is estimated.

[0049] As Figure 1 and Figure 2As shown, in the attitude estimation device 1, the image data captured by the camera 2 is analyzed by the attitude estimation device 1, whereby the attitude of the object is estimated. In Figure 1 In the example shown, the camera 2 is arranged to capture the entire work site 3 or the object (workpiece) 4 and the like in an overhead manner, whereby the object 4 is disposed at a position relatively far from the camera 2.

[0050] The camera 2 is a camera device that captures a prescribed region in the work site 3 from above. The camera 2 is fixed to, for example, a ceiling or the like in the work site 3 and is arranged to capture the entire work site 3 in an overhead manner. The camera 2 can be configured similarly to the camera 2 known in the art, and can be configured to detect or estimate the depth or the depth of the object 4 from the captured image data.

[0051] For example, the camera 2 can be a stereo camera having two cameras and capable of calculating the distance to the object 4 from the parallax based on the image data captured by each camera, the distance between the cameras, and the focal length of each camera. Alternatively, the camera 2 can be a monocular camera as long as it is capable of sensing the depth. For example, the camera 2 can be configured by a camera that moves the camera 2 and detects the depth of the object 4 based on the image data before and after the movement and the difference in movement, a camera that detects the distance to the object 4 based on the time from when light is emitted to the object 4 to when the reflected light is detected, or a camera that emits a prescribed light pattern such as a stripe pattern to the object 4 and detects the distance to the object 4 based on the distortion of the prescribed light pattern.

[0052] The attitude estimation device 1 has a processor, a communication unit, a storage unit, and the like as main structures. The attitude estimation device 1 is configured to perform an operation in accordance with a prescribed program using data acquired from the outside and data stored in advance and the like, and to output the result of the operation as a control instruction signal. For example, the attitude estimation device 1 loads a program stored in a recording medium to a work area of the storage unit by the processor and executes, performs various controls by the execution of the program, and thereby performs a function in accordance with a prescribed purpose.

[0053] The processor is, for example, a CPU or a DSP. This processor is configured to control the attitude estimation device 1 and perform various information processing operations. It should be noted that the processor may also include a GPU (Graphics Processing Unit) capable of high-speed image processing. The storage unit includes, for example, RAM and ROM. As described above, a working area for the processor to execute programs is formed in the storage unit. Additionally, the storage unit may include, for example, an auxiliary storage unit such as an EPROM or a hard disk drive. This auxiliary storage unit may also include a so-called removable medium, which is a removable recording medium. Furthermore, the auxiliary storage unit freely stores various programs, data, and tables on the recording medium through reading or writing. It should be noted that the auxiliary storage unit may also store an operating system. The communication unit is a wireless communication circuit that connects to external communication devices, such as the aforementioned camera 2, in a manner capable of data communication using wireless communication. It should be noted that the attitude estimation device 1 is configured to perform machine learning processing based on neural networks, including a detection unit, a calculation unit, and a learning unit, using appropriate elements from the above-described components.

[0054] Next, use Figure 2 and Figure 3 The structure of the attitude estimation device 1 in the embodiment of the present invention will be described. In the attitude estimation device 1, the attitude of the object 4 contained in the image data can be estimated by parsing the image data as described above. Figure 2 This is a diagram illustrating the overall structure of the learning model or architecture for presuming the pose of object 4 as a concept. Figure 3 A block diagram representing the functional structure is shown.

[0055] like Figure 2 As shown, the attitude estimation device 1, which is used to estimate the attitude of object 4, includes: a first processing unit 5, which estimates the size of object 4, what kind of workpiece (label) object 4 is, and other characteristics of object 4 itself; and a second processing unit 6, which estimates the position of object 4 relative to camera 2. Moreover, it is configured to estimate the attitude of object 4 based on the characteristics of object 4 itself and the position of object 4 estimated by the first processing unit 5 and the second processing unit 6.

[0056] Next, use Figure 3 The specific contents of the first processing unit 5 and the second processing unit 6, as well as the specific functional structures used to estimate the posture of object 4, are explained. For example... Figure 3As shown, the posture estimation device 1 is provided with an image acquisition section 7, an image information preprocessing section 8, a super-resolution processing section 9, a normal image generation section 10, a corrected image generation section 11, an object information estimation section 12, an object position estimation section 13, an object posture estimation section 14, and an output section 15. Note that a file having data related to the position and posture of the object 4 or a file having data related to the position and posture of the object 4 is input to the posture estimation device 1 or a learning model in advance. In addition, a figure visualizing the image generated by the processing performed by these functional structures or the image of the object 4 output is shown in Figs. 14, 15, and 16. Figure 4 , Figure 5 and Figure 6

[0057] The image acquisition section 7 acquires image data or video data captured by the camera 2. The image acquisition section 7 can be configured to input arbitrary image data by a user at the time of acquisition of the image data, or can be configured to automatically take in from the camera 2.

[0058] The image information preprocessing section 8 performs processing for making it possible to acquire more information related to the object 4 from the acquired image data or easily extracting data related to the object 4. For example, in the image information preprocessing section 8, preprocessing such as normalization of the image data, adjustment of the size, cropping of the object 4, and the like is performed. Note that the normalization processing performed here is normalization processing performed as preprocessing of the image data, and is processing of adjusting the distribution of data by scaling the pixel value of the image to a prescribed range to adjust the mean and variance.

[0059] Specifically, in the image information preprocessing section 8, first, the depth information of the object 4 is acquired from the image data and information from the camera 2. In the image information preprocessing section 8, a depth image of the object 4 is generated based on the depth and the depth of field of the object 4 detected by the camera 2. Although the illustration is omitted, the depth image is an image in which data related to the three-dimensional spatial coordinates of the object 4 from the camera 2 as a starting point can be acquired by assigning a color to a distance value and arranging the output or changing the gradation of the color. Note that, Figure 2 The image included in the first processing section 5 shown is an example of a depth image.

[0060] ​Note that, by learning a depth image for learning in advance, a depth image can be generated based on input image data, and random noise is added to the generated depth image. The random noise is added to ensure robustness and consistency of the image data, and by learning in a state where such noise is applied to the depth image, the prediction performance for data that the learning model has not learned can be improved. In addition, such a structure that randomly adds noise to image data or performs a standardization process on image data uses a structure that is disclosed and provided (shared) through an API (Application Programming Interface).

[0061] In the image information preprocessing section 8, the objects 4 included in the thus generated depth image are recognized, and segmentation is performed in which a label is added to each object 4 on a pixel-by-pixel basis. By the segmentation, a segment region that recognizes a region occupied by the object 4 in the image data is generated. In a case where a plurality of objects are included in the image data in addition to the object 4 whose pose is to be estimated, a segment region is generated for all of the objects.

[0062] In addition, the image information preprocessing section 8 transforms the generated segment region of each object into a rectangle. For example, the segment region of each object is transformed into a rectangle by generating a quadrangular region (bounding box) of each object. Note that, at this time, the object to be recognized can be only an object 4 included in a label list set to the learning model, that is, an object 4 that is to be recognized by the learning model or an object 4 that is to be learned by the learning model.

[0063] Then, the image information preprocessing section 8 masks and crops a segment image to which a segment region is added for each object. That is, only the object 4 whose pose is to be estimated to which a segment region is added is left, and the other objects are masked. Then, a cropped image in which only the object 4 to which a segment region is added is cropped from the masked image in which the object 4 is masked is generated.

[0064] The super-resolution processing section 9 performs super-resolution on the cropped image generated by the image information preprocessing section 8 using a convolutional neural network. The super-resolution is a process performed to restore the resolution of image data to a high resolution, and is performed by a convolutional neural network such as SRCNN (Super-Resolution Convolutional Neural Network) that can transform such image data into high-resolution image data.

[0065] Further, the SRCNN is constructed by combining a convolution layer and a ReLU activation function (ReLU function). In the SRCNN, an object 4 included in the image data is enlarged by, for example, a bicubic interpolation of a pixel supplemented by a weighted average of values of 16 pixels around the pixel at the center of the pixel to be generated, and the like. Then, a high-resolution image is generated by a neural network constituted by the convolution layer and the ReLU function. Further, a loss function of the SRCNN uses a mean square error. Further, the super-resolution processing section 9 corresponds to a super-resolution image generation section in the embodiment of the present application.

[0066] The normal image generation section 10 generates a normal vector image (normal image) from the super-resolution image generated by the super-resolution processing section 9, that is, a depth image after super-resolution. The normal image is an image that represents the orientation of a face of the object 4 by attaching a normal vector perpendicular to the surface of the object 4. In the normal image generation section 10, a normal image in which a normal vector for representing the outline of the object 4 is attached to the super-resolution image is generated. Note that such a structure for finding a normal vector of the surface of the object 4 uses a generally disclosed and provided API. Further, the normal image generation section 10 corresponds to a normal data acquisition section in the embodiment of the present application.

[0067] The correction image generation section 11 generates a correction image based on the super-resolution image and the normal image respectively generated by the super-resolution processing section 9 and the normal image generation section 10. In the correction image generation section 11, first, a feature amount of the object 4 is extracted from the super-resolution image and the normal image, respectively. The feature amount is a value that quantitatively represents a qualitative feature of the object 4 included in each of the super-resolution image and the normal image. Further, the extraction of the feature amount is to take out an element for determining the object 4, such as data that acquires a luminance distribution, an edge, a color occurrence rate, and the like, using deep learning (deep learning). Further, the extraction of the feature amount is performed by generating a feature map or the like using a learning model. At this time, the learning model can be subjected to additional learning in advance on a depth image for learning that includes the object 4 whose attitude is intended to be estimated. Note that by extracting the feature amount from the super-resolution image, data related to the distance (size) of the object 4 is mainly extracted. Further, by extracting the feature amount from the normal image, data related to the surface and the outline of the object 4 is extracted.

[0068] In the correction image generation section 11, the super-resolution image and the normal image are respectively input to a convolutional neural network for image classification, whereby the extraction of the feature amount is performed. As the convolutional neural network for image classification, for example, DenseNet (Densely Connected Convolutional Network) or the like is used.

[0069] DenseNet is a configuration characterized by repeatedly fusing and configuring convolution layers. In DenseNet, an architecture called a Dense block is provided, which is a convolution block that inputs all the outputs of the layers located before each layer as feature maps. DenseNet has convolution layers and pooling layers between its Dense blocks. In addition, a Transition-Layer is provided as a layer for compressing (downsampling) the number of channels that has become large due to the Dense blocks. That is, DenseNet directly inputs the feature maps that have extracted the feature amounts of the object 4 generated in each layer to the next layer, whereby the loss of information between layers is less, and it is possible to alleviate the so-called vanishing gradient problem and achieve reinforcement of deep feature propagation, efficient use of features, and the like.

[0070] Then, the correction image generation section 11 generates a correction image in which the object 4 is corrected, on the basis of the feature amounts extracted or generated by inputting the super-resolution image and the normal line image to DenseNet, respectively. That is, in the correction image generation section 11, a correction image that more clearly exhibits the features of the object 4 is generated on the basis of the feature amounts extracted by the image classification convolutional neural network.

[0071] Specifically, the correction image generation section 11 fuses the feature amounts of the object 4 extracted from the super-resolution image generated by the super-resolution processing section 9 and the feature amounts of the object 4 extracted from the normal line image generated by the normal line image generation section 10 in the channel direction. As an example of such fusion, first, in deep learning, a feature map that spatially shows the feature amounts on the basis of the features of the object 4 extracted from the super-resolution image, and a feature map that spatially shows the feature amounts on the basis of the features of the object 4 extracted from the normal line image are generated. Each feature tensor having each feature map generated from each image is fused with each other in the channel direction (that is, the values after aligning the dimensions of the data with each other). The feature tensor is, for example, a multi-dimensional tensor including a set of feature maps or other features, and is composed of multiple dimensions such as a batch size, a number of channels, and the like.

[0072] Then, in the correction image generation section 11, a first synthesized feature amount obtained by fusing the feature tensors with each other in the channel direction is generated. Then, by inputting the generated first synthesized feature amount to a decoder, restoration of image data is performed. In the decoder, the size is increased mainly by reducing the increased number of channels through the extraction and synthesis of the feature amounts described above. That is, by performing transposed convolution (inverse convolution) or the like, the size of the feature map is reduced. That is, upsampling for restoring the downsampled feature map is performed. In this way, image data restored by reconstructing the features of the object 4 is generated as a correction image (restored image).

[0073] It should be noted that, in generating the corrected image, a so-called weight sharing mechanism is employed to optimize learning efficiency or reduce processing weights, sharing similar weights (filters) across all images. Furthermore, this process of correcting the surface (curved surface) of object 4 based on the super-resolution image and normal image is performed using a block composed of multiple layers such as convolutional layers and activation functions used for this process. The corrected image generation unit corresponds to the feature extraction unit and the corrected image generation unit in the embodiments of the present invention.

[0074] The object information estimation unit 12 estimates the features of the object 4 itself based on the corrected image generated by the corrected image generation unit 11. First, the object information estimation unit 12 extracts the feature quantities of the object 4 contained in the corrected image. In the object information estimation unit 12, similar to the image processing used by the image information preprocessing unit 8, the feature quantities of the object 4 are extracted by image recognition processing using a convolutional neural network such as DenseNet. That is, by inputting the corrected image into the CNN, convolution processing and pooling processing are repeatedly performed, thereby extracting the feature quantities of the object 4. It should be noted that, by extracting feature quantities from the corrected image, data related to the shape, size, and orientation of the object 4 are mainly extracted. In addition, the processing of extracting the feature quantities of the object 4 from the corrected image of the object 4 is performed by a block (intermediate layer block) composed of multiple layers such as convolutional layers and activation functions used for such processing.

[0075] Furthermore, the object information estimation unit 12 outputs information related to the features of the final object 4 itself. In the object information estimation unit 12, the size (scale), label (category or class), rotation amount (rotation angle), and thickness of the object 4 are estimated based on the extracted feature quantities. It should be noted that size refers to the dimensions of the object 4, label refers to the type of object 4 (workpiece) such as a wrench, and rotation amount (rotation angle) refers to the degree to which the object 4 has rotated relative to the camera 2. In the object information estimation unit 12, the output is provided with descriptions corresponding to each parameter. For example, size and rotation amount are described using their coordinates (x, y, z or x, y, z, w), and labels are described using matrices (1xN), etc. Additionally, different output headers are used for each parameter when outputting them.

[0076] For example, such as Figure 2 As shown, the parameters are estimated using a scale header (for estimating the size), a Rabel header (for estimating the label), and a rotational header (for estimating the rotation amount). It should be noted that the rotational header is constructed from a combination of a linear transformation layer and a ReLU function as the activation function. Additionally, as... Figure 2As shown, the amount of rotation (rotation angle) of the image is configured to be fitted by a least squares method such as the Gauss-Newton method described later, to the output result. Note that the object information estimation unit 12 corresponds to the state estimation unit in the embodiment of the present application.

[0077] In the above processing, it is estimated what kind of workpiece the object 4 is, the workpiece size thereof, and the like. Next, a process for estimating the position of the object 4 using the image generated in the above processing will be described.

[0078] The object position estimation unit 13 extracts a feature quantity from the occlusion image generated by the image information preprocessing unit 8 to estimate the position of the object 4. In the object position estimation unit 13, first, the occlusion image generated by the image information preprocessing unit 8 is acquired. The occlusion image is an image in which the image data (depth image) is occluded except for the object 4 whose posture is to be estimated as described above. That is, in the occlusion image, random noise is added to the object 4, and the segment region of this object 4 is transformed into a rectangle, and becomes a depth image that does not include other objects 4.

[0079] The object position estimation unit 13 extracts a feature quantity of the object 4 from the occlusion image. As described above, the feature quantity is a value that quantitatively represents the qualitative feature of the object 4. The feature quantity is extracted, for example, by a CNN in which a convolution layer and a pooling layer are alternately arranged. Specifically, the feature is extracted by applying a convolution process using a filter that reacts to a specific shape of the image data, and a pooling process that aggregates the image data to be smaller by extracting one value from the values of a window of a prescribed range in the values after convolution. Then, based on the plurality of features extracted in this way, the feature quantity is extracted by generation of a feature map that spatially represents the feature quantity, or a feature tensor that represents a set of feature maps, and the like.

[0080] In addition, the object position estimation unit 13 generates a second synthesized feature quantity that fuses the feature quantity thus extracted from the occlusion image and the feature quantity extracted from the corrected image generated by the corrected image generation unit 11 in the channel direction. That is, a feature map or a feature tensor based on the features of the object 4 extracted from the occlusion image, and a feature map or a feature tensor based on the features of the object 4 extracted from the corrected image are generated. By fusing these two feature maps with each other or the feature tensors with each other in the channel direction (i.e., the values after aligning the dimensions of the data with each other), the second synthesized feature quantity is generated.

[0081] Then, the object position estimation unit 13 estimates the position of object 4 based on the generated second synthetic feature quantity. The object position estimation unit 13 uses existing image recognition algorithms to estimate the position of object 4. That is, it uses the aforementioned DenseNet, ResNet (Residual Network), etc., which are CNNs used for image recognition, to estimate the position of object 4. It should be noted that ResNet has an architecture also known as residual blocks. By performing skip connections that append the signal (image data) input to a specified layer to the output of a layer higher than the specified layer, the propagated error can be transmitted without attenuation, thus enabling high-precision image recognition.

[0082] Then, the object position estimation unit 13 estimates and outputs the position of object 4 based on the generated second synthetic feature quantity. The position of object 4, as its position relative to camera 2, is represented by three-dimensional coordinates (x, y, z). Furthermore, when outputting these coordinates, as shown... Figure 2 As shown, an output head, such as a Transformer-Header (TransHeader), is used to estimate the position of object 4. By inputting a second synthetic feature value into this Transformer-Header, the position of object 4 relative to camera 2 is estimated and output. That is, in the object position estimation unit 13, the position is estimated based on a first synthetic feature value generated from a corrected image, etc., and a second synthetic feature value generated by the object position estimation unit 13. It should be noted that the processing of extracting the feature value of object 4 from the top-view image of object 4 is performed by a block composed of multiple layers such as convolutional layers and activation functions used for such processing. It should be noted that the object position estimation unit 13 corresponds to the position estimation unit in the embodiment of the present invention.

[0083] The object attitude estimation unit 14 estimates the attitude of object 4 based on feature data regarding the size, label, and rotation amount of object 4 estimated by the object information estimation unit 12, and position data regarding the three-dimensional position of object 4 relative to camera 2 estimated by the object position estimation unit 13. Based on these parameters, the object attitude estimation unit 14 estimates what kind of workpiece object 4 is, its placement method, orientation, etc. The estimated result is output as a probability value, for example. That is, the estimated result assigns a high probability to the label and attitude with the highest probability, and conversely, assigns a low probability to those with low probability.

[0084] Further, in the object posture estimation section 14, fitting is performed based on the obtained feature data and position data. That is, the object posture estimation section 14 adjusts the learning model of the object 4, the parameters of the function for estimation of the object 4, based on the input data (Grand Truth) related to the actual object 4 and the output result estimated to be related to the object 4. At this time, the matching of the input data and the learning model is optimized by using the least square method, for example, the Gauss-Newton method as a nonlinear least square method, thereby performing the process of constructing a model with high construction accuracy. Further, the degree of shift (error) of them is found (evaluated) by the Loss function which is configured to evaluate and output by comparing the output result (prediction result) with the actual parameter (teacher data). Note that the object posture estimation section 14 corresponds to the position estimation section and the optimization section in the embodiment of the present application.

[0085] The output section 15 displays the result estimated by the object posture estimation section 14. In the output section 15, for example, the item with the highest possibility with respect to the position, label, orientation, etc. of the estimated object 4 is displayed. At this time, the possibility of the item is displayed quantitatively as a probability. Note that the output section 15 can be, for example, a monitor of a PC or a display of a portable terminal, etc.

[0086] Next, the process or step (procedure) performed by the posture estimation apparatus 1 and the estimation estimation method in the embodiment of the present application will be described. In Figure 7 , Figure 8 and Figure 9 , as one example of the process or procedure, a flowchart (procedure chart) of the process performed when estimating the posture of the object 4 disposed at a position farther from the camera 2 in the work site 3 is shown.

[0087] As shown in Figure 7 , first, in step S1, initial setting in the learning model is performed. The initial setting performed in step S1 includes setting of data related to the object 4, the target object, setting of data related to the degree of learning in the learning model, etc.

[0088] After the initial setting of the learning model is completed, the process proceeds to step S2, and reading of image data is performed. In step S2, image data captured by the camera 2 is input. The input of the image data can be performed manually by a user or the like, or can be performed by automatic transfer from the camera 2. Note that in step S2, the input image data is subjected to a preprocessing such as normalization processing.

[0089] After inputting the image data, the processing proceeds to step S3, where noise is added to the image data. As described above, camera 2 can infer the depth and proximal dimension of object 4 based on the image data, and this information is added to the image data. That is, the image data input in step S2 is also a depth image (depth_img) with data related to the depth or proximal dimension of object 4 appended. In step S3, noise is randomly added to the image data that serves as the depth image or to the object 4 contained within the depth image. By adding noise to the depth image, the performance and robustness of the learning model are improved.

[0090] After adding noise to the depth image, the process proceeds to step S4, where the region containing object 4 in the image data is estimated. In step S4, firstly, based on the label contained in the data of object 4 set in the initial configuration, a segment region of object 4 corresponding to that label is generated or appended. That is, in step S4, as... Figure 4 As shown in (a), all additional segment regions of the workpiece, etc., in the learning model are set as objects for which the pose is to be estimated. In addition, in step S4, each object is detected, for example, based on the feature quantities of each object containing object 4 that have been learned in advance, or the pixel regions of object 4 in the image data are detected, thereby generating the segment regions of object 4.

[0091] Then, in step S4, the segment regions of each generated object 4 are transformed into rectangles. For example, in step S4, as... Figure 4 As shown in (b), quadrilateral regions (bounding boxes) of each object 4 are generated by object 4 detection, thereby transforming the segment regions of each object into rectangles.

[0092] After generating the segment image (seg_img) in step S4, the process proceeds to step S5, where the depth image generated in step S3 is subjected to occlusion processing. In step S5, as... Figure 4 As shown in (c), a masking image (headview_img) is generated. In this masking image (headview_img), among multiple objects that have undergone segment processing, the region of the specified object 4 whose posture is to be estimated is retained, and the other parts are masked. For example, if multiple objects (workpieces) are detected in step S4, these multiple objects 4 become different layers, so the layer of the specified object 4 whose posture is to be estimated is extracted.

[0093] After masking the image data, the process proceeds to step S6, where the depth image is cropped according to the defined segment region of object 4. That is, in step S6, as... Figure 5 As shown in (a), a cropped image (crop_depth) is generated in a manner that cropping is performed to retain only the specified object 4 contained in the depth image.

[0094] After the cropping processing is thus performed, the processing proceeds to step S7 to super-resolve the cropped image generated in step S6. In step S7, the above-described SRCNN or the like, which is a neural network for super-resolving the object 4, is used. Thereby, features are extracted from the cropped image input to the SRCNN to supplement missing (deficient) data, and a super-resolved cropped image, i.e., a super-resolution image, as shown in (b) of FIG. 8 is generated. As described above, since the object 4 is located at a position relatively far from the camera 2, the stereoscopic sense of the object 4 is not sometimes expressed in the captured image data, or the hole, gap, or the like of the object 4 is flattened. In step S7, by the cropped image being super-resolved, the stereoscopic sense, hole, or the like of the object 4 is restored to a point cloud with relatively high accuracy. Figure 5

[0095] After the super-resolved cropped image is generated, the processing proceeds to step S8 to generate a normal image (crop_norm) of the object 4 from the super-resolution image. As shown in (c) of FIG. 8, the normal image is an image that represents the orientation of the surface of the object 4 by additionally attaching data related to a vector perpendicular to the surface of the object 4. In step S7, by the cropped image being super-resolved, the surface of the object 4 is restored with relatively high accuracy. Therefore, the normal image generated in step S8 also can express the surface of the object 4 with relatively high accuracy. Figure 5

[0096] After the normal image is generated, the processing proceeds to step S9 to extract a feature amount (feature amount 1) of the object 4 from the super-resolution image. In step S9, a feature portion of the object 4 included in the super-resolution image is extracted, and a qualitative parameter is extracted as a quantitative parameter, i.e., a feature amount. Such processing is performed using DenseNet or the like, which is a convolutional neural network for image classification. In addition, the feature amount extracted at this time is mainly a feature amount for detecting or estimating the distance of the camera 2 from the object 4. Note that an example of data input to the convolutional neural network for image classification is shown in (a) of FIG. 9. Figure 6

[0097] After the feature amount of the object 4 is extracted from the super-resolution image, the processing proceeds to step S10, and a feature amount (feature amount 2) of the object 4 is extracted from the normal image generated in step S8. In step S10, as with the processing performed in step S9, a feature portion of the object 4 included in the normal image is extracted using a CNN for image classification, and a qualitative parameter is extracted as a quantitative parameter, i.e., a feature amount. In addition, the feature amount extracted at this time is mainly a feature amount for detecting or estimating the surface, outline, or the like of the object 4. Note that an example of data input to the convolutional neural network for image classification is shown in (b) of FIG. 10. Figure 6

[0098] ​​​​After the feature quantity of the object 4 is extracted from the normal image, the process proceeds to step Sll, and a first synthesized feature quantity (feature quantity 3) is generated by fusing the feature quantities extracted in step S9 and step S10. In step Sll, the feature quantity extracted from the super-resolution image and the feature quantity extracted from the normal image are fused. At the time of fusion, the generated respective feature quantity tensors are fused with each other in the channel direction (i.e., the values after the dimensions of data are aligned with each other). In this way, the first synthesized feature quantity is generated by fusing the feature quantities of the super-resolution image and the normal image in the channel direction.

[0099] After the first synthesized feature quantity is generated, the process proceeds to step S12, and a correction image is generated from the first synthesized feature quantity. In step S12, based on the first synthesized feature quantity, a correction image in which the prescribed object 4 is corrected (depth_imprior) is generated. In step S12, by inputting the first synthesized feature quantity to a decoder or the like, the correction image shown in (c) of FIG. 13 is generated. That is, when the object 4 is expanded by deconvolution of the first synthesized feature quantity, the missing part is interpolated using the copied value of the periphery. That is, by performing up-sampling processing (Upsampling) that enlarges (expands) the feature map, the object 4 is restored, and the correction image is generated. Figure 6

[0100] After the correction image is generated, the process proceeds to step S13, and a feature quantity (feature quantity 4) of the object 4 is extracted from the generated correction image. In step S13, as with the process performed in step S9 and step S10, the feature quantity is extracted by a CNN used in image processing such as DenseNet, from the feature part of the object 4 contained in the correction image, whereby the feature quantity is extracted. The feature quantity extracted in step S13 is mainly a feature quantity for detecting or estimating the shape and size of the object 4, the orientation of the object 4, and the like.

[0101] ​After the feature amount of the object 4 is extracted from the corrected image, the process proceeds to step S14, and detailed information related to the object 4 is estimated from the extracted feature amount of the object 4 in the corrected image. In step S14, data such as the size (dimension), label (classification), and rotation amount of the object 4 are estimated from the feature amount of the object 4. That is, based on the feature amount of the object 4 extracted from the corrected image, values estimated by estimating these data are output as final output. The output values are represented by coordinates or matrices. For example, the size of the object 4 is represented by coordinates in the X-axis direction, Y-axis direction, and Z-axis direction, and the rotation amount of the object 4 is further represented by adding coordinates of the W-axis orthogonal to the Z-axis. In addition, the label is represented by a matrix corresponding to the number of labels of the estimated object 4 and the like. Note that the size, label, and rotation amount of the object 4 are respectively estimated by inputting to the head network. That is, the detailed data of the object 4 are estimated by the head that estimates the size, the head that estimates the label, and the head that estimates the rotation amount, from the input feature amount.

[0102] After the data of the object 4 is estimated, the process proceeds to estimate the position of the object 4. After the process in step S14, the process proceeds to step S15, and a feature amount (feature amount 5) of the object 4 is extracted from the masked image to which the masking process is performed on the image data generated in step S5. In step S15, similarly to the process performed in step S9 and the like, by inputting the masked image to DenseNet, a feature portion of the prescribed object 4 included in the masked image is extracted, and a qualitative parameter is extracted as a quantitative parameter, that is, a feature amount. With respect to the feature amount extracted at this time, a feature related to the position of the object 4 in the field of view angle of the image data, that is, the relative position with respect to the camera 2 is mainly extracted.

[0103] After the feature amount related to the position of the object 4 is extracted, the process proceeds to step S16, and the extracted feature amount of the object 4 is synthesized or fused. In step S16, a second synthesized feature amount (feature amount 6) after the feature amount related to the position of the object 4 extracted in step S15 and the feature amount related to the object 4 itself extracted in step S13 are synthesized is calculated. In step S16, similarly to the process in step S11, values after the generated respective feature amount tensors are aligned with each other in the channel direction, that is, the dimension of the data, are fused with each other.

[0104] After the second synthesized feature amount is generated, the process proceeds to step S17, and the position of the object 4 in the image data is estimated. In step S17, the position of the object 4 is estimated based on the second synthesized feature amount of the object 4 generated in step S16. In step S17, for example, the position of the object 4 is estimated and output from each patch obtained by subdividing the image data based on the architecture of the Transformer. The position of the object 4 is output by coordinates in the X-axis direction, Y-axis direction, and Z-axis direction.

[0105] Thus, by the processes of Step S1 to Step S17, the size (dimension), label (category), amount of rotation (rotation angle), and position of the object 4 included in the image data are detected or estimated, and thus the pose of the object 4 is estimated.

[0106] After the pose of the object 4 is estimated, the process proceeds to Step S18, and fitting is performed based on the estimated position and pose of the object 4. In the fitting, the parameters of the learning model of the object 4 and the parameters of the function for the estimation of the object 4 are adjusted from the parameters of the object 4 based on the input image data and the parameters of the object 4 based on the estimated output result. Specifically, using a least squares method such as the Gauss-Newton method, the parameters of the learning model are adjusted in order to minimize the error function of the input actual data (Grand Truth) and the estimated result. Thus, in Step S18, by optimizing the matching of the input data and the learning model, the process of constructing a model with high precision is performed.

[0107] Thus, in the pose estimation device 1 and the pose estimation method of the object 4 of the embodiment of the present application, in estimating the pose of the object 4 included in the image data, first, a depth image and a normal vector image (normal image) of the object 4 are generated based on the input image data. From the generated depth image, a super-resolution image that is super-resolution-processed is generated by a convolutional neural network (SRCNN) that improves the resolution of the object 4 included in the image data. Using a CNN suitable for image classification such as DenseNet, feature amounts of the object 4 are extracted from the super-resolution image and the normal image, respectively. A first synthetic feature amount in which the extracted respective feature amounts are aligned and fused in the channel direction is generated, and a correction image is generated based on the first synthetic feature amount. Then, a result of estimating the features of the object 4 itself such as the size (dimension), label (category), and amount of rotation (rotation angle) from the feature amounts of the object 4 extracted from the correction image is output.

[0108] After that, a feature amount of the object 4 is extracted from a depth-occluded image in which only the object 4 is left from the input image data. A second synthetic feature amount in which the feature amount extracted from the depth-occluded image and the feature amount of the object 4 extracted from the correction image are fused in the channel direction is generated. The position of the object 4 from the camera 2 is estimated from the second synthetic feature amount. That is, the pose of the object 4 is estimated by the data output based on the first synthetic feature amount and the second synthetic feature amount.

[0109] For example, in a case where the object 4 has to be arranged at a position far from the camera 2 due to the situation of the work site 3 or the like, the object 4 included in the captured image data becomes small. In this case, the resolution of the object 4 is low, the feature amount of the object 4 is small, or the shape of the object 4 is recognized as a flat surface or the like because the distance between the positions at which the feature amounts are extracted from each other becomes long, and it can be impossible to obtain sufficient feature amounts for estimating the posture of the object 4.

[0110] Even in such a case, in the posture estimation device 1 or the posture estimation method in the embodiment of the present application, the shape of the object 4 is restored with high accuracy by the super-resolution image obtained by super-resolving the input image data (depth image), the normal vector image of the object 4 generated from the super-resolution image, and each image after normalization. Therefore, the surface, the contour, the thickness, and the distance of the object 4 can be restored with high accuracy, and thus the features of the object 4 itself can be estimated with high accuracy.

[0111] In addition, the position of the object 4 is estimated based on the occlusion image in which the object 4 is occluded from the image data (depth image) captured in the overhead manner and the above-described correction image. The above-described super-resolution image is a segment image in which only the object 4 is cropped, and thus although the features of the object 4 itself are extracted with high accuracy, the positional relationship between the camera 2 and the object 4 cannot be extracted with high accuracy. Therefore, by further using the feature amount of the object 4 extracted from the object 4 in the occlusion image in which the image of the object 4 is occluded from the input image data captured in the overhead manner, the learning model can learn the actual positional relationship with high accuracy, and the estimation accuracy of the posture of the object 4 can be further improved.

[0112] The above describes the embodiment of the present application, but the present application is not limited to the above-described example, and can be appropriately changed within a range in which the object of the present application is achieved. For example, the input image data is not limited to the above-described depth image, and can be point cloud data in which the shape of the object 4 is expressed by a point cloud by 3D scanning or the like, or image data in which three channel data of RGB expressing the color of the object 4 is attached in addition to the depth image. In addition, the depth image is not limited to a structure in which the depth data is acquired by the camera 2, and can be configured to acquire data (depth data) related to the depth of the object by a depth data acquisition unit that extracts the depth data by analyzing the acquired image data. Furthermore, the posture estimation system including the posture estimation device 1 and the camera 2 illustrated in FIG. 1 can be configured. Figure 1 The posture estimation system including the posture estimation device 1 and the camera 2 illustrated in FIG. 1 can be configured.

Claims

1. An attitude estimation device of an object, which uses a convolutional neural network to analyze image data obtained by photographing an object, and estimates an attitude of the object included in the image data, characterized by comprising: a first neural network configured to estimate an attitude of the object included in the image data, wherein the first neural network is configured to estimate the attitude of the object included in the image data by using a first neural network model, and wherein the first neural network model is configured to estimate the attitude of the object included in the image data by using a first neural network model. Possessing: a depth data acquisition section that acquires a depth image that contains depth data that is data related to the depth of the object contained in the image data; a normal data acquisition section that acquires a normal image that is an image obtained by finding the normal vector of the surface of the object from the depth image; a super-resolution image generation section that generates a super-resolution image that is an image obtained by increasing the resolution of the object cropped out from the depth image; a feature quantity extraction section that extracts a feature quantity that quantitatively indicates the features of the object contained in the image data, the super-resolution image, and the normal image; a corrected image generation section that generates a corrected image that is an image obtained by restoring the shape of the object based on the feature quantity extracted from the super-resolution image and the normal image; a state estimation section that estimates a parameter related to the state of the object based on the corrected image; a position estimation section that estimates the position of the object from the image data; and a posture estimation section that estimates the posture of the object based on the parameter related to the state of the object estimated by the state estimation section and the position of the object estimated by the position estimation section.

2. The posture estimation device of the object according to claim 1, characterized in that the posture estimation device of the object further possesses an optimization section that fits the output data of the position of the object estimated by the position estimation section and the posture of the object estimated by the posture estimation section and the actual data related to the position and posture of the object by a Gauss-Newton method.

3. The posture estimation device of the object according to claim 1 or 2, characterized in that the corrected image generation section is configured to generate a synthetic feature quantity that is obtained by fusing the feature quantity of the object extracted from the depth image and the feature quantity of the object extracted from the normal image in the channel direction, restore the object by implementing up-sampling processing that interpolates a missing portion using a copy value of the surroundings based on the synthetic feature quantity, and generate the corrected image.

4. A posture estimation method of an object that analyzes image data obtained by photographing an object using a convolutional neural network and estimates the posture of the object contained in the image data, characterized by acquiring a depth image that contains depth data that is data related to the depth of the object contained in the image data, acquiring a normal image that is an image obtained by finding the normal vector of the surface of the object from the depth image, generating a super-resolution image that is an image obtained by increasing the resolution of the object cropped out from the depth image, extracting a feature quantity that quantitatively indicates the features of the object contained in the image data, the super-resolution image, and the normal image, generating a corrected image that is an image obtained by restoring the shape of the object based on the feature quantity extracted from the super-resolution image and the normal image, ​ estimating a parameter related to a state of the object based on the corrected image, estimating a position of the object from the image data, estimating a posture of the object based on the estimated parameter related to the state of the object and the estimated position of the object.

5. The posture estimation method of an object according to claim 4, wherein fitting output data related to the estimated position of the object and the estimated posture of the object and actual data related to the position and the posture of the object by a Gauss-Newton method.

6. The posture estimation method of an object according to claim 4 or 5, wherein generating a synthetic feature amount in which the feature amount of the object extracted from the depth image and the feature amount of the object extracted from the normal image are fused in a channel direction, based on the synthetic feature amount, restoring the object by performing up-sampling processing using a copy value of a periphery to interpolate a missing portion, and generating the corrected image.