Storage Medium for Apparatus, Method, and Program for Speculation, Learning, and Production of Teaching Data

By extending the blank area and using the learning model, the difficulty of inferring skeleton information in the prior art in images that do not contain certain joints is solved, and stable and accurate skeleton speculation under privacy protection is achieved.

CN115244578BActive Publication Date: 2025-06-27OMRON CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180018560.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-20
Filing Date
2021-03-05
Publication Date
2025-06-27
Estimated Expiration
2041-03-05

AI Technical Summary

Technical Problem

The prior art is difficult to stably infer skeleton information in images that do not contain certain joints, especially in scenarios of privacy protection and job analysis.

Method used

By obtaining an image containing a portion of the joints of the objector, expanding the blank area to generate a new image, and using the learned speculative model to infer the skeleton information, including the joint position located in the blank area.

Benefits of technology

It realizes the skeleton information that contains the joint position of the blank area without exposing the privacy of the object, which improves the accuracy and automation of the skeleton speculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115244578B_ABST
    Figure CN115244578B_ABST
Patent Text Reader

Abstract

The present invention provides a speculation device, a learning device, a teaching data production device, a speculation method, a learning method, a teaching data production method, and a storage medium for a program. In an image including a blank area, skeleton information of a part having a joint in the blank area is speculated. The speculation device (1) includes: an input unit (130) that acquires a first image including a first joint of an object person and not including a second joint; a blank area expansion unit (101) that generates a second image obtained by expanding the first image with the blank area; and a speculation unit (12) that speculates the skeleton information using the second image and a learned speculation model, the skeleton information including a joint position of the second joint located in the blank area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a speculation device that speculates the skeleton position of a subject using an image captured of the subject, and the like, and particularly relates to a speculation, learning, teaching data production device, method, and storage medium for a program. Background Art

[0002] Factory worker (subject) operation improvement is carried out by analyzing human movements. Conventionally, a person watches a moving image captured by a camera to measure the operation time to promote operation improvement. Attempts to automate human movement analysis are also being advanced, and as a conventional technique, speculation of joint positions and corresponding relationships using deep learning, that is, skeleton speculation (Non-Patent Document 1) is known.

[0003] Prior Art Documents

[0004] Non-Patent Documents

[0005] Non-Patent Document 1: “OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019 (“OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”, IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2019) Summary of the Invention

[0006] Problems to be Solved by the Invention

[0007] However, the conventional technique as described above requires shooting from an angle (for example, from above) that does not capture the face of the worker (subject) in view of privacy protection and work analysis on the workbench.

[0008] As an existing deep learning-based skeleton speculation, OpenPose described in Non-Patent Document 1 can be cited. In OpenPose, due to the structure of the processing range, it is not possible to stably speculate the skeleton based on an image that does not include joints such as shoulders or necks.

[0009] An object of an embodiment of the present invention is to stably speculate skeleton information of joint positions including a joint of a subject based on an image that does not include a part of the joints of the subject.

[0010] Technical Means for Solving the Problems

[0011] To solve the above problems, a speculation device according to an embodiment of the present invention includes: an image acquisition unit that acquires a first image including a first joint of a subject and not including a second joint; a blank area expansion unit that generates a second image obtained by expanding the first image with a blank area; and a speculation unit that uses the second image and a learned speculation model to speculate skeleton information, the skeleton information including joint positions of the second joint located in the blank area.

[0012] To solve the above problems, a learning device according to an embodiment of the present invention includes: an image acquisition unit that acquires a first image including a first joint of a subject and not including a second joint; a blank area expansion unit that generates a second image obtained by expanding the first image with a blank area; a teaching data storage unit that stores teaching data including skeleton information and the second image, the skeleton information including the second joint located in the blank area; and a learning unit that uses the teaching data to learn a speculation model of skeleton information based on the skeleton information and the second image.

[0013] To solve the above problems, a teaching data production device according to an embodiment of the present invention includes: an image acquisition unit that acquires a first image including a first joint of a subject and not including a second joint; a blank area expansion unit that generates a second image obtained by expanding the first image with a blank area; a display control unit that displays the second image; an input unit that receives an input of joint positions of the second joint from a user for the blank area in the second image; and a teaching data production unit that produces teaching data associating skeleton information including joint positions of the first joint and the second joint with the second image.

[0014] A speculation method according to an embodiment of the present invention includes: an image acquisition step of acquiring a first image including a first joint of a subject and not including a second joint; a blank area expansion step of generating a second image obtained by expanding the first image with a blank area; and a speculation step of using the second image and a learned speculation model to speculate skeleton information, the skeleton information including joint positions of the second joint located in the blank area.

[0015] A learning method according to an embodiment of the present invention includes: an image acquisition step of acquiring a first image including a first joint of a subject and not including a second joint; a blank area expansion step of generating a second image obtained by expanding the first image with a blank area; a teaching data acquisition step of acquiring teaching data including skeleton information and the second image, the skeleton information including the second joint located in the blank area; and a learning step of using the teaching data to learn a speculation model of skeleton information based on the skeleton information and the second image.

[0016] A method for creating teaching data according to an embodiment of the present invention includes: an image acquisition step of acquiring a first image that includes a first joint of a subject and does not include a second joint; a blank area expansion step of generating a second image obtained by expanding the first image with a blank area; a display control step of displaying the second image; an input step of receiving an input of the joint position of the second joint from a user for the blank area in the second image; and a teaching data creation step of creating teaching data that associates skeleton information including the joint positions of the first joint and the second joint with the second image.

[0017] Effects of the Invention

[0018] According to an embodiment of the present invention, it is possible to infer skeleton information including joints located in a blank area. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 FIG. is a block diagram showing an example of components in the learning operation of the inference device according to Embodiment 1 of the present invention.

[0020] Figure 2 FIG. is a block diagram showing an example of components in the inference operation of the inference device according to Embodiment 1 of the present invention.

[0021] Figure 3 FIG. is a model diagram showing the data state of components in the learning operation of the inference device according to Embodiment 1 of the present invention.

[0022] Figure 4 FIG. is a schematic diagram of skeleton information inferred using the inference device according to Embodiment 1 of the present invention.

[0023] Figure 5 FIG. is a flowchart showing the learning process of the inference device according to Embodiment 1 of the present invention.

[0024] Figure 6 FIG. is a flowchart showing the inference process of the inference device according to Embodiment 1 of the present invention.

[0025] Figure 7 FIG. is a block diagram showing an example of the main part structure of the teaching data creation device according to Embodiment 2 of the present invention.

[0026] Figure 8 FIG. is an example of a user interface for designating a joint position in the teaching data creation unit according to Embodiment 2 of the present invention.

[0027] Figure 9 FIG. is a flowchart showing the operation of the teaching data creation device according to Embodiment 2 of the present invention.

[0028] Figure 10It is a schematic diagram of an image of an operator taken from above in Embodiment 3 of the present invention.

[0029] Figure 11 It is a schematic diagram of an image of an operator taken from the side in Embodiment 3 of the present invention.

[0030] [Description of symbols]

[0031] 1: Estimation device

[0032] 2: Teaching data creation device

[0033] 10: Control unit

[0034] 11: Learning unit

[0035] 12: Estimation unit

[0036] 20: Storage unit

[0037] 51: User interface

[0038] 101: Blank area expansion unit

[0039] 102: Teaching data creation unit

[0040] 103: Data expansion unit

[0041] 104: Excess and deficiency area correction unit (deficiency area correction unit)

[0042] 111: Estimation model acquisition unit

[0043] 112: Feature quantity extraction unit

[0044] 113: Joint estimation unit

[0045] 114: Bonding degree estimation unit

[0046] 121: Skeleton estimation unit

[0047] 122: Estimation model learning unit

[0048] 130: Input unit (image acquisition unit)

[0049] 140: Output unit

[0050] 150: Display control unit

[0051] 201: Teaching data storage unit

[0052] 202: Estimation model storage unit

[0053] 300, 300a, 300b: Images (first images)

[0054] 301: Blank completed image (second image)

[0055] 302: Skeleton-containing data

[0056] 303: Image processing completed data

[0057] 303a: Image after image processing (third image)

[0058] 304: Teaching data

[0059] 304a: Image with over and under correction completed (new second image)

[0060] 312: Teaching skeleton information

[0061] 311, 311a, 311b, 311c, 311d, 311e, 311f, 311g, 311h, 314: Blank area 313: Deficient pixel area

[0062] 501: Image display unit

[0063] 511: Operator list

[0064] 512: Editing operator

[0065] 513: Operator addition button

[0066] 521: Joint list

[0067] 522: Editing joint

[0068] 531: Coordinate display indicator

[0069] 541: Editing joint position

[0070] 532: X - coordinate display indicator

[0071] 542: Determined joint position

[0072] 533: Y - coordinate display indicator

[0073] 543: Binding information

[0074] 601: Operator

[0075] 602: Workbench

[0076] 603: Object of operation

[0077] 604: Shielded area Detailed implementation mode

[0078] Hereinafter, an implementation mode of one aspect of the present invention (hereinafter also referred to as "this implementation mode") will be described based on the accompanying drawings.

[0079] 〔Embodiment 1〕

[0080] §1. Application Examples

[0081] The estimation device is a device that estimates the skeletal information of an operator (subject) using an image taken of the operator. The skeletal information includes information on the positions of the joints of the operator. The joint positions of the operator represent the posture of the operator corresponding to the operation action.

[0082] Before the estimation, the estimation device first learns an estimation model for the estimation. Specifically, the estimation device associates the skeletal information with an image taken from above of the operator and generates teaching data, where the skeletal information includes the joint positions contained in or inferable by humans in the image. The estimation device uses the generated teaching data to learn the estimation model.

[0083] Since the image taken from above of the operator does not include the face of the operator, it does not include some joints of the operator, such as the joints of the neck. During the learning, an image with a blank area set for the image lacking some joints is used for learning. Thus, an estimation model for estimating the missing joint positions in the blank area is generated.

[0084] As described above, after generating the estimation model, an image lacking some joints is input to the estimation device, and the estimation model is used for estimation. Thus, the skeletal information can be estimated in a state where privacy is protected.

[0085] §2. Structural Examples

[0086] Based on Figures 1 to 4 the structural example of the estimation device 1 will be described. Figure 1 It is a block diagram showing an example of the components of the estimation device 1 that play a role in the learning operation. Figure 2 It is a block diagram showing an example of the components of the estimation device 1 that play a role in the estimation operation. Figure 3 It is a model diagram showing the data state of the components in the learning operation of the estimation device 1. Figure 4 It is a schematic diagram of the skeletal information estimated using the estimation device 1.

[0087] As Figure 1 , Figure 2 shown, the estimation device 1 (learning device) includes: a control unit 10 that comprehensively controls each part of the estimation device 1; and a storage unit 20 that stores various data used by the estimation device.

[0088] In the control unit 10, there are included a blank area expansion unit 101, a teaching data creation unit 102, a data expansion unit 103, an over- and under-area correction unit 104, a speculation model acquisition unit 111, a feature quantity extraction unit 112, a joint speculation unit 113, a combination degree speculation unit 114, a skeleton speculation unit 121, a speculation model learning unit 122, an input unit 130, and an output unit 140. Moreover, in the storage unit 20, there are included a teaching data storage unit 201 and a speculation model storage unit 202.

[0089] In the control unit 10, the learning unit 11 functions during the learning operation (refer to Figure 1 ), and the speculation unit 12 functions during the speculation operation (refer to Figure 2 ). In the learning unit 11, there are included a feature quantity extraction unit 112, a joint speculation unit 113, a combination degree speculation unit 114, a skeleton speculation unit 121, and a speculation model learning unit 122. In the speculation unit 12, there are included a speculation model acquisition unit 111, a feature quantity extraction unit 112, a joint speculation unit 113, a combination degree speculation unit 114, and a skeleton speculation unit 121.

[0090] The input unit 130 (image acquisition unit) accepts data input including an image and user input for the speculation device 1. The input unit 130 accepts the input of the image 300 and outputs the image 300 to the blank area expansion unit 101. Moreover, the input unit 130 can also acquire the image 300 from a camera connected to the speculation device 1, an external server via a network, or a storage device within the speculation device 1. The input unit 130 can also support not only static images but also the input of dynamic images.

[0091] The image 300 (first image) is an image that reflects a part of the joints of the operator, such as the joints of the elbow or hand, but lacks another part of the joints of the operator, such as the neck or shoulders. The image 300 does not include, for example, the face of the operator, so the privacy of the operator can be protected. As the image 300, it can also be an image taken from above the operator. In this case, it is easy to capture the situation where the operator is performing an operation. By taking a picture from above, for example, it is also possible to capture the position change of the work object on the workbench as the situation where the operation is in progress. The image 300 taken from above also does not include another part of the joints of the operator, such as the neck or shoulders.

[0092] Moreover, the input unit 130 also accepts operation inputs from the user to the estimation device 1 via an input device. The input device may be, for example, an indicating device such as a mouse or a touch panel, or a cross key. The input unit 130 accepts from the user the specification of the joint positions for the image. As user input, it may also be the position specification of the completed image 301 using the indicating device, the position specification using the cross key, or the direct specification of the pixel coordinates of the completed image 301. The input unit 130 outputs the information of the input joint positions (teaching skeleton information 312) to the teaching data creation unit 102.

[0093] The blank area expansion unit 101 creates a completed image 301 with an expanded image range (image size) by attaching a blank area 311 adjacent to at least one side of the input image 300. Subsequently, the blank area expansion unit 101 outputs the completed image 301 to the teaching data creation unit 102 or the feature amount extraction unit 112.

[0094] The completed image 301 (second image) is an image formed by integrating the image 300 and the blank area 311. Figure 3 In, the blank area 311 is represented by diagonal hatching in the lower right, and is actually filled with a specific single color. The specific single color is, for example, black or white, but is not limited thereto. In addition, the blank area 311 may not be a single color, but a region with a specific pattern or hatching (the same pattern or hatching for multiple completed images 301). And, Figure 3 In, the blank area 311 is adjacent to one side of the image 300, but may also be adjacent to two or more sides.

[0095] The size and configuration of the blank area may also be parameters that can be set by user input.

[0096] The teaching data creation unit 102 creates skeleton-containing data 302 (teaching data) that associates the teaching skeleton information 312 with the completed image 301. Subsequently, the teaching data creation unit 102 outputs the skeleton-containing data 302 to the data expansion unit 103.

[0097] The skeleton-containing data 302 is data that includes the completed image 301 and the teaching skeleton information 312. The teaching skeleton information 312 includes information on the positions (joint positions) of multiple parts of the operator (neck, right shoulder, right elbow, right hand, left shoulder, left elbow, left hand, right waist, and left waist) corresponding to the completed image 301.

[0098] The joints (parts) set as the teaching skeleton information 312 are set as nine parts here, namely the neck, right shoulder, right elbow, right hand, left shoulder, left elbow, left hand, right waist, and left waist, but are not limited thereto. The head or feet may also be set, or some joints (such as the right hand) may be lacking. Figure 3In this case, for the teaching skeleton information 312, joints are represented by black circles and the connections between joints are represented by line segments.

[0099] The data expansion unit 103 performs arbitrary image processing on the input skeleton-containing data 302 to produce processed image data 303. Subsequently, the processed image data 303 is output to the over- and under-area correction unit 104.

[0100] The processed image data 303 is data that includes a processed image 303a obtained by applying image processing to the blank-completed image 301 and teaching skeleton information 312 that has undergone the same geometric transformation as the image processing. Examples of geometric transformation include horizontal flipping, rotation, scaling, horizontal movement, vertical movement, and projective transformation.

[0101] Examples of image processing include brightness change, color change, and geometric transformation. The image processing is not limited to applying one type at a time and can also apply multiple types in sequence. Moreover, it is also possible to directly output the skeleton-containing data 302 as the processed image data 303 to the over- and under-area correction unit 104 without performing any image processing.

[0102] The over- and under-area correction unit 104 (under-area correction unit) sets the area of pixels within the region of the original image size (the image size of the blank-completed image 301) in the processed image 303a of the input processed image data 303 that has no image information as the under-pixel region 313. The over- and under-area correction unit 104 also re-sets the region that is the same as the blank region 311 of the original blank-completed image 301 as the blank region 314. The over- and under-area correction unit 104 ignores the part of the region in the processed image data 303 that exceeds the original image size. The over- and under-area correction unit 104 corrects the under-pixel region 313 with a blank area (fills it with the same single color as the blank region 311) in the processed image 303a, thereby generating the over- and under-corrected image 304a.

[0103] The teaching data 304 includes the over- and under-corrected image 304a (second image) and teaching skeleton information 312 that has undergone the same geometric transformation as the image processing. The over- and under-area correction unit 104 outputs the teaching data 304 to the teaching data storage unit 201.

[0104] Figure 3In [the figure], the insufficient pixel region 313 is hatched diagonally to the upper right, but is actually filled with a single color that is the same as the blank region 311. Moreover, in the processing of the data expansion unit 103, since geometric transformation is performed as image processing, pixels that become image-less among the pixels of the image before image processing are also included in the insufficient pixel region 313. Further, pixels that exceed the original image region due to geometric transformation are not included in the over- and under-corrected image 304a.

[0105] The teaching data storage unit 201 stores the input teaching data 304. Moreover, based on an instruction from the control unit 10, the teaching data 304 is output to the feature quantity extraction unit 112 and the estimation model learning unit 122.

[0106] The estimation model acquisition unit 111 acquires the stored estimation model from the estimation model storage unit 202. The estimation model acquisition unit 111 outputs the estimation model (its multiple parameters) to the feature quantity extraction unit 112, the joint estimation unit 113, and the binding degree estimation unit 114.

[0107] The feature quantity extraction unit 112 extracts feature quantities from the over- and under-corrected image 304a or the blank-completed image 301 that constitutes the input teaching data 304. The feature quantity extraction unit 112 outputs the extracted feature quantities to the joint estimation unit 113 and the binding degree estimation unit 114.

[0108] The joint estimation unit 113 creates a joint estimation result representing the positions of multiple joints based on the input feature quantities. The joint estimation unit 113 outputs the joint estimation result to the skeleton estimation unit 121.

[0109] The binding degree estimation unit 114 obtains a binding degree estimation result representing the binding degree between joints based on the input feature quantities. The binding degree estimation unit 114 outputs the binding degree estimation result to the skeleton estimation unit 121.

[0110] The skeleton estimation unit 121 estimates the estimated skeleton information based on the input joint estimation result and binding degree estimation result. The skeleton estimation unit 121 outputs the estimated skeleton information to the estimation model learning unit 122 and the output unit 140.

[0111] The estimated skeleton information includes information on the joint positions corresponding to the operator of the image used for estimation (the over- and under-corrected image 304a or the blank-completed image 301). The estimated skeleton information may include the positions of a part of the joints (such as the neck or shoulders) located in the blank region (311 or 313) in the image used for estimation (the over- and under-corrected image 304a or the blank-completed image 301).

[0112] As Figure 4As shown, in the estimation device 1, nine joints, namely the neck, right shoulder, right elbow, right hand, left shoulder, left elbow, left hand, right waist, and left waist, are estimated as the joint estimation results. Moreover, in the estimation device 1, the degree of connection between eight joints, namely between the neck and the right shoulder, between the right shoulder and the right elbow, between the right elbow and the right hand, between the neck and the left shoulder, between the left shoulder and the left elbow, between the left elbow and the left hand, between the neck and the right waist, and between the neck and the left waist, is estimated as the connection degree estimation result.

[0113] The skeleton estimation unit 121 determines the positions of the respective joints based on the relative positional relationships of the joints, the connection rules set for each joint, and the relative strengths of the connection degrees between the joints. Particularly, in the case where the image for estimation includes multiple workers, the skeleton estimation unit 121 uses the connection degree estimation result to determine which worker each joint corresponds to.

[0114] The estimation model learning unit 122 compares the input estimated skeleton information with the teaching skeleton information 312 in the teaching data 304. If sufficient estimation accuracy is not obtained, learning continues, and the parameters of the feature quantity extraction unit 112, the joint estimation unit 113, and the connection degree estimation unit 114 are corrected. Subsequently, for learning again, the current feature quantity, the current joint estimation result, and the current connection degree estimation result are output to the feature quantity extraction unit 112, the joint estimation unit 113, and the connection degree estimation unit 114, causing them to perform processing again.

[0115] If sufficient estimation accuracy has been obtained, the estimation model learning unit 122 ends the learning, and stores the parameters of the feature quantity extraction unit 112, the joint estimation unit 113, and the connection degree estimation unit 114 as the estimation model in the estimation model storage unit 202.

[0116] The estimation model storage unit 202 stores the parameters of the feature quantity extraction unit 112, the joint estimation unit 113, and the connection degree estimation unit 114, that is, the estimation model. Moreover, the learned estimation model is output to the estimation model acquisition unit 111.

[0117] The output unit 140 performs the display of the estimated skeleton information by the estimation device 1 and the data output from the estimation device 1.

[0118] §3. Action Example

[0119] (Learning Process)

[0120] Based on Figure 5 to illustrate the learning process of the estimation device 1. Figure 5 is a flowchart showing the learning process of the estimation device 1.

[0121] The input unit 130 acquires the image 300 (S11) and outputs the image 300 to the blank area expansion unit 101. The blank area expansion unit 101 expands the image 300 with the blank area 311, thereby generating the blank-completed image 301 (S12). The blank area expansion unit 101 outputs the blank-completed image 301 to the teaching data production unit 102.

[0122] The teaching data production unit 102 associates the teaching skeleton information 312 corresponding to the image 300 with the blank-completed image 301, thereby producing the skeleton-containing data 302 (S13). The teaching data production unit 102 outputs the skeleton-containing data 302 to the data expansion unit 103.

[0123] The data expansion unit 103 applies image processing to the image (blank-completed image 301) of the skeleton-containing data 302, thereby generating the image-processed image 303a (S14). At this time, in image processing where the pixel coordinates of joint positions such as rotation, zooming, left-right movement, up-down movement, and projective transformation change, the data expansion unit 103 also performs the same deformation on the teaching skeleton information 312. Thus, the data expansion unit 103 generates the teaching skeleton information 312 corresponding to the image-processed image 303a. The data expansion unit 103 outputs the image-processed data 303 including the image-processed image 303a and the teaching skeleton information 312 to the over-and-under area correction unit 104.

[0124] The over-and-under area correction unit 104 sets the area within the region of the original image size (the image size of the blank-completed image 301) where there is no image information as the insufficient pixel area 313 for the image-processed image 303a of the image-processed data 303. The over-and-under area correction unit 104 corrects the over-and-under area (the insufficient pixel area 313 and the area exceeding the original image size region). The area exceeding the original image size region is ignored. The over-and-under area correction unit 104 fills the insufficient pixel area 313 with a single color that is the same color as the blank area. Moreover, the over-and-under area correction unit 104 also re-sets the area that is the same as the blank area 311 of the original blank-completed image 301 as the blank area 314 for the image-processed image 303a. Thus, the over-and-under area correction unit 104 produces the over-and-under corrected image 304a (S15). The insufficient pixel area 313 and the blank area 314 are the same color, but they can also be different colors.

[0125] The over-and-under area correction unit 104 produces the teaching data 304 by combining the geometrically deformed teaching skeleton information 312 and the over-and-under corrected image 304a (S15). The over-and-under area correction unit 104 stores the produced teaching data 304 in the teaching data storage unit (S16).

[0126] If the number of teaching data stored in the teaching data storage unit is less than the specified number (No in S17), the process of creating the teaching data 304 based on the skeleton-containing data 302 (data expansion process S14 to S16) is repeated again until the number of teaching data reaches the specified number. At this time, the image processing in the data expansion unit 103 is different each time. The type and amount of change of the image processing can be determined using random number parameters. In addition, if no image processing is performed, the skeleton-containing data 302 becomes the teaching data 304.

[0127] If the number of teaching data stored in the teaching data storage unit 201 is equal to or more than the specified number (Yes in S17), the control unit 10 ends the process of creating the teaching data and transfers to the learning process.

[0128] The learning unit 11 reads the teaching data from the teaching data storage unit 201 and performs the learning process. The inference model learning unit 122 obtains an inference model including a plurality of parameters from the inference model storage unit 202. The inference model is a model that takes an image as input and outputs the inferred skeleton information. At the time of non-learning, the inference model includes initial parameters. The inference model learning unit 122 outputs corresponding multiple parameters to the feature amount extraction unit 112, the joint inference unit 113, and the binding degree inference unit 114.

[0129] The feature amount extraction unit 112 extracts feature amounts from the over- and under-correction completed image 304a in the teaching data 304 using the feature amount extraction parameters (S18). The joint inference unit 113 uses the joint inference parameters to obtain a joint inference result indicating the joint position based on the extracted feature amounts (S19). The binding degree inference unit 114 uses the binding degree inference parameters to obtain a binding degree inference result indicating the binding degree between joints based on the extracted feature amounts (S20). The skeleton inference unit 121 obtains the inferred skeleton information based on the joint inference result and the binding degree inference result (S21).

[0130] The inference model learning unit 122 determines whether there is sufficient accuracy of the inferred skeleton information with respect to the teaching skeleton information 312 in the teaching data 304 (S22). If the difference between the inferred skeleton information and the teaching skeleton information is within a specified reference, the inference model learning unit 122 determines that there is sufficient accuracy.

[0131] If sufficient accuracy is not achieved (No in S22), the estimation model learning unit 122 corrects the parameters of the feature quantity extraction unit 112, the joint estimation unit 113, and the combination degree estimation unit 114 (S23). When correcting, the estimation model learning unit 122 corrects the parameters (learns the estimation model) to reduce the error between the estimated skeleton information and the teaching skeleton information 312. The estimation model learning unit 122 outputs the corrected parameters to the feature quantity extraction unit 112, the joint estimation unit 113, and the combination degree estimation unit 114. Subsequently, the learning unit 11 repeats the processes of S18 to S21. This process is performed for a plurality of teaching data 304.

[0132] If sufficient accuracy is achieved (Yes in S22), the learning process ends, and the estimation model learning unit 122 stores the learned estimation model in the estimation model storage unit 202 (S24). The output unit 140 displays a message indicating that the learning has ended on the display device.

[0133] Figure 5 In S22, a trigger for ending the learning process with sufficient accuracy is adopted, but it is not limited to this. The learning can also end after a specified number of repeated learning (the processes from S18 to S23).

[0134] Moreover, as a similar process from S18 to S24, OpenPose can also be used.

[0135] (Estimation process)

[0136] Based on Figure 6 the estimation process of the estimation device 1 will be described. Figure 6 is a flowchart showing the estimation process of the estimation device 1.

[0137] The input unit 130 acquires the image 300 (S31) and outputs the image 300 to the blank area expansion unit 101. The blank area expansion unit 101 expands the image 300 with the blank area 311, thereby generating the blank-completed image 301 (S32). The blank area expansion unit 101 outputs the blank-completed image 301 to the feature quantity extraction unit 112.

[0138] The estimation model acquisition unit 111 reads the learned estimation model from the estimation model storage unit 202 (S33). The estimation model acquisition unit 111 outputs a plurality of parameters included in the learned estimation model to the feature quantity extraction unit 112, the joint estimation unit 113, and the combination degree estimation unit 114.

[0139] The feature quantity extraction unit 112 extracts feature quantities from the blank-completed image 301 using the feature quantity extraction parameters (S34). The joint estimation unit 113 uses the joint estimation parameters to obtain a joint estimation result indicating the joint positions based on the extracted feature quantities (S35). The degree of combination estimation unit 114 uses the degree of combination estimation parameters to obtain a degree of combination estimation result indicating the degree of combination between joints based on the extracted feature quantities (S36). The skeleton estimation unit 121 obtains estimated skeleton information based on the joint estimation result and the degree of combination estimation result (S37). Subsequently, the output unit 140 displays the skeleton information on the display device (S38). Additionally, the output unit 140 may output the skeleton information to an external server.

[0140] Moreover, as similar processing from S33 to S38, OpenPose can also be used.

[0141] §4. Function / Effect

[0142] As described above, after the input unit 130 of the estimation device 1 in the first embodiment acquires the image 300, a blank area 311 is added to the image 300, and the teaching skeleton information 312 input by the user is taught. Subsequently, the data expansion unit 103 and the excess and deficiency area correction unit 104 create teaching data 304 until the number of data required for machine learning is obtained. Machine learning is performed using the created teaching data 304 to learn the estimation model.

[0143] After learning, after the input unit 130 acquires a new image 300, the learned estimation model is used to estimate the estimated skeleton information based on the blank-completed image 301 expanded with the blank area 311.

[0144] In this way, the estimation device 1 generates a blank-completed image 301 or an excess and deficiency corrected image 304a with blank areas 311 and 314 added to the image 300 (the first image) lacking the second joint as the second image. The estimation device 1 uses the second image including the blank area as the input to the estimation model, and uses the teaching skeleton information including the joint positions of the second joint located in the blank area as the output of the estimation model for learning. Thus, the estimation device 1 can create an estimation model that can estimate the skeleton information including the joint positions of the second joint based on the image 300 lacking the second joint. The estimation device 1 expands the image range to the area where the second joint is considered to be located for learning, and thereby can appropriately estimate the joint positions of the second joint in the blank area 311 without image information. Therefore, the estimation device 1 can use the learned estimation model to estimate the estimated skeleton information including the joint positions of the first joint and the second joint based on the image 300 lacking the second joint.

[0145] In the prior art, for example, if the joint positions of the neck or shoulders connecting the right arm and the left arm cannot be estimated (if the image does not include the joint positions of the neck or shoulders), it is impossible to stably and accurately estimate the joint positions of the right arm and the left arm.

[0146] Even if the image does not include the joint positions of the neck (or shoulders), the estimation device 1 can estimate the joint positions of the neck (or shoulders) connecting the right arm and the left arm together with the joint positions of the right arm and the left arm. Thus, the joint positions of the right arm and the left arm can also be stably and accurately estimated.

[0147] Moreover, the input unit can input privacy-conscious images lacking the head, neck, shoulders, etc., and images showing the situation on the workbench. By using these images, in addition to the time change of the skeleton information caused by the time change of the subject, the time change of the work object on the workbench can also be analyzed together. Through their simultaneous analysis, the work analysis of the subject can be automated.

[0148] (Modification example)

[0149] In addition, the joint estimation unit 113 and the combination degree estimation unit 114 can also be configured in multiple stages. For example, after obtaining the joint estimation result and the combination degree estimation result based on the feature amount, the joint estimation unit 113 and the combination degree estimation unit 114 can use the joint estimation result and the combination degree estimation result to perform estimation again. At this time, the joint estimation unit 113 outputs the first joint estimation result to the combination degree estimation unit 114, and the combination degree estimation unit 114 outputs the first combination degree estimation result to the joint estimation unit 113. The joint estimation unit 113 uses the feature amount, the first joint estimation result, and the first combination degree estimation result to obtain the second joint estimation result. The combination degree estimation unit 114 uses the feature amount, the first joint estimation result, and the first combination degree estimation result to obtain the second combination degree estimation result. At this time, the joint estimation unit 113 and the combination degree estimation unit 114 use parameters different from the first time. Moreover, the joint estimation unit 113 and the combination degree estimation unit 114 can also perform estimation processing three or more times. Thus, the estimation accuracy is improved.

[0150] 〔Embodiment 2〕

[0151] Hereinafter, based on Figures 7 to 9 Another embodiment of the present invention will be described. In addition, for the sake of convenience of explanation, components having the same functions as those described in the above embodiment are denoted by the same reference numerals and their descriptions will not be repeated.

[0152] §1. Structural example

[0153] Based on Figure 7 the structure of the teaching data creation device 2 of the present embodiment will be described.Figure 7 This is a block diagram showing an example of the main part structure of the teaching data production device 2.

[0154] As Figure 7 shown, the teaching data production device 2 includes: a control unit 10 that comprehensively controls each part of the learning data production device; and a storage unit 20 that stores various data used by the teaching data production device 2. Additionally, the storage unit 20 can also be a machine external to the teaching data production device 2.

[0155] The control unit 10 includes a blank area expansion unit 101, a teaching data production unit 102, an input unit 130, and a display control unit 150.

[0156] The display control unit 150 has the function of displaying the state of the teaching data production device 2 and displaying images. The display device (not shown) that is the object controlled by the display control unit 150 can also be a machine external to the teaching data production device 2.

[0157] Based on Figure 8 this, the user interface of the learning data production device of this embodiment will be described. Figure 8 This is an example of the user interface 51 for designating joint positions in the teaching data production unit.

[0158] The user interface 51 includes an image display unit 501, an operator list 511, an operator addition button 513, a joint list 521, and a coordinate display indicator 531. The display of the user interface 51 is controlled by the display control unit 150.

[0159] The image display unit 501 is an area for displaying the completed blank image 301 following the instructions of the display control unit 150. Moreover, it displays the edited joint position 541, the determined joint position 542, and the combination information 543 input through user input. The edited joint position 541 and the determined joint position 542 are displayed differently so as to be distinguishable. The combination information 543 is a display of the combination between joints indicating the foregoing joint combination relationship. Figure 8 In

[0160] it, an arrow is used for display, but it is not limited to this, and it can also be a line segment.

[0161] The joint list 521 is a list of joints for which the operator must make settings. The joints to be set are nine joints: the neck, right shoulder, right elbow, right hand, left shoulder, left elbow, left hand, right waist, and left waist, but are not limited thereto, and it is also possible to enable setting of the head or feet. The joints in the joint list for which settings are made (corresponding to the edit joint position 541) can also be highlighted as edit joints 522.

[0162] The coordinate display indicator 531 indicates the pixel coordinates in the setting of the edit joint 522, and includes an X coordinate display indicator 532 for the X coordinate and a Y coordinate display indicator 533 for the Y coordinate. For both the X coordinate and the Y coordinate, in addition to displaying the current set pixel, user input is also accepted. When the coordinates change due to user input, the display control unit 150 also changes the display position of the edit joint position 541.

[0163] When the user changes the coordinates of the edit joint 522 using the input unit 130 and the coordinate display indicator 531, the coordinates of the entry into the blank area 311 can also be specified as the coordinates.

[0164] §2. Action Example

[0165] Based on Figure 9 the operation of the teaching data production device 2 of the present embodiment will be described. Figure 9 is a flowchart of the operation of the teaching data production device 2.

[0166] The input unit 130 acquires the image 300 and outputs the image 300 to the blank area expansion unit 101 (S41). The blank area expansion unit 101 attaches the blank area 311 adjacent to the edge of the input image 300 and produces the blank-completed image 301. The blank area expansion unit 101 outputs the blank-completed image 301 to the teaching data production unit 102 and the display control unit 150 (S42).

[0167] The display control unit 150 causes the display device to display the user interface 51. The display control unit 150 displays the blank-completed image 301 on the image display unit 501 (S43). The input unit 130 accepts inputs from the user via a mouse, keyboard, touch panel, etc. For example, the input unit 130 accepts the addition made by the operator through the operator addition button 513, the selection made by the operator in the operator list 511, the selection of the joint to be edited in the joint list 521, and the specification of the joint position to be edited on the coordinate display indicator 531 or the image display unit 501 (S44). At this time, the input unit 130 accepts not only the specification of the joint position of the joint located on the image 300, but also the specification of the joint position of the joint located in the blank area 311. The input unit 130 outputs the information of the operator, joint, and joint position inputted in association with each other to the teaching data creation unit 102.

[0168] The teaching data creation unit 102 changes the display performed by the display control unit 150 based on the user input. Moreover, the teaching data creation unit 102 generates teaching skeleton information 312 including information of one or more operators, multiple joints, and multiple joint positions. The multiple joints include a first joint located in the area of the image 300 and a second joint located in the blank area 311. The teaching data creation unit 102 associates the teaching skeleton information 312 with the blank-completed image 301 to create the skeleton-containing data 302 (S45).

[0169] The teaching data creation unit 102 stores the skeleton-containing data 302 in the teaching data storage unit 201.

[0170] §3. Function / Effect

[0171] As described above, after the image 300 is inputted to the teaching data creation device 2 of the present Embodiment 2, the blank area 311 is expanded, and the blank-completed image 301 is displayed on the user interface 51. The teaching data creation device 2 accepts the input of the joint position of the joint corresponding to the blank area 311 from the user, and thus, can create the teaching skeleton information 312 corresponding to the blank-completed image 301.

[0172] The joint position constituting the teaching skeleton information can also be set in the blank area 311. Thus, it is possible to input an image considering privacy lacking a head, neck, shoulders, etc. using the input unit, and thus, it is possible to create the teaching skeleton information 312 which is the learning data used in Embodiment 1.

[0173] 〔Embodiment 3〕

[0174] Hereinafter, based on Figure 10 、 Figure 11To illustrate another embodiment of the present invention. In addition, for the sake of convenience in explanation, components having the same functions as those described in the above embodiment are denoted by the same reference numerals and their descriptions will not be repeated.

[0175] Based on Figure 10 、 Figure 11 To illustrate an example of the input image of this embodiment. Figure 10 It is a schematic diagram of an image of an operator taken from above. Figure 11 It is a schematic diagram of an image of an operator taken from the side.

[0176] (Image taken from above)

[0177] As Figure 10 shown, by taking a picture from above, an image 300a can be taken that does not reflect the face of the operator and can easily grasp the situation on the workbench. Therefore, by taking a picture from above, an image suitable for work analysis can be obtained while protecting privacy.

[0178] The image 300a includes an operator 601, a workbench 602, and a work object 603. Moreover, blank areas 311a, 311b, 311c, and 311d extend adjacent to the four sides of the image 300a. Figure 10 In this case, blank areas are set to be adjacent to all four sides, but it is not limited to this. As long as blank areas are set to be adjacent to any one or more sides, it is acceptable.

[0179] The operator 601 is the operator who is the object of work analysis. Elements related to privacy such as the face of the operator may not be included in the image. The workbench 602 is the space where the operator performs work. It can be not only a workbench with a single-layer shelf but also a table with multiple-layer shelves. The work object 603 is the object on which the operator performs work. For the workbench 602 or the work object 603, marks or the like may be added for work analysis.

[0180] As Figure 10As shown, by expanding the blank area 311a at the upper part of the image 300a, the face of the operator or the like is not reflected, and thus the skeleton information can be inferred. By expanding the blank area 311b at the right side of the image 300a, even when the operator moves to the left (the right side in the image) or reaches for a work object in the space outside the image 300a, the skeleton information can be inferred. By expanding the blank area 311c at the lower side of the image 300a, even when the operator reaches for a work object outside the image 300a, the skeleton information can be inferred. By expanding the blank area 311d at the left side of the image 300a, even when the operator moves to the right (the left side in the image) or reaches for a work object in the space outside the image 300a, the skeleton information can be inferred.

[0181] Therefore, by taking a picture from above, the skeleton information can be inferred while protecting privacy, and the work analysis on the flat workbench can be made easier. Moreover, by attaching a plurality of blank areas adjacent to multiple sides of the image 300a, even when the operator moves outside the image 300a, the skeleton information can be inferred.

[0182] (Side view image)

[0183] As Figure 11 shown, by taking a picture from the side, an image 300b that does not reflect the face of the operator and can easily grasp the state of the multi-layer workbench can be taken. Therefore, by taking a picture from the side, an image suitable for work analysis can be obtained while protecting privacy.

[0184] The image 300b includes an operator 601, a workbench 602, a work object 603, and a shielding area 604. Moreover, blank areas 311e, 311f, 311g, and 311h are expanded adjacent to the four sides of the image 300a. Figure 11 In [description], the blank areas are set to be adjacent to all four sides, but it is not limited to this. As long as the blank areas are set to be adjacent to any one or more sides.

[0185] The shielding area 604 can also be set when it is necessary to reflect elements that invade privacy such as the face of the operator according to the installation position and installation direction of the camera. The shielding area 604 is painted with the same single color as the blank area 311. Moreover, it can be set to any size at any position adjacent to the blank area.

[0186] As Figure 11As shown, by expanding the blank area 311e at the upper part of the image 300b, even when parts or tools are arranged on the multi-layered shelves or on the upper outer side of the image 300b, the skeleton information can be inferred. By expanding the blank area 311f on the right side of the image 300b, the skeleton information when reaching for an operation object arranged on the inner side as observed by the operator on the workbench can be inferred. By expanding the blank area 311g on the lower side of the image 300b, the skeleton information when reaching for an operation object arranged on the lower side of the workbench can be inferred. By expanding the blank area on the left side of the image 300b, the skeleton information can be inferred without reflecting the face of the operator or the like.

[0187] Therefore, by taking a side view, the skeleton information can be inferred while protecting privacy, and the operation analysis on the three-dimensional workbench can be made easier.

[0188] As the shooting angle, it is not limited to the top and the side, and it can also be taken from the upper rear side of the operator. At this time, the planar operation analysis at the time of top shooting and the three-dimensional operation analysis at the time of side shooting can be achieved simultaneously.

[0189] 〔Implementation example with software〕

[0190] The control blocks of the inference device 1 and the teaching data production device 2 (especially the blank area expansion unit 101, the teaching data production unit 102, the data expansion unit 103, the over and under area correction unit 104, the inference model acquisition unit 111, the feature amount extraction unit 112, the joint inference unit 113, the combination degree inference unit 114, the skeleton inference unit 121, the inference model learning unit 122, the input unit 130, the output unit 140, and the display control unit 150 in the control unit 10) can be implemented either by a logic circuit (hardware) formed on an integrated circuit (IC (Integrated Circuit) chip) or the like, or by software.

[0191] In the latter case, the estimation device 1 and the teaching data creation device 2 include a computer that executes commands of software, i.e., a program, for implementing each function. The computer includes, for example, one or more processors, and includes a computer-readable recording medium storing the program. In the computer, the program is read from the recording medium and executed by the processor, thereby achieving the object of the present invention. As the processor, for example, a Central Processing Unit (CPU) can be used. As the recording medium, a "non-transitory tangible medium" can be used. For example, in addition to a Read Only Memory (ROM) and the like, a tape, a disk, a card, a semiconductor memory, a programmable logic circuit, etc. can also be used. Moreover, a Random Access Memory (RAM) for expanding the program and the like can be further included. Further, the program can also be provided to the computer via any transmission medium (such as a communication network or a broadcast wave) capable of transmitting this program. Additionally, an embodiment of the present invention can also be implemented in the form of a data signal embedded in a carrier wave that realizes the program by electronic transmission.

[0192] 〔Summary〕

[0193] To solve the above problems, an estimation device according to an embodiment of the present invention includes: an image acquisition unit that acquires a first image including a first joint of a subject and not including a second joint; a blank area expansion unit that generates a second image obtained by expanding the first image with a blank area; and an estimation unit that estimates skeleton information using the second image and a learned estimation model, the skeleton information including joint positions of the second joint located in the blank area.

[0194] According to the above structure, by adding a blank area to the first image lacking the second joint, it is possible to stably estimate the skeleton information including the joint positions of the second joint located in the blank area.

[0195] The blank area expansion unit may also make the blank area adjacent to one side of the first image.

[0196] According to the above structure, it is possible to estimate the joint positions of the second joint of the subject located at a position exceeding one side of the first image.

[0197] The second joint may also include a joint of the neck.

[0198] According to the above structure, the first image can be set, for example, to an image that reflects the shoulders, elbows, hands, and waist but does not reflect the neck. Therefore, the first image that does not reflect the face of the subject can be used to stably and accurately infer the skeleton information including the neck.

[0199] The image acquisition unit may also acquire the first image taken from above.

[0200] According to the above structure, by using the image taken from above, for example, the face of the subject is not reflected, and operation analysis can be performed by using the skeleton information of the operation object on the workbench and the subject (operator) together.

[0201] To solve the above problems, a learning device according to an embodiment of the present invention includes: an image acquisition unit that acquires a first image including a first joint of a subject and not including a second joint; a blank area expansion unit that generates a second image obtained by expanding the first image with a blank area; a teaching data storage unit that stores teaching data including skeleton information and the second image, where the skeleton information includes the second joint located in the blank area; and a learning unit that uses the teaching data to learn an inference model of skeleton information based on the skeleton information and the second image.

[0202] According to the above structure, by adding a blank area to the first image lacking the second joint, the second joint located in the blank area is included in the teaching data, so that the skeleton information having the second joint in the blank area can be learned.

[0203] The blank area expansion unit may also make the blank area adjacent to one side of the first image.

[0204] The second joint may also include the joints of the neck.

[0205] The learning device may also include: a data expansion unit that performs image processing of geometric transformation on the second image to generate a third image; and a deficiency area correction unit that corrects the deficiency pixel area in the area corresponding to the second image in the third image with a blank area as a new second image for learning.

[0206] According to the above structure, even based on a small number of images, a plurality of teaching data can be produced, so that learning can be carried out efficiently.

[0207] The image acquisition unit may also acquire the first image taken from above.

[0208] According to the above structure, by using the image taken from above, the inference model can be learned based on the skeleton information of the operation object on the workbench and the subject (operator).

[0209] In order to solve the above problems, a teaching data production device according to an embodiment of the present invention includes: an image acquisition unit that acquires a first image including a first joint of a subject and not including a second joint; a blank area expansion unit that generates a second image obtained by expanding the first image with a blank area; a display control unit that displays the second image; an input unit that receives an input of the joint position of the second joint from a user for the blank area in the second image; and a teaching data production unit that produces teaching data associating skeleton information including the joint positions of the first joint and the second joint with the second image.

[0210] According to the above structure, by adding a blank area to the first image lacking the second joint, teaching data including the joint position of the second joint located in the blank area can be produced.

[0211] The blank area expansion unit may also make the blank area adjacent to one side of the first image.

[0212] The second joint may also include a neck joint.

[0213] A speculation method according to an embodiment of the present invention includes: an image acquisition step of acquiring a first image including a first joint of a subject and not including a second joint; a blank area expansion step of generating a second image obtained by expanding the first image with a blank area; and a speculation step of using the second image and a learned speculation model to speculate skeleton information, the skeleton information including the joint position of the second joint located in the blank area.

[0214] A learning method according to an embodiment of the present invention includes: an image acquisition step of acquiring a first image including a first joint of a subject and not including a second joint; a blank area expansion step of generating a second image obtained by expanding the first image with a blank area; a teaching data acquisition step of acquiring teaching data including skeleton information and the second image, the skeleton information including the second joint located in the blank area; and a learning step of using the teaching data to learn a speculation model of skeleton information based on the skeleton information and the second image.

[0215] A teaching data production method according to an embodiment of the present invention includes: an image acquisition step of acquiring a first image including a first joint of a subject and not including a second joint; a blank area expansion step of generating a second image obtained by expanding the first image with a blank area; a display control step of displaying the second image; an input step of receiving an input of the joint position of the second joint from a user for the blank area in the second image; and a teaching data production step of producing teaching data associating skeleton information including the joint positions of the first joint and the second joint with the second image.

[0216] The inference device of each embodiment of the present invention can also be implemented by a computer. In this case, by causing the computer to operate as each part (software element) included in the inference device, the inference program of the inference device implemented by the computer and a computer-readable recording medium recording the inference program are also included in the scope of the present invention.

[0217] The learning device of each embodiment of the present invention can also be implemented by a computer. In this case, by causing the computer to operate as each part (software element) included in the learning device, the learning program of the learning device implemented by the computer and a computer-readable recording medium recording the learning program are also included in the scope of the present invention.

[0218] The teaching data production device of each embodiment of the present invention can also be implemented by a computer. In this case, by causing the computer to operate as each part (software element) included in the teaching data production device, the teaching data production program of the teaching data production device implemented by the computer and a computer-readable recording medium recording the teaching data production program are also included in the scope of the present invention.

[0219] 〔Supplementary Notes〕

[0220] The present invention is not limited to the above-described embodiments, and various modifications can be made within the scope shown in the claims. Embodiments obtained by appropriately combining technical components separately disclosed in different embodiments are also included in the technical scope of the present invention.

Claims

1. A speculation device, comprising: An image acquisition unit that acquires a first image including a first joint of a subject and not including a second joint; A blank area expansion unit that generates a second image obtained by expanding the first image with a blank area; And A speculation unit that uses the second image and a learned speculation model to speculate skeleton information, the skeleton information including the joint position of the second joint located in the blank area.

2. The speculation device according to claim 1, wherein The blank area expansion unit makes the blank area adjacent to one side of the first image.

3. The speculation device according to claim 1 or 2, wherein The second joint includes a joint of the neck.

4. The speculation device according to claim 1 or 2, wherein The image acquisition unit acquires the first image taken from above.

5. A learning device, comprising: An image acquisition unit that acquires a first image including a first joint of a subject and not including a second joint; A blank area expansion unit that generates a second image obtained by expanding the first image with a blank area; A teaching data storage unit that stores teaching data including skeleton information and the second image, the skeleton information including the second joint located in the blank area; And A learning unit that uses the teaching data to learn a speculation model of skeleton information based on the skeleton information and the second image.

6. The learning device according to claim 5, wherein The blank area expansion unit makes the blank area adjacent to one side of the first image.

7. The learning device according to claim 5 or 6, wherein The second joint includes a joint of the neck.

8. The learning device according to claim 5 or 6, wherein The learning device includes: A data expansion unit that performs image processing of geometric deformation on the second image to generate a third image; And A shortage area correction unit that corrects a shortage pixel area in a region corresponding to the second image in the third image with a blank area as a new second image for learning.

9. The learning device according to claim 5 or 6, wherein The image acquisition unit acquires the first image taken from above.

10. A teaching data production device, comprising: An image acquisition unit that acquires a first image including a first joint of a subject and not including a second joint; A blank area expansion unit that generates a second image obtained by expanding the first image with a blank area; A display control unit that displays the second image; An input unit that receives an input of the joint position of the second joint from a user for the blank area in the second image; And A teaching data production unit that produces teaching data associating skeleton information including the joint positions of the first joint and the second joint with the second image.

11. The teaching data production device according to claim 10, wherein The blank area expansion unit makes the blank area adjacent to one side of the first image.

12. The teaching data production device according to claim 10 or 11, wherein The second joint includes a joint of the neck.

13. A speculation method, comprising: An image acquisition step of acquiring a first image that includes the first joint of the person of interest and does not include the second joint; A blank area expansion step of generating a second image obtained by expanding the first image with a blank area; And A speculation step of speculating skeleton information using the second image and a learned speculation model, the skeleton information including the joint position of the second joint located in the blank area.

14. A learning method, comprising: An image acquisition step of acquiring a first image that includes the first joint of the person of interest and does not include the second joint; A blank area expansion step of generating a second image obtained by expanding the first image with a blank area; A teaching data acquisition step of acquiring teaching data including the skeleton information and the second image, the skeleton information including the second joint located in the blank area; And A learning step of using the teaching data to learn a speculation model of the skeleton information based on the skeleton information and the second image.

15. A teaching data production method, comprising: An image acquisition step of acquiring a first image that includes the first joint of the person of interest and does not include the second joint; A blank area expansion step of generating a second image obtained by expanding the first image with a blank area; A display control step of displaying the second image; An input step of receiving an input of the joint position of the second joint from the user for the blank area in the second image; And A teaching data production step of producing teaching data associating the skeleton information including the joint positions of the first joint and the second joint with the second image.

16. A storage medium stores a speculation program for causing a computer to function as the speculation device according to claim 1, wherein, The speculation program is used to cause a computer to function as the image acquisition unit, the blank area expansion unit, and the speculation unit.

17. A storage medium stores a learning program for causing a computer to function as the learning device according to claim 5, wherein, The learning program is used to cause a computer to function as the image acquisition unit, the blank area expansion unit, and the learning unit.

18. A storage medium stores a teaching data production program for causing a computer to function as the teaching data production device according to claim 10, wherein, The teaching data production program is used to cause a computer to function as the image acquisition unit, the blank area expansion unit, the display control unit, the input unit, and the teaching data production unit.

Citation Information

Patent Citations

  • Human body attitude estimation method based on cascade error correction mechanism

    CN107220596A

  • A pedestrian re-identification method combining deep learning and metric learning

    CN109447175A