Skeleton estimation apparatus

The skeletal estimation device corrects three-dimensional coordinate errors by adjusting for deviations in installation height, distance, and inclination, ensuring accurate skeletal point estimation without model reconstruction.

JP2025121083APending Publication Date: 2025-08-19TOYOTA INDUSTRIES CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024016291
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-06
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Discrepancies in the installation height, distance, and inclination of the imaging device relative to the subject cause errors in the three-dimensional coordinates of skeleton points.

Method used

A skeletal estimation device that includes an imaging device, a storage device for a trained model, and a control unit to correct three-dimensional coordinates based on deviations in installation height, distance, and inclination, using a reference tilt from training data to minimize errors.

Benefits of technology

Reduces errors in three-dimensional coordinates of skeleton points by geometrically correcting them based on angular misalignment, without the need to reconstruct the model, thus maintaining accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025121083000001_ABST
    Figure 2025121083000001_ABST
Patent Text Reader

Abstract

To reduce an error of three-dimensional coordinates of a skelton point.SOLUTION: A skeleton estimation apparatus according to the present invention has an imaging device for capturing a subject to generate an image data, a storage device for deriving a two-dimensional coordinate of the skeleton point of the subject captured in the image data and for storing a leaned model for use in conversion of the derived two-dimensional coordinate into a three-dimensional coordinate, and a control unit. The control unit acquires image data from the imaging device, derives the two-dimensional coordinate of the skeleton point of the subject captured in the image data, and inputs the derived two-dimensional coordinate to the learned model, thus acquiring the three-dimensional coordinate. The control unit corrects the acquired three-dimensional coordinate on the basis of difference between an inclination of the imaging device relative to the subject and reference inclination. The reference inclination is an inclination of a virtual camera relative to the virtual subject. The virtual camera is virtually set based on a learned two-dimensional coordinate which is a two-dimensional coordinate used upon generation of the learned model.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a skeleton estimation device. [Background technology]

[0002] The task estimation device disclosed in Patent Document 1 estimates the worker's posture by using a trained model. The task estimation device then calculates a load value for the worker's posture. The trained model receives image data as input and outputs three-dimensional coordinates of skeleton points of a subject depicted in the image data. The task estimation device estimates the worker's posture from the three-dimensional coordinates of the skeleton points. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-48017 Summary of the Invention [Problem to be solved by the invention]

[0004] There may be a discrepancy in at least one of the following: the installation height of the imaging device and the reference installation height; the distance between the imaging device and the subject and the reference distance; and the inclination of the imaging device relative to the subject and the reference inclination. In this case, an error occurs in the three-dimensional coordinates. The reference installation height is the installation height of the virtual camera assumed from the training two-dimensional coordinates, which are the two-dimensional coordinates used when creating the trained model. The reference distance is the distance between the virtual subject having the training two-dimensional coordinates and the virtual camera. The reference inclination is the inclination of the virtual camera relative to the virtual subject. [Means for solving the problem]

[0005] A skeletal estimation device that solves the above problem comprises an imaging device that images a subject and generates image data; a storage device that derives two-dimensional coordinates of skeletal points of the subject that appear in the image data and stores a trained model that converts the derived two-dimensional coordinates into three-dimensional coordinates; and a control unit that acquires the three-dimensional coordinates converted by the trained model and corrects the three-dimensional coordinates based on at least one of the deviation between the installation height of the imaging device and a reference installation height, the deviation between the distance between the imaging device and the subject and a reference distance, and the deviation between the inclination of the imaging device with respect to the subject and a reference inclination, where the reference installation height is the installation height of a virtual camera that is assumed from training data used when generating the trained model, the reference distance is the distance between the training subject of the training data and the virtual camera, and the reference inclination is the inclination of the virtual camera with respect to the training subject.

[0006] The control unit corrects the three-dimensional coordinates based on at least one of the following: a difference between the installation height of the image capture device and a reference installation height; a difference between the distance between the image capture device and the subject and a reference distance; and a difference between the inclination of the image capture device with respect to the subject and a reference inclination. Even if an error occurs in the three-dimensional coordinates of the skeleton points output by the trained model due to these differences, this error can be reduced.

[0007] A skeletal estimation device that solves the above problem comprises an imaging device that images a subject and generates image data; a storage device that derives two-dimensional coordinates of skeletal points of the subject that appear in the image data and stores a trained model that converts the derived two-dimensional coordinates into three-dimensional coordinates; and a control unit that calculates a coefficient to correct the trained model based on at least one of the deviation between the installation height of the imaging device and a reference installation height, the deviation between the distance between the imaging device and the subject and the reference distance, and the deviation between the inclination of the imaging device with respect to the subject and the reference inclination, and corrects the trained model using the calculated coefficient, wherein the reference installation height is the installation height of a virtual camera assumed from training data used when generating the trained model, the reference distance is the distance between the training subject of the training data and the virtual camera, and the reference inclination is the inclination of the virtual camera with respect to the training subject.

[0008] The control unit corrects the trained model based on at least one of the deviation between the installation height of the image capture device and the reference installation height, the deviation between the distance between the image capture device and the subject and the reference distance, and the deviation between the inclination of the image capture device with respect to the subject and the reference inclination. As a result, the trained model outputs three-dimensional coordinates with small errors. Therefore, it is possible to reduce errors in the three-dimensional coordinates of the skeleton points.

[0009] In the skeleton estimation device, the three-dimensional coordinates may be coordinates of a three-axis Cartesian coordinate system in which the vertical direction is the Y-axis, the axis perpendicular to the Y-axis and pointing from the imaging device to the subject is the X-axis, and the axis perpendicular to the X-axis and the Y-axis is the Z-axis.

[0010] In the skeleton estimation device, the control unit may evaluate a posture of the subject based on the three-dimensional coordinates. [Effects of the Invention]

[0011] According to the present invention, errors in the three-dimensional coordinates of skeleton points can be reduced. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a schematic diagram of a skeletal structure estimation device. [Figure 2] FIG. 2 is a diagram showing the positional relationship between the imaging device and the subject. [Figure 3] FIG. 3 is a diagram showing image data input to the skeleton estimation model and the skeleton model output. [Figure 4] FIG. 4 is a diagram showing the positional relationship between the virtual camera and the virtual subject. [Figure 5] FIG. 5 is a flowchart showing the control performed by the control unit. DETAILED DESCRIPTION OF THE INVENTION

[0013] An embodiment of a skeleton estimation device will be described. <Bone structure estimation device> As shown in FIGS. 1 and 2, a skeleton estimation device 10 includes an imaging device 20 and an information processing device 30. The imaging device 20 captures an image of a subject M1 and generates image data, and is disposed so as to capture the image of the subject M1. The position and angle of the imaging device 20 are fixed. The subject M1 in this embodiment is, for example, a person performing work in a predetermined work area A1. The work is, for example, assembly work or inspection work in a factory. The imaging device 20 is disposed so as to capture an image of the work area A1 in which the subject M1 performs work.

[0014] Subject M1 works within working area A1. Therefore, the position where subject M1 works can be considered to be a fixed position. In the following description, the distance between image capture device 20 and subject M1, the installation height of image capture device 20, and the inclination of image capture device 20 relative to subject M1 are assumed to be constant.

[0015] The distance between the imaging device 20 and the subject M1 (hereinafter referred to as the actual distance) is the distance between a predetermined point P1 on the imaging device 20 and a predetermined point P2 on the subject M1. The predetermined point P1 on the imaging device 20 can be set arbitrarily. The predetermined point P2 on the subject M1 can be set arbitrarily. The actual distance is, for example, the Euclidean distance in the horizontal direction between the predetermined point P1 on the imaging device 20 and the predetermined point P2 on the subject M1. In a three-axis Cartesian coordinate system described below, the Z coordinate of the predetermined point P1 on the imaging device 20 and the Z coordinate of the predetermined point P2 on the subject M1 are the same. In this case, since the Euclidean distance does not include a Z-axis component, the actual distance is the distance in the X-axis direction in the three-axis Cartesian coordinate system described below. When point P3 is the intersection of a virtual horizontal line L11 extending from point P1 toward subject M1 and a virtual orthogonal line L12 that is perpendicular to virtual horizontal line L11 and extends from point P2 in the Y-axis direction, the distance in the X-axis direction is the length of the virtual line connecting point P1 and point P3.

[0016] The installation height of the imaging device 20 (hereinafter referred to as the actual height) is the height from the surface directly below the imaging device 20 or the surface on which the subject M1 is standing to a point P1 of the imaging device 20 when the surface directly below the imaging device 20 and the surface on which the subject M1 is standing are at the same height. When the surface directly below the imaging device 20 and the surface on which the subject M1 is standing are at different heights, the installation height of the imaging device 20 is the height from the surface on which the subject M1 is standing to the point P1 of the imaging device 20. In other words, the actual height can be said to be the height from the surface on which the subject M1 is standing to the point P1 of the imaging device 20. In the example shown in FIG. 1, the surface directly below the imaging device 20 and the surface on which the subject M1 is standing are horizontal and are the same surface. The actual height is the difference between the Y coordinate of the point P1 of the imaging device 20 and the Y coordinate of the surface on which the subject M1 is standing in a three-axis Cartesian coordinate system described below.

[0017] The tilt of the imaging device 20 with respect to the subject M1 (hereinafter referred to as the actual tilt) is, for example, the tilt of the optical axis 20A of the imaging device 20 with respect to the surface on which the subject M1 stands. In this embodiment, if the angle formed between the optical axis 20A and a virtual line extending from a point P1 of the imaging device 20 that is parallel to the surface on which the subject M1 stands is formed vertically below the virtual line, the angle is set to -θ.

[0018] The actual distance, actual height, and actual inclination must be recognizable at least when the control unit 31 executes the correction described below. In this embodiment, the actual distance is α, the actual height is β, and the actual tilt is −θ, and the actual distance, actual height, and actual tilt are stored in advance in a storage device 34, which will be described later. In this case, when the photographer uses the imaging device 20 to capture an image of the subject M1, the photographer needs to instruct the photographer to capture the image at the actual distance, actual height, and actual tilt stored in the storage device 34, but it may be difficult for the photographer to faithfully reproduce these values.

[0019] The actual distance, actual height, and actual tilt may be estimated from image data or obtained from the photographer, or if the imaging device 20 is capable of obtaining the actual distance, actual height, and actual tilt, the information processing device 30 may obtain the actual distance, actual height, and actual tilt directly from the imaging device 20.

[0020] The imaging device 20 generates image data. The imaging device 20 may use still image data obtained by imaging as image data. The imaging device 20 may use video data obtained by imaging as image data. In this embodiment, the image data is video data. The imaging device 20 may be, for example, a monocular camera, a stereo camera, or a TOF (Time of Flight) camera as long as it is able to generate image data.

[0021] The information processing device 30 includes a control unit 31. The control unit 31 includes a processor 32 and a storage unit 33. The processor 32 is, for example, a central processing unit (CPU), a graphics processing unit (GPU), or a digital signal processor (DSP). The storage unit 33 includes a random access memory (RAM) and a read-only memory (ROM). The storage unit 33 stores program code or instructions configured to cause the processor 32 to execute processing. The storage unit 33, i.e., a computer-readable medium, includes any available medium accessible by a general-purpose or special-purpose computer. The control unit 31 may be configured by a hardware circuit such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). The control unit 31, which is a processing circuit, may include one or more processors operating according to a computer program, one or more hardware circuits such as an ASIC or FPGA, or a combination thereof.

[0022] The information processing device 30 includes a storage device 34. The storage device 34 stores information that can be read by the control unit 31. The storage device 34 is, for example, a hard disk drive or a solid state drive. The storage device 34 stores a skeleton estimation model 40 as a trained model.

[0023] As shown in FIG. 3, a skeleton estimation model 40 receives image data IM and outputs a three-dimensional skeleton model M2 of a subject M1 depicted in the image data IM. The three-dimensional skeleton model M2 includes skeleton points J1 of the subject M1 and linear links L1 connecting the skeleton points J1 to each other. The skeleton estimation model 40 is created, for example, from given training data based on an existing algorithm. The training data is, for example, a collection of three-dimensional coordinates of multiple skeleton points J1 of the training subject MI1 shown in FIG. 4 (hereinafter referred to as training three-dimensional coordinates) and a collection of training two-dimensional coordinates derived from the training three-dimensional coordinates. The algorithm is, for example, Open Pose.

[0024] It is assumed that the storage device 34 stores the distance between the virtual camera 21 and the training subject MI1 (hereinafter referred to as the reference distance) assumed from the training data, the installation height of the virtual camera 21 (hereinafter referred to as the reference installation height), and the inclination of the virtual camera 21 relative to the training subject MI1 (hereinafter referred to as the reference inclination).

[0025] The skeleton estimation model 40 includes a two-dimensional skeleton model generation unit 41 that derives the two-dimensional coordinates of each skeleton point J1 of the subject M1 appearing in the image data IM and generates a two-dimensional skeleton model by connecting the two-dimensional coordinates of each skeleton point J1 with links L1, and a three-dimensional skeleton model generation unit 42 that generates a three-dimensional skeleton model M2 by converting the two-dimensional coordinates of each skeleton point J1 in the two-dimensional skeletal model into three-dimensional coordinates.

[0026] Here, the reference distance, reference installation height, and reference tilt will be described using FIG. 4 . The reference distance is, for example, the distance between a predetermined point P11 of the virtual camera 21 and a predetermined point P12 of the learning subject MI1. The predetermined point P11 of the virtual camera 21 can be set arbitrarily. The predetermined point P12 of the learning subject MI1 can be set arbitrarily. However, in order to accurately grasp the deviation between the reference distance and the actual distance, it is preferable that the point P11 of the virtual camera 21 corresponds to the point P1 of the imaging device 20. For example, if the point P1 of the imaging device 20 is set to the center position of the lens of the imaging device 20, it is preferable that the point P11 of the virtual camera 21 corresponds to the center position of the lens of the virtual camera 21. Furthermore, it is preferable that the point P12 of the learning subject MI1 corresponds to the point P2 of the subject M1. For example, if the point P2 of the subject M1 is set to the waist of the subject M1, it is preferable that the point P12 of the learning subject MI1 corresponds to the waist of the learning subject MI1.

[0027] The reference distance in this embodiment is the Euclidean distance in the horizontal direction between a predetermined point P12 of the training subject MI1 and a predetermined point P11 of the virtual camera 21. In this embodiment, in a three-axis Cartesian coordinate system described later, the Z coordinate of the predetermined point P11 of the virtual camera 21 and the Z coordinate of the predetermined point P12 of the training subject MI1 are the same. In this case, the Euclidean distance does not include a Z-axis component, so the reference distance is the distance in the X-axis direction in the three-axis Cartesian coordinate system described later. The reference distance in this embodiment is α. Note that in this embodiment, the coordinate system based on the imaging device 20 and the coordinate system based on the virtual camera 21 are the same.

[0028] When the plane directly below the virtual camera 21 (hereinafter referred to as the first virtual plane SI1) and the plane on which the learning subject MI1 is standing (hereinafter referred to as the second virtual plane SI2) are at the same height, the reference installation height is the height from the first virtual plane SI1 or the second virtual plane SI2 to a point P11 of the virtual camera 21. When the first virtual plane SI1 and the second virtual plane SI2 are at different heights, the height from the second virtual plane SI2 to a point P11 of the virtual camera 21 is the installation height of the virtual camera 21. In other words, the reference installation height can be said to be the height from the second virtual plane SI2 to a point P11 of the virtual camera 21. The first virtual plane SI1 and the second virtual plane SI2 are horizontal and are the same plane. In a three-axis Cartesian coordinate system described later, the reference installation height is the difference between the Y coordinate of the point P11 of the virtual camera 21 and the Y coordinate of the second virtual plane SI2. In this embodiment, the reference installation height is β.

[0029] The reference tilt is, for example, the tilt of the optical axis 21A of the virtual camera 21 with respect to the second imaginary plane SI2. In this embodiment, the second imaginary plane SI2 and the optical axis 21A are parallel to each other. Therefore, the reference tilt in this embodiment is 0.

[0030] <Correction performed by the control unit> The correction performed by the control unit 31 after image data IM is input to the information processing device 30 will be described with reference to FIGS.

[0031] The control unit 31 of the present invention performs correction when there is at least one of a deviation between the actual height and the reference installation height, a deviation between the actual distance and the reference distance, and a deviation between the actual tilt and the reference tilt. In this embodiment, however, a case where there is a deviation between the actual tilt and the reference tilt will be described.

[0032] The control unit 31 executes steps S1 to S5 shown in FIG. 5 at a predetermined control cycle. In step S1, the control unit 31 inputs the image data IM to the two-dimensional skeletal model generation unit 41. As a result, the two-dimensional skeletal model generation unit 41 derives the two-dimensional coordinates of each skeletal point J1 and generates a two-dimensional skeletal model by connecting the two-dimensional coordinates of each skeletal point J1 with links L1. Each skeletal point J1 indicates the position of a joint such as the neck, shoulders, elbows, wrists, waist, hips, knees, and ankles. The two-dimensional coordinates are coordinates in the image data IM.

[0033] Next, in step S2, the control unit 31 inputs the two-dimensional skeletal model generated in step S1 to the three-dimensional skeletal model generation unit 42. As a result, the three-dimensional skeletal model generation unit 42 generates a three-dimensional skeletal model M2 by converting the two-dimensional coordinates of each skeletal point J1 in the two-dimensional skeletal model into three-dimensional coordinates. The skeletal points J1 are assigned ID codes, and the control unit 31 can recognize from the ID codes which joints each skeletal point J1 corresponds to. For example, the control unit 31 can recognize from the ID codes which skeletal points J1 correspond to the right elbow joint and which skeletal points J1 correspond to the left elbow joint.

[0034] As shown in FIGS. 1 and 2, the three-dimensional coordinate system is a triaxial Cartesian coordinate system in which the vertical direction is the Y-axis, the X-axis is an axis perpendicular to the Y-axis and extending from the image capture device 20 to the subject M1, and the Z-axis is an axis perpendicular to the X-axis and Y-axis. The three-dimensional coordinate system (x1, y1, z1) is represented by a Z-coordinate z1, which is the coordinate on the Z-axis of the triaxial Cartesian coordinate system, an X-coordinate x1, which is the coordinate on the X-axis of the triaxial Cartesian coordinate system, and a Y-coordinate y1, which is the coordinate on the Y-axis of the triaxial Cartesian coordinate system. The three-dimensional skeletal model generation unit 42 derives the three-dimensional coordinate of the skeleton point J1 relative to a reference position such as the waist. The origin of the triaxial Cartesian coordinate system can be set at any position. In this embodiment, the origin of the triaxial Cartesian coordinate system is a position where the Z-coordinate z1 and the X-coordinate x1 are the position of the waist, and the Y-coordinate y1 is the plane on which the subject M1 is standing.

[0035] Next, in step S3, the control unit 31 compares the actual height with the reference installation height, the actual distance with the reference distance, and the actual tilt with the reference tilt. In this embodiment, the actual height and the reference installation height are α, the actual distance and the reference distance are β, the actual tilt is −θ1, and the reference tilt is 0.

[0036] Therefore, the control unit 31 recognizes that only the actual tilt and the reference tilt are different. Next, in step S4, the control unit 31 corrects the three-dimensional coordinates of each skeleton point J1 in the three-dimensional skeletal model M2 generated in step S3 based on the deviation between the tilt of the image capture device 20 with respect to the subject M1 and the reference tilt. Because the three-dimensional skeletal model generation unit 42 generates the three-dimensional skeletal model M2 based on the reference tilt, if there is a deviation between the actual tilt and the reference tilt, this deviation will cause an error in the three-dimensional coordinates of each skeleton point J1 in the three-dimensional skeletal model M2. The correction performed in step S4 is intended to reduce this error. The deviation between the actual tilt and the reference tilt will be referred to as an angle deviation, as appropriate.

[0037] When an angular misalignment occurs, an error occurs in the Y coordinate y1 of the three-dimensional coordinates. The control unit 31 corrects the Y coordinate y1 based on a correction value calculated using the actual distance β and the amount of angular misalignment, and corrects the three-dimensional skeletal model. This will be explained in detail below.

[0038] The control unit 31 sets the value obtained by β×tan(radθ) as a correction value, and corrects the Y coordinate y1 of the three-dimensional coordinates of each skeleton point J1 in the three-dimensional skeleton model M2 generated in step S2 based on the correction value. radθ is the amount of angular deviation. In other words, radθ is the angle between the optical axis 21A of the virtual camera 21 and the optical axis 20A of the image capture device 20. In this embodiment, the reference tilt is 0 and the actual tilt is -θ, so radθ=θ.

[0039] If the reference tilt is greater than the actual tilt, the Y coordinate y1 is corrected by adding a correction value to the Y coordinate y1, and if the reference tilt is smaller than the actual tilt, the Y coordinate y1 is corrected by subtracting a correction value from the Y coordinate y1.

[0040] Next, in step S5, the control unit 31 evaluates the posture of the subject M1 based on the three-dimensional skeletal model corrected in step S4. In this embodiment, the control unit 31 performs the evaluation by assigning a score to the posture of the subject M1. Because the Y coordinate y1 was corrected in step S4, the control unit 31 assigns a score to the posture of the subject M1 based on the corrected three-dimensional coordinates (x1, y1, z1). For example, the score increases as the physical load increases.

[0041] The control unit 31 calculates, for example, the waist angle, right knee angle, and left knee angle from the three-dimensional coordinates of each skeleton point J1. The waist angle is, for example, the angle between the thigh and the waist. The right knee angle and left knee angle are the angles between the left and right thighs and lower legs, respectively.

[0042] Scoring of the posture of subject M1 can be performed by any method. For example, a correspondence relationship in which scores are assigned to combinations of waist angle, right knee angle, and left knee angle may be determined in advance, and the posture of subject M1 may be scored based on this correspondence. A calculation formula may be determined in advance, such as increasing the score as the waist angle increases, and the posture of subject M1 may be scored based on this calculation formula. The posture of subject M1 may be scored using a trained model that scores the posture of subject M1.

[0043] The posture score of subject M1 is displayed, for example, on a display unit. This allows a manager or the like to recognize the physical load on subject M1. The display format on the display unit is arbitrary, but for example, the posture score may be displayed in a graph format that links the posture score with the passage of time.

[0044] [Effects of this embodiment] (1) The control unit 31 corrects the three-dimensional coordinates based on the deviation between the actual tilt and the reference tilt. Even if an error occurs in the three-dimensional coordinates of each skeleton point J1 in the three-dimensional skeleton model M2 output by the three-dimensional skeleton model generation unit 42 due to an angle deviation, this error can be reduced.

[0045] It is also possible to reconstruct the three-dimensional skeletal model generation unit 42 according to the angular misalignment deviation amount radθ each time the actual tilt is changed. In this case, collecting and annotating training data is required to reconstruct the three-dimensional skeletal model generation unit 42, which increases the effort and cost. It is also possible to reconstruct a trained model that corrects errors in the three-dimensional coordinates of each skeleton point J1, but similar problems may arise even in this case. In contrast, when the three-dimensional coordinates are geometrically corrected based on the angular misalignment deviation amount radθ as in the embodiment, it is not necessary to reconstruct the three-dimensional skeletal model generation unit 42 or a trained model for correcting errors. Therefore, it is possible to reduce errors in the three-dimensional coordinates of each skeleton point J1 without increasing the effort and cost.

[0046] (2) The control unit 31 uses the value obtained by β×tan(radθ) as a correction value, and corrects the Y coordinate y1 of the three-dimensional coordinates in the three-dimensional skeletal model M2 based on this correction value. This makes it possible to reduce errors in the three-dimensional coordinates caused by angular misalignment. When angular misalignment occurs, errors in the three-dimensional coordinates tend to become large. Therefore, by reducing errors in the three-dimensional coordinates caused by angular misalignment, it is possible to reduce errors more appropriately.

[0047] (3) The skeleton estimation device 10 scores the posture of the subject M1 based on the three-dimensional coordinates of each corrected skeleton point J1, thereby making it possible to grasp the physical load on the subject M1. [Example of change] The embodiment can be modified as follows: The embodiment and the following modifications can be combined with each other to the extent that they are not technically inconsistent.

[0048] In the above embodiment, correction is performed using the calculated correction value. However, this is not limiting. The control unit 31 may correct the three-dimensional coordinates by multiplying the three-dimensional coordinates in the three-dimensional skeleton model generated by the three-dimensional skeleton model generation unit 42 by coefficients associated with the height deviation amount (hereinafter referred to as the height deviation amount), which is the deviation between the actual height and the reference installation height, the distance deviation amount (hereinafter referred to as the distance deviation amount), which is the deviation between the actual distance and the reference distance, and the angle deviation amount. The coefficients may be multiplied by individual values for each skeleton point J1. The coefficients may be set in advance according to the height deviation amount, the distance deviation amount, and the angle deviation amount radθ. The coefficients may be derived according to the height deviation amount, the distance deviation amount, and the angle deviation amount radθ. In this case, the coefficients are derived using, for example, a map or a calculation formula.

[0049] In step S4, the control unit 31 may correct the three-dimensional coordinates in the three-dimensional skeletal model based on two or more deviations among the deviation between the actual height and the reference installation height, the deviation between the actual distance and the reference distance, and the deviation between the actual inclination and the reference inclination.

[0050] In the above embodiment, the three-dimensional coordinates of the three-dimensional skeletal model generated by the three-dimensional skeletal model generation unit 42 are corrected. However, this is not limiting. The control unit 31 may also correct the trained model. For example, if there is at least one of a deviation between the actual height and the reference installation height, a deviation between the actual distance and the reference distance, and a deviation between the actual tilt and the reference tilt, the trained model may be corrected by multiplying the function used by the three-dimensional skeletal model generation unit 42 to convert two-dimensional coordinates into three-dimensional coordinates by a coefficient corresponding to the amount of deviation in height, distance, or angle. This causes the three-dimensional skeletal model generation unit 42 to output three-dimensional coordinates with minimal error. Therefore, the error in the three-dimensional coordinates of the skeleton point J1 can be reduced.

[0051] The skeleton estimation device 10 may be used to obtain the three-dimensional coordinates of each skeleton point J1 of a subject M1 other than a worker. The bone structure estimation device 10 may be used to analyze the movement of the subject M1, and may not necessarily be used to score the movement.

[0052] The control unit 31 may evaluate the posture of the subject M1 by a method other than a score, such as classifying the posture into types such as uncomfortable posture, slightly uncomfortable posture, and comfortable posture, and then evaluating the posture. The position of the image capture device 20 may be changeable. In this case, at least one of the distance deviation and the height deviation changes in response to the change in the position of the image capture device 20. The control unit 31 may correct the three-dimensional coordinates or the trained model in response to the changed deviation. In this case, the skeleton estimation device 10 may be provided with a sensor that detects the amount of change in the position of the image capture device 20.

[0053] The inclination of the image capture device 20 with respect to the subject M1 may be changeable. In this case, the amount of angular misalignment radθ changes in response to a change in the inclination of the image capture device 20 with respect to the subject M1. The control unit 31 may correct the three-dimensional coordinates or the trained model in response to the changed amount of misalignment radθ. In this case, the skeleton estimation device 10 may be provided with a sensor that detects the amount of change in the inclination of the image capture device 20 with respect to the subject M1.

[0054] The coordinate system based on the image capture device 20 and the coordinate system based on the virtual camera 21 may be different coordinate systems. In this case, the coordinates in each coordinate system can be converted into each other based on the deviation between the origins of the respective coordinate systems and the deviation between the coordinate axes of the respective coordinate systems.

[0055] The reference inclination can be set as appropriate, for example, to the angle between the horizontal line extending from the point P11 and the optical axis 21A. [Explanation of symbols]

[0056] IM...image data, J1...skeleton points, M1...subject, MI1...learning subject, 10...skeleton estimation device, 20...imaging device, 21...virtual camera, 31...control unit, 34...storage device.

Claims

1. an imaging device that captures an image of a subject and generates image data; a storage device that derives two-dimensional coordinates of skeleton points of the subject captured in the image data and stores a trained model that converts the derived two-dimensional coordinates into three-dimensional coordinates; a control unit that acquires the three-dimensional coordinates converted by the trained model, and corrects the three-dimensional coordinates based on at least one of a difference between an installation height of the imaging device and a reference installation height, a difference between a distance between the imaging device and the subject and a reference distance, and a difference between an inclination of the imaging device with respect to the subject and a reference inclination, a reference installation height that is an installation height of a virtual camera assumed from training data used when generating the trained model, a reference distance that is a distance between a training subject of the training data and the virtual camera, and a reference tilt that is an inclination of the virtual camera relative to the training subject.

2. an imaging device that captures an image of a subject and generates image data; a storage device that derives two-dimensional coordinates of skeleton points of the subject captured in the image data and stores a trained model that converts the derived two-dimensional coordinates into three-dimensional coordinates; calculating a coefficient for correcting the trained model based on at least one of a difference between an installation height of the imaging device and a reference installation height, a difference between a distance between the imaging device and the subject and a reference distance, and a difference between an inclination of the imaging device with respect to the subject and a reference inclination; a control unit that corrects the trained model using the calculated coefficients; a reference installation height that is an installation height of a virtual camera assumed from training data used when generating the trained model, a reference distance that is a distance between a training subject of the training data and the virtual camera, and a reference tilt that is an inclination of the virtual camera relative to the training subject.

3. 3. The skeleton estimation device according to claim 1, wherein the three-dimensional coordinates are coordinates in a three-axis Cartesian coordinate system in which the vertical direction is the Y axis, the axis perpendicular to the Y axis and pointing from the imaging device to the subject is the X axis, and the axis perpendicular to the X axis and the Y axis is the Z axis.

4. The skeletal structure estimation device according to claim 1 or 2, wherein the control unit evaluates a posture of the subject based on the three-dimensional coordinates.

Citation Information

Patent Citations

  • Work estimation apparatus, method and program

    JP2022048017A