Skeleton estimation device and skeleton estimation program

JP2026142581APending Publication Date: 2026-09-08MITSUBISHI ELECTRIC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025029617
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2026-09-08

AI Technical Summary

Benefits of technology

【0007】 本開示によれば、機械学習を用いて、カメラ画像に基づき、カメラ画像上の人の2次元骨格点から人の3次元骨格点を推定する技術において、3次元骨格点の推定精度が低減することを防ぐことができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026142581000001_ABST
    Figure 2026142581000001_ABST
Patent Text Reader

Abstract

The present invention provides a skeleton estimation device that prevents a reduction in the estimation accuracy of three-dimensional skeleton points. [Solution] The system includes an image acquisition unit (11) that acquires a camera image of a subject captured by a camera (2), a corrected two-dimensional skeleton estimation unit (12) that estimates the two-dimensional skeleton points of the subject in the camera image, corrected based on the environmental parameters of the camera (2), from the camera image acquired by the image acquisition unit (11), and a three-dimensional skeleton estimation unit (13) that estimates the three-dimensional skeleton points of the subject in real space using a machine learning model for three-dimensional skeleton estimation, based on the two-dimensional skeleton points of the subject estimated by the corrected two-dimensional skeleton estimation unit (12).
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] The present disclosure relates to a skeleton estimation device and a skeleton estimation program. [[Background Art]]

[0002] Conventionally, there has been known a technology that uses machine learning such as a neural network to estimate a person's skeletal points (hereinafter referred to as "two-dimensional skeletal points") in a captured image obtained by capturing an image of a person with a camera (hereinafter referred to as "camera image"), and then estimates the person's skeletal points in a three-dimensional real space (hereinafter referred to as "three-dimensional skeletal points") from the estimated two-dimensional skeletal points (see, for example, Patent Document 1). The two-dimensional skeletal points and three-dimensional skeletal points are represented by coordinates on the camera image (hereinafter referred to as "two-dimensional skeletal point coordinates") and coordinates in the real space (hereinafter referred to as "three-dimensional skeletal point coordinates"), respectively. [[Prior Art Literature]] [[Patent Literature]]

[0003] [[Patent Document 1]] Japanese Unexamined Patent Publication No. 2022-49521 [[Summary of the Invention]] [[Problem to be Solved by the Invention]]

[0004] A camera image is obtained by converting information in a three-dimensional real space into two dimensions using various environmental parameters. Here, the environmental parameters refer to parameters representing elements related to the shooting environment of the camera that affect the geometric characteristics of the camera image. The environmental parameters include, for example, parameters related to camera distortion, or parameters related to the positional relationship between the camera and a person who is the subject. Therefore, the 2D skeleton points estimated from camera images are affected by environmental parameters, such as camera distortion or the loss of actual depth information. As a result, when estimating 3D skeleton points from 2D skeleton points, it becomes difficult to accurately reconstruct the 3D skeleton points from the 2D skeleton points on the camera image, which poses a challenge as it may reduce the accuracy of 3D skeleton point estimation based on camera images. In the technology disclosed in Patent Document 1, a computer converts 2D skeletal coordinates obtained by skeletal detection on a 2D image into 3D skeletal coordinates. It then refers to influence data, which associates the degree of influence each joint has on the coordinate conversion error from 2D to 3D for each tilt class, which is divided according to the tilt of the body's axes. The computer calculates an estimated error from the influence of each joint corresponding to the tilt class to which the 3D skeletal coordinate belongs, and from the 3D skeletal coordinate. The computer then performs a process to exclude 3D skeletal coordinates whose estimated error value is above a predetermined threshold. However, such technology does not take into account the fundamental problem that 2D skeletal coordinates are affected by environmental parameters, and therefore the above problem remains unresolved.

[0005] This disclosure was made to solve the above-mentioned problems, and aims to provide a skeleton estimation device that prevents a reduction in the estimation accuracy of 3D skeleton points in a technology that uses machine learning to estimate a person's 3D skeleton points from 2D skeleton points on a camera image. [Means for solving the problem]

[0006] The skeleton estimation device according to this disclosure comprises: an image acquisition unit that acquires camera images of a subject captured by a camera; a corrected two-dimensional skeleton estimation unit that estimates the two-dimensional skeleton points of the subject in the camera image, corrected based on the camera's environmental parameters, from the camera image acquired by the image acquisition unit; and a three-dimensional skeleton estimation unit that estimates the three-dimensional skeleton points of the subject in real space using a machine learning model for three-dimensional skeleton estimation, based on the two-dimensional skeleton points of the subject estimated by the corrected two-dimensional skeleton estimation unit. [Effects of the Invention]

[0007] According to this disclosure, in a technique that uses machine learning to estimate a person's 3D skeletal points from 2D skeletal points on a camera image, it is possible to prevent a decrease in the estimation accuracy of the 3D skeletal points. [Brief explanation of the drawing]

[0008] [Figure 1] This figure shows an example of the configuration of the skeleton estimation device according to Embodiment 1. [Figure 2] This figure shows a detailed configuration example of the corrected 2D skeleton estimation unit of the skeleton estimation device according to Embodiment 1. [Figure 3] Figures 3A, 3B, and 3C illustrate an example of the effect of strain parameters on two-dimensional skeletal points estimated from camera images. [Figure 4] This is a flowchart illustrating the operation of the skeletal estimation device according to Embodiment 1. [Figure 5] This is a flowchart illustrating the detailed operation of the corrected 2D skeleton estimation unit in step ST2 of Figure 4. [Figure 6] Figures 6A, 6B, and 6C illustrate an example of the influence of positional parameters on two-dimensional skeletal points estimated from camera images. [Figure 7] Figures 7A and 7B show an example of the hardware configuration of the skeleton estimation device according to Embodiment 1. [Modes for carrying out the invention]

[0009] The embodiments of this disclosure will be described in detail below with reference to the drawings. Embodiment 1. The skeleton estimation device 1 according to Embodiment 1 (see Figure 1 below) estimates the 2D skeleton points of a person (hereinafter referred to as "the subject") captured by camera 2 (see Figure 1 below) from the camera image, corrected based on the environmental parameters of camera 2, and then estimates the 3D skeleton points of the subject in 3D real space using a pre-trained machine learning model (hereinafter referred to as "machine learning model") based on the estimated 2D skeleton points of the subject. In the following description, when "real space" is used, it is assumed that the real space is a 3D real space. In detail, the skeleton estimation device 1 uses a machine learning model (hereinafter referred to as the "2D skeleton estimation machine learning model") to estimate the 2D skeleton points of the subject on the camera image captured by the camera, and corrects the estimated 2D skeleton points. Then, the skeleton estimation device 1 uses a machine learning model (hereinafter referred to as the "3D skeleton estimation machine learning model") to estimate the 3D skeleton points of the subject from the corrected 2D skeleton points. The skeleton estimation device 1 corrects the 2D skeleton points of the subject estimated from the camera image based on environmental parameters, thereby removing the influence of the environmental parameters from the corrected 2D skeleton points. The skeleton estimation device 1 prevents a reduction in the estimation accuracy of the 3D skeleton points by estimating the subject's 3D skeleton points using a machine learning model for 3D skeleton estimation, based on the 2D skeleton points corrected to remove the influence of the environmental parameters.

[0010] In Embodiment 1, the machine learning model for 2D skeleton estimation is a machine learning model that takes a camera image as input and outputs information indicating the 2D skeleton points of a person on the camera image. The machine learning model for 2D skeleton estimation may be a machine learning model that has been trained using camera images that are affected by the environmental parameters of camera 2 as training data, without considering the influence of the environmental parameters of camera 2. On the other hand, in Embodiment 1, the machine learning model for 3D skeleton estimation is a machine learning model that takes information indicating 2D skeleton points as input and outputs information indicating 3D skeleton points corresponding to those 2D skeleton points. The machine learning model for 3D skeleton estimation is a machine learning model that has been trained using information indicating 2D skeleton points that are not affected by the environmental parameters of camera 2, or more specifically, 2D skeleton points estimated from camera images that are not affected by the environmental parameters of camera 2, as training data. In Embodiment 1, "not affected" is not limited to being completely unaffected, but also includes being unaffected even if there is an influence within an acceptable range that does not require consideration. For example, a machine learning model for 3D skeleton estimation may be a machine learning model trained using information that shows 2D skeleton points after correction of 2D skeleton points estimated using a machine learning model for 2D skeleton estimation based on camera images; in other words, information that shows 2D skeleton points corrected based on environmental parameters.

[0011] The 2D and 3D machine learning models for skeletal estimation are pre-generated and stored in a location accessible to the skeletal estimation device 1.

[0012] In Embodiment 1, the skeletal points are characteristic points corresponding to the locations of parts of the human body. The specific points on the human body that will be designated as skeletal points corresponding to the locations of those body parts are predetermined. The skeletal estimation device 1 can estimate multiple two-dimensional skeletal points on a camera image, such as skeletal points corresponding to the position of the head and skeletal points corresponding to the position of the neck, and can estimate multiple three-dimensional skeletal points in real space based on the estimated two-dimensional skeletal points.

[0013] As mentioned above, the camera image is a 2D representation of real-space information using the environmental parameters of camera 2. Therefore, the 2D skeletal points estimated from the camera image are affected by the environmental parameters, making it difficult to accurately reconstruct the 3D skeletal points from the 2D skeletal points on the camera image. This may reduce the accuracy of the 3D skeletal point estimation based on the camera image. In contrast, the skeleton estimation device 1 according to Embodiment 1 is capable of estimating 2D skeleton points that have been corrected based on the environmental parameters of the camera 2, or more specifically, corrected to eliminate the influence of those environmental parameters, and then estimates 3D skeleton points from the corrected 2D skeleton points using a machine learning model for 3D skeleton estimation. As a result, the skeleton estimation device 1 eliminates the influence of environmental parameters and improves the accuracy of estimating the 3D skeleton points of the subject. In Embodiment 1, the environmental parameters of camera 2 refer to parameters that represent elements related to the shooting environment of camera 2 that affect the geometric characteristics of the camera image, as described above. The geometric characteristics of the camera image refer to the spatial transformation characteristics when an object in three-dimensional space is projected onto a two-dimensional image plane. Environmental parameters include, for example, parameters related to camera distortion (hereinafter referred to as "distortion parameters") or parameters related to the positional relationship between camera 2 and the subject (a person) (hereinafter referred to as "position parameters").

[0014] In the following Embodiment 1, as an example, the subject is assumed to be the driver of a vehicle. The skeleton estimation device 1 uses a two-dimensional skeleton estimation machine learning model to estimate the corrected two-dimensional skeleton points of the driver from camera images captured by a camera 2, which is provided to capture at least the area in the vehicle interior where the driver may be present. Based on the estimated two-dimensional skeleton points of the driver, the device uses a three-dimensional skeleton estimation machine learning model to estimate the three-dimensional skeleton points of the driver. Furthermore, in the following Embodiment 1, as an example, the environmental parameter is assumed to be a distortion parameter, and the skeleton estimation apparatus 1 will be described with an example in which the skeleton estimation apparatus 1 estimates the three-dimensional skeleton points of a driver by removing the influence of the distortion parameter.

[0015] FIG. 1 is a diagram showing a configuration example of the skeleton estimation apparatus 1 according to Embodiment 1. The skeleton estimation apparatus 1 is mounted on, for example, a vehicle (not shown). The skeleton estimation apparatus 1 is connected to a camera 2 via a network, and the skeleton estimation apparatus 1 and the camera 2 constitute a skeleton estimation system 100.

[0016] The camera 2 is provided so as to be capable of capturing an image of at least an area in the vehicle compartment where a driver can exist. Here, as an example, the camera 2 is provided, for example, at a central portion in the vehicle width direction of a dashboard (not shown), and is provided so as to be capable of simultaneously capturing an image of the driver and an occupant in the passenger seat (hereinafter referred to as a "passenger seat occupant") from the front in the traveling direction of the vehicle relative to the driver seat and the passenger seat. In Embodiment 1, the term "central" is not limited to strictly the center, but includes substantially the center. This is merely an example, and the camera 2 only needs to be provided so as to be capable of capturing an image of the driver who is the subject. For example, the camera 2 may be shared with a so-called "Driver Monitoring System (DMS)". The camera 2 outputs captured camera images to the skeleton estimation apparatus 1.

[0017] As shown in FIG. 1, the skeleton estimation apparatus 1 includes an image acquisition unit 11, a corrected two-dimensional skeleton estimation unit 12, a three-dimensional skeleton estimation unit 13, a three-dimensional skeleton information output unit 14, a first storage unit 15, and a second storage unit 16.

[0018] The image acquisition unit 11 acquires a camera image obtained by capturing an image of a subject, here the driver, by the camera 2. The image acquisition unit 11 outputs the acquired camera image to the corrected two-dimensional skeleton estimation unit 12.

[0019] The corrected 2D skeleton estimation unit 12 estimates the 2D skeleton points of the subject in the camera image, which have been corrected based on the environmental parameters of the camera 2, in this case the distortion parameters, from the camera image acquired by the image acquisition unit 11. Here, Figure 2 shows a detailed example of the configuration of the corrected two-dimensional skeleton estimation unit 12 of the skeleton estimation device 1 according to Embodiment 1. As shown in Figure 2, the corrected 2D skeleton estimation unit 12 includes a 2D skeleton estimation unit 121 and a 2D skeleton correction unit 122.

[0020] The 2D skeleton estimation unit 121 estimates the driver's 2D skeleton points using a 2D skeleton estimation machine learning model based on the camera image acquired by the image acquisition unit 11. Specifically, the 2D skeleton estimation unit 121 estimates the driver's 2D skeleton points by inputting the camera image into the 2D skeleton estimation machine learning model and obtaining information indicating the 2D skeleton points. More specifically, the information indicating the 2D skeleton points is information indicating the coordinates of the 2D skeleton points in the camera image. The information indicating the 2D skeleton points is, for example, information that associates the camera image, the coordinates of the 2D skeleton points on the camera image, and information indicating which part of the body (e.g., head, right shoulder, left shoulder, neck, right elbow, left elbow, right hip, left hip, etc.) the 2D skeleton point indicated by the 2D skeleton point coordinates corresponds to. The 2D skeleton points are represented, for example, by coordinates on the camera image. The 2D skeleton estimation unit 121 outputs information indicating the estimated 2D skeleton points of the driver (hereinafter referred to as "estimated 2D skeleton point information") to the 2D skeleton correction unit 122. For example, the 2D skeleton estimation unit 121 may use information indicating 2D skeleton points obtained from a 2D skeleton estimation machine learning model as the estimated 2D skeleton point information.

[0021] The 2D skeleton correction unit 122 corrects the 2D skeleton points estimated by the 2D skeleton estimation unit 121 based on the environmental parameters of the camera 2, in this case, the distortion parameters. The 2D skeleton correction unit 122 can correct the 2D skeleton points on the camera image estimated by the 2D skeleton estimation unit 121, or more specifically, the 2D skeleton point coordinates, using known distortion correction methods. For example, the 2D skeletal correction unit 122 corrects the 2D skeletal point coordinates on the camera image by transforming the coordinates of each pixel using distortion parameters in a so-called barrel correction method to restore the linearity of the camera image. For example, a mapping table may be generated in advance, based on the distortion parameter, indicating how much a certain pixel coordinate in the camera image should be corrected to in order to correct the distortion. The 2D skeleton correction unit 122 may then correct the 2D skeleton point coordinates based on this mapping table. As a result, the 2D skeleton correction unit 122 corrects the 2D skeleton points estimated by the 2D skeleton estimation unit 121, more specifically the coordinates of the 2D skeleton points, to 2D skeleton points that are not affected by the distortion parameter, in other words, 2D skeleton points from which the influence of the distortion parameter has been removed, more specifically the coordinates of the 2D skeleton points. That is, the 2D skeleton correction unit 122 corrects the 2D skeleton points estimated by the 2D skeleton estimation unit 121, more specifically the coordinates of the 2D skeleton points, to 2D skeleton points that would be estimated from an image not affected by the distortion parameter, more specifically the coordinates of the 2D skeleton points.

[0022] The distortion parameters are pre-set, for example, at the factory when the camera 2 is shipped, and are stored in a memory unit (not shown) located in a place accessible to the skeleton estimation device 1. The two-dimensional skeleton correction unit 122 simply needs to retrieve the distortion parameters from this memory unit.

[0023] Furthermore, the 2D skeletal correction unit 122 may calculate distortion parameters based on camera images, for example. For example, the 2D skeletal correction unit 122 can calculate distortion parameters based on the distortion in the camera image of the parameter calculation object that is captured in the camera image. Which object will be used as the parameter calculation object is determined in advance by the developers, etc. The developers, etc. decide that the parameter calculation object will be an object that is expected to be captured in the camera image and that is free of distortion. Here, a free of distortion object is assumed to be a linear structure that is not expected to deform. Examples of parameter calculation objects include the B-pillar, door frame, or the straight edge of the center console. For example, a table (hereinafter referred to as the "distortion parameter calculation table") is pre-defined by the developer or others, which defines how much distortion the object used for parameter calculation should be when it is captured in a camera image from a reference shape, and this table is stored in a memory unit (not shown). The reference shape refers to the shape of the object used for parameter calculation in a camera image when it is captured without distortion. The 2D skeletal correction unit 122 can, for example, use known image recognition technology to extract objects for parameter calculation captured in the camera image, and then calculate the distortion parameters by comparing the shape of the extracted objects for parameter calculation with a distortion parameter calculation table.

[0024] The corrected 2D skeleton estimation unit 12 outputs information indicating the 2D skeleton points of the driver after correction by the 2D skeleton correction unit 122 (hereinafter referred to as "corrected estimated 2D skeleton point information") to the 3D skeleton estimation unit 13. The corrected estimated two-dimensional skeletal point information is, in detail, information that associates, for example, the two-dimensional skeletal point coordinates (hereinafter referred to as "corrected two-dimensional skeletal point coordinates") of the corrected two-dimensional skeletal point (hereinafter referred to as "corrected two-dimensional skeletal point") with information indicating which part of the body the corrected two-dimensional skeletal point indicated by the corrected two-dimensional skeletal point coordinates corresponds to. The corrected estimated two-dimensional skeletal point information may include, for example, the captured image output from the image acquisition unit 11.

[0025] The 3D skeleton estimation unit 13 estimates the 3D skeleton points of the driver in real space using a machine learning model for 3D skeleton estimation, based on the 2D skeleton points of the driver estimated by the corrected 2D skeleton estimation unit 12, or more specifically, the corrected 2D skeleton points of the driver corrected based on the strain parameters. In detail, the 3D skeleton estimation unit 13 inputs the corrected estimated 2D skeleton point information output from the corrected 2D skeleton estimation unit 12 into a machine learning model for 3D skeleton estimation, and estimates the driver's 3D skeleton points by obtaining information indicating the 3D skeleton points corresponding to the corrected 2D skeleton points indicated in the corrected estimated 2D skeleton point information. More specifically, the information indicating the 3D skeleton points is information indicating the coordinates of the 3D skeleton points in real space. The information indicating the 3D skeleton points is, for example, information that associates the coordinates of the 3D skeleton points with information indicating which part of the body the 3D skeleton point indicated by the coordinates corresponds to. For example, in the information indicating the 3D skeleton points, distortion-corrected camera images included in the corrected estimated 2D skeleton point information may also be associated with this information. The 3D skeleton points are represented, for example, by coordinates in real space. The 3D skeleton estimation unit 13 outputs information indicating the estimated 3D skeleton points of the driver (hereinafter referred to as "estimated 3D skeleton point information") to the 3D skeleton information output unit 14.

[0026] The 3D skeleton information output unit 14 outputs the estimated 3D skeleton point information output from the 3D skeleton estimation unit 13 to a device (not shown) that executes various applications using the estimated 3D skeleton point information. For example, the 3D skeletal information output unit 14 outputs estimated 3D skeletal point information to an action detection device (not shown) that detects the actions of the vehicle's occupants. Based on the estimated 3D skeletal point information, the action detection device detects the actions of the occupant, in this case the driver, such as whether the driver is operating the navigation device (not shown) or reaching for the drink holder (not shown). Since the technology for detecting a person's actions from their 3D skeletal points is a well-known technology, a detailed explanation is omitted. The action detection device then performs control according to the detected driver's actions. An example of control according to the driver's actions is the control of turning on the interior lights. Furthermore, for example, the 3D skeletal information output unit 14 may output estimated 3D skeletal point information to a physique detection device (not shown) that detects the physique of the vehicle occupant. The physique detection device detects the occupant's, in this case the driver's, physique based on the estimated 3D skeletal point information. Since the technology for detecting a person's physique from their 3D skeletal points is a known technology, a detailed explanation is omitted. The physique detection device performs control according to the detected physique of the driver. Examples of control according to the driver's physique include airbag control. Furthermore, for example, the 3D skeletal information output unit 14 may output estimated 3D skeletal point information to a posture detection device (not shown) that detects the posture of the vehicle occupant. The posture detection device detects the posture of the occupant, in this case the driver, based on the estimated 3D skeletal point information. Since the technology for detecting a person's posture from their 3D skeletal points is a known technology, a detailed explanation is omitted. The posture detection device performs control according to the detected posture of the driver. Examples of control according to the driver's posture include alarm output control when posture collapse is detected, or automatic driving control.

[0027] The function of the 3D skeleton information output unit 14 may also be provided by the 3D skeleton estimation unit 13. In that case, the skeleton estimation device 1 is not required to include the 3D skeleton information output unit 14.

[0028] The first memory unit 15 stores a machine learning model for two-dimensional skeleton estimation. In Figure 1, the first storage unit 15 is shown to be located within the skeleton estimation device 1. However, the first storage unit 15 may be located outside the skeleton estimation device 1, in a location accessible to the skeleton estimation device 1.

[0029] The second memory unit 16 stores a machine learning model for 3D skeleton estimation. In Figure 1, the second storage unit 16 is shown to be located within the skeleton estimation device 1. However, the second storage unit 16 may be located outside the skeleton estimation device 1, in a location accessible to the skeleton estimation device 1.

[0030] Alternatively, for example, the first memory unit 15 and the second memory unit 16 may be the same memory unit, and this memory unit may store a machine learning model for 2D skeleton estimation and a machine learning model for 3D skeleton estimation.

[0031] Here, we will explain, with the help of a diagram, the influence of the distortion parameter on the two-dimensional skeletal points estimated from the camera image, and the effect of the skeletal estimation device 1 correcting the two-dimensional skeletal points, using an example. Figures 3A, 3B, and 3C illustrate an example of the effect of strain parameters on two-dimensional skeletal points estimated from camera images. Figure 3A shows a camera image of a driver seated in the reference position facing camera 2, with no distortion in camera 2. Figure 3B shows a camera image of a driver rotated by camera 2, from a position facing camera 2 in the reference position, with a yaw angle of +45 degrees around the central axis of the body passing through the skeletal point corresponding to the neck. In Embodiment 1, "reference position" refers to a standard position set in advance by the developer, etc. The driver captured in the camera image shown in Figure 3A is seated in the driver's seat at the standard position, which is the reference position, inside the vehicle. The yaw angle of the driver captured in the camera image shown in Figure 3A is 0 degrees. Here, the camera image is an image of the driver and passenger seat taken by camera 2 from the center of the dashboard in the vehicle width direction. For simplicity of explanation, only the left half of the camera image of the driver is shown. For the purposes of this example, the vehicle is assumed to be a right-hand drive vehicle, and in Figures 3A and 3B, the center of the camera image is indicated by "O". Also, here, the yaw angle to the right is represented as positive, and the yaw angle to the left is represented as negative, with respect to the body's central axis.

[0032] On the other hand, Figure 3C shows a camera image taken by camera 2, which has distortion, of a driver seated in a reference position facing camera 2. In Figure 3C, as with Figures 3A and 3B, for the sake of simplicity of explanation, only the left half of the camera image in which the driver was captured is shown. For example, the vehicle is a right-hand drive vehicle, and the center of the camera image is indicated by "O". In other words, the difference between the camera image shown in Figure 3A and the camera image shown in Figure 3C is whether or not camera 2, which captured the camera image, has distortion. Figures 3A, 3B, and 3C illustrate the two-dimensional skeletal points of the driver estimated based on the camera images, respectively. Here, it is assumed that the two-dimensional skeletal points corresponding to the positions of the driver's nose, neck, right shoulder, left shoulder, right hip, and left hip were estimated based on the camera images.

[0033] For example, as shown in Figures 3A and 3B, when comparing the two-dimensional skeletal points of a driver estimated from a camera image of a driver seated at a reference position facing camera 2 with a distortion-free camera 2 (referred to as "facing two-dimensional skeletal points"; see Figure 3A) with the two-dimensional skeletal points of a driver estimated from a camera image of a driver rotated by a yaw angle of +45 degrees relative to camera 2 with a distortion-free camera 2 (referred to as "rotated facing two-dimensional skeletal points"; see Figure 3B), the distance between skeletal points is compressed vertically as the rotating facing two-dimensional skeletal points are further away from camera 2, in other words, as they are further from camera 2, i.e., as they are further from the center of the camera image. On the other hand, as shown in Figures 3B and 3C, for example, when comparing the rotated, oriented 2D skeleton points with the 2D skeleton points of the driver estimated from the camera image captured by the distorted camera 2 with the driver facing camera 2 (referred to as "distorted 2D skeleton points"; see Figure 3C), the distorted 2D skeleton points, like the rotated, oriented 2D skeleton points, are imaged with the distance between skeleton points compressed vertically as they are further away from camera 2. In other words, if it is unknown whether or not there is distortion in the camera image, or to what extent the distortion parameter of camera 2 is, the yaw angle of the driver being captured by the camera image cannot be determined. As a result, a problem arises in which the estimation accuracy of the 3D skeleton points estimated from the 2D skeleton points estimated based on the camera image, which is affected by the distortion parameter, may be reduced.

[0034] In contrast, the skeleton estimation device 1, as described above, has a corrected 2D skeleton estimation unit 12 that estimates the 2D skeleton points of the driver in the camera image, corrected based on distortion parameters, and a 3D skeleton estimation unit 13 that estimates the 3D skeleton points of the driver in real space using a machine learning model for 3D skeleton estimation, based on the 2D skeleton points estimated by the corrected 2D skeleton estimation unit 12. More specifically, the 2D skeleton estimation unit 121 estimates the 2D skeleton points of the driver based on the camera image using a 2D skeleton estimation machine learning model, the 2D skeleton correction unit 122 corrects the 2D skeleton points estimated by the 2D skeleton estimation unit 121 based on the distortion parameters of the camera 2, and the 3D skeleton estimation unit 13 estimates the 3D skeleton points of the driver in real space based on the 2D skeleton points corrected by the 2D skeleton correction unit 122 using a 3D skeleton estimation machine learning model. As a result, the skeleton estimation device 1 can solve the above problem and prevent a reduction in the estimation accuracy of 3D skeleton points.

[0035] For example, if the camera image is a distorted camera 2, as shown in Figure 3C, and the driver is seated at a reference position facing camera 2, then in the skeleton estimation device 1, the 2D skeleton estimation unit 121 estimates the driver's 2D skeleton points (at this point, these 2D skeleton points are distorted 2D skeleton points) from the camera image using a 2D skeleton estimation machine learning model. Then, the 2D skeleton correction unit 122 corrects the distorted 2D skeleton points based on the distortion parameters so that they become undistorted 2D skeleton points, as shown in Figure 3A, i.e., 2D skeleton points facing the driver seated at the reference position. More specifically, the 2D skeleton correction unit 122 corrects the coordinates of the distorted 2D skeleton points to match the coordinates of the 2D skeleton points facing the camera. As a result, the distorted 2D skeleton points are corrected based on the distortion parameters to remove the influence of the distortion parameters. The 2D skeleton points facing the camera are 2D skeleton points from which the influence of the distortion parameters has been removed. Then, the 3D skeleton estimation unit 13 estimates the 3D skeleton points of the driver using a machine learning model for 3D skeleton estimation, based on the 2D skeleton points facing the driver, from which the influence of the strain parameter has been removed. Therefore, the skeleton estimation device 1 can estimate the subject's 3D skeleton points while excluding the influence of environmental parameters, thereby preventing a decrease in the estimation accuracy of the 3D skeleton points.

[0036] Furthermore, in the skeleton estimation device 1, the 3D skeleton estimation unit 13 estimates the 3D skeleton points of the driver using a 3D skeleton estimation machine learning model that takes as input information indicating 2D skeleton points unaffected by strain parameters and outputs information indicating 3D skeleton points corresponding to those 2D skeleton points. In other words, the 3D skeleton estimation machine learning model is trained on information regarding 2D skeleton points unaffected by strain parameters. That is, the 3D skeleton estimation machine learning model is trained on training data (information regarding 2D skeleton points) that excludes the influence of strain parameters. Therefore, the skeleton estimation device 1 not only improves the estimation accuracy of 3D skeleton points as described above, but also enables the estimation of the 3D skeleton points of the driver even if the 3D skeleton estimation machine learning model is trained on a small amount of training data. Furthermore, the skeleton estimation device 1 can estimate the 3D skeleton points of a driver even if the 3D skeleton estimation machine learning model is one that has been trained using information about 2D skeleton points based on images that do not take the real environment into account, such as computer graphics (CG), as training data.

[0037] Furthermore, in the skeleton estimation device 1, the two-dimensional skeleton estimation machine learning model used by the two-dimensional skeleton estimation unit 121 when estimating the two-dimensional skeleton points of the driver may be a machine learning model trained on camera images affected by distortion parameters. For example, the machine learning model for 2D skeleton estimation may be a machine learning model trained on camera images collected from the Web with unknown distortion parameters. Even in that case, the 2D skeleton correction unit 122 corrects the 2D skeleton points of the driver estimated by the 2D skeleton estimation unit 121, and the 3D skeleton estimation unit 13 estimates the 3D skeleton points of the driver using a machine learning model for 3D skeleton estimation from the information indicating the corrected 2D skeleton points. Therefore, the skeleton estimation device 1 can prevent a reduction in the estimation accuracy of the driver's 3D skeleton points. To solve the problems described above, one might consider preparing camera images of the driver captured at all yaw angles for all distortion parameters as training data, and then training a 2D skeleton estimation machine learning model with this training data. However, such a solution is not practical because it is costly and in some cases inherently difficult to solve.

[0038] The operation of the skeleton estimation device 1 according to Embodiment 1 will be described. Figure 4 is a flowchart illustrating the operation of the skeleton estimation device 1 according to Embodiment 1.

[0039] The image acquisition unit 11 acquires camera images of the subject, in this case the driver, captured by the camera 2 (step ST1). The image acquisition unit 11 outputs the acquired camera image to the corrected 2D skeleton estimation unit 12.

[0040] The corrected 2D skeleton estimation unit 12 estimates the 2D skeleton points of the subject in the camera image, which have been corrected based on the environmental parameters of the camera 2, in this case the distortion parameters, from the camera image acquired by the image acquisition unit 11 in step ST1 (step ST2). The corrected 2D skeleton estimation unit 12 outputs the corrected estimated 2D skeleton point information to the 3D skeleton estimation unit 13.

[0041] The 3D skeleton estimation unit 13 estimates the 3D skeleton points of the driver in real space using a machine learning model for 3D skeleton estimation, based on the 2D skeleton points of the driver estimated by the corrected 2D skeleton estimation unit 12 in step ST2 (step ST3). More specifically, the 3D skeleton estimation unit 13 inputs the corrected estimated 2D skeleton point information output from the corrected 2D skeleton estimation unit 12 into a machine learning model for 3D skeleton estimation, and estimates the 3D skeleton points of the driver by obtaining information indicating the 3D skeleton points corresponding to the corrected 2D skeleton points shown in the corrected estimated 2D skeleton point information. The 3D skeleton estimation unit 13 outputs estimated 3D skeleton point information to the 3D skeleton information output unit 14.

[0042] The 3D skeleton information output unit 14 outputs the estimated 3D skeleton point information output from the 3D skeleton estimation unit 13 in step ST3 to a device that executes various applications using the estimated 3D skeleton point information (step ST4).

[0043] Figure 5 is a flowchart illustrating the detailed operation of the corrected 2D skeleton estimation unit 12 in the processing of step ST2 in Figure 4.

[0044] The 2D skeleton estimation unit 121 estimates the 2D skeleton points of the driver using a 2D skeleton estimation machine learning model based on the camera image acquired by the image acquisition unit 11 in step ST1 of Figure 4 (step ST11). Specifically, the 2D skeleton estimation unit 121 inputs the camera image into the 2D skeleton estimation machine learning model and obtains information indicating the 2D skeleton points to estimate the 2D skeleton points of the driver. The 2D skeleton estimation unit 121 outputs estimated 2D skeleton point information to the 2D skeleton correction unit 122.

[0045] The 2D skeleton correction unit 122 corrects the 2D skeleton points of the driver estimated by the 2D skeleton estimation unit 121 in step ST11 based on the environmental parameters of the camera 2, in this case, the distortion parameters (step ST12). The 2D skeleton correction unit 122 outputs the corrected estimated 2D skeleton point information to the corrected 2D skeleton estimation unit 12, and the corrected 2D skeleton estimation unit 12 outputs the corrected estimated 2D skeleton point information to the 3D skeleton estimation unit 13.

[0046] Thus, the skeleton estimation device 1 includes an image acquisition unit 11 that acquires camera images of a subject (in this case, a driver) captured by camera 2, and an image acquisition unit 11 that estimates the two-dimensional skeleton points of the subject (in this case, a driver) in the camera images, corrected based on the environmental parameters (in this case, distortion parameters) of camera 2 from the camera images acquired by the image acquisition unit 11. More specifically, the skeleton estimation device 1 uses a machine learning model for 2D skeleton estimation that takes a camera image as input and outputs information indicating the 2D skeleton points of a person on the camera image to estimate 2D skeleton points, and then corrects the estimated 2D skeleton points based on environmental parameters. Then, the skeleton estimation device 1 estimates the 3D skeleton points of the subject in real space using a machine learning model for 3D skeleton estimation, based on the corrected 2D skeleton points. The skeleton estimation device 1 can estimate the 3D skeleton points of a subject while excluding the influence of environmental parameters, thereby preventing a decrease in the estimation accuracy of the 3D skeleton points.

[0047] In the above embodiment 1, the subject was, for example, the driver of a vehicle. However, this is only an example. For example, the subject may be a passenger in the front seat, or a passenger seated in the back seat. Furthermore, the subject is not limited to a passenger in a vehicle, but may be a passenger in various moving objects. For example, the subject may be a forklift driver. Also, for example, the subject may be a person in a living room, or any person in real space, not limited to a passenger in a vehicle. For example, camera 2 is a surveillance camera that monitors a monitoring area, and the subject may be a person present within that monitoring area.

[0048] Furthermore, in the above embodiment 1, the skeleton estimation device 1 is an in-vehicle device mounted on a vehicle, and the image acquisition unit 11, the 2D skeleton estimation unit 121, the 2D skeleton correction unit 122, the 3D skeleton estimation unit 13, and the 3D skeleton information output unit 14 are provided on the in-vehicle device. However, this is merely one example. For example, the skeleton estimation device 1 may be configured such that some of the image acquisition unit 11, the 2D skeleton estimation unit 121, the 2D skeleton correction unit 122, the 3D skeleton estimation unit 13, and the 3D skeleton information output unit 14 are mounted on the in-vehicle device, and the rest are provided on a server connected to the in-vehicle device via a network, thereby configuring the system with the in-vehicle device and the server. Alternatively, the image acquisition unit 11, the 2D skeleton estimation unit 121, the 2D skeleton correction unit 122, the 3D skeleton estimation unit 13, and the 3D skeleton information output unit 14 may all be provided on the server.

[0049] <Example (1)> In the above embodiment 1, the skeleton estimation device 1 estimates the two-dimensional skeleton points of the driver using a two-dimensional skeleton estimation machine learning model based on the camera image, and then corrects the estimated two-dimensional skeleton points based on the distortion parameter. However, this is just one example. For instance, the skeleton estimation device 1 may perform distortion correction on the camera image based on distortion parameters, and then estimate the two-dimensional skeleton points of the driver using a two-dimensional skeleton estimation machine learning model based on the distortion-corrected camera image (hereinafter referred to as the "corrected camera image"). In this case, the estimated two-dimensional skeleton points of the driver will be two-dimensional skeleton points that have been corrected based on distortion parameters, and whose influence on distortion parameters has been removed. In the skeleton estimation device 1, the corrected 2D skeleton estimation unit 12 only needs to be configured to estimate the 2D skeleton points of the driver in the camera image, which have been corrected based on the distortion parameters of the camera 2, from the camera image acquired by the image acquisition unit 11. In this case, the corrected 2D skeleton estimation unit 12 can perform distortion correction on the camera image using known techniques for correcting image distortion, and the 2D skeleton correction unit 122 can do so. This corrects the 2D skeleton correction unit 122 for any 2D skeleton points that may exist on the camera image. Then, the corrected 2D skeleton estimation unit 12 can input the corrected camera image into a machine learning model for 2D skeleton estimation to obtain information indicating the 2D skeleton points, thereby estimating the 2D skeleton points of the driver. In this case, for the example configuration of the corrected 2D skeleton estimation unit 12 shown in Figure 2, the arrow connecting the 2D skeleton estimation unit 121 and the 2D skeleton correction unit 122 becomes an arrow from the 2D skeleton correction unit 122 to the 2D skeleton estimation unit 121. In this case, the order of steps ST11 and ST12 is reversed in the detailed operation of the corrected 2D skeleton estimation unit 12 as explained using Figure 5. In step ST12, the 2D skeleton correction unit 122 performs distortion correction on the camera image. Then, in step ST11, the 2D skeleton estimation unit 121 estimates the 2D skeleton points of the driver using a machine learning model for 2D skeleton estimation, based on the corrected camera image after distortion correction by the 2D skeleton correction unit 122. The corrected 2D skeleton estimation unit 12 outputs information indicating the 2D skeleton points of the driver estimated by the 2D skeleton estimation unit 121 to the 3D skeleton estimation unit 13 as corrected estimated 2D skeleton point information. Even with this configuration, the skeleton estimation device 1 can estimate the subject's 3D skeleton points while excluding the influence of environmental parameters, thus preventing a reduction in the estimation accuracy of the 3D skeleton points.

[0050] <Modification (2)> In the above embodiment 1, the environmental parameter was assumed to be a strain parameter, as an example. The following example illustrates how the skeleton estimation device 1 prevents a reduction in the estimation accuracy of the three-dimensional skeleton points by correcting the two-dimensional skeleton points to eliminate the influence of environmental parameters other than strain parameters, using environmental parameters other than strain parameters as examples. For example, environmental parameters are location parameters.

[0051] First, we will explain, using a diagram, the influence of positional parameters on the two-dimensional skeletal points estimated from camera images, using an example. In the example explained using Figure 6 below, the subject is assumed to be, for example, a person present in a living room. Figures 6A, 6B, and 6C illustrate an example of the influence of positional parameters on two-dimensional skeletal points estimated from camera images. Figure 6A shows a camera image of a child who is far away from camera 2, and Figure 6B shows a camera image of the same child as in Figure 6A, but taken when the child is closer to camera 2 than when the initial image was taken. As shown in Figures 6A and 6B, in camera images, the closer a subject is to camera 2, the larger the subject appears in the image. From the size of the subject in the camera image, it is difficult to determine the subject's physique, for example, whether the subject is an "adult" or a "child." For example, from the camera image shown in Figure 6B, it is difficult to determine whether the subject being captured in the camera image is an "adult" or a "child positioned closer than the standard position." This is because the camera image has lost information about the subject's position relative to camera 2, in other words, its depth.

[0052] For example, when a body size detection device attempts to detect a subject's body size based on three-dimensional skeletal points estimated from two-dimensional skeletal points of the subject estimated from a camera image, the body size detection device determines the subject's body size by considering the positional relationship between the three-dimensional skeletal points. Therefore, if the two-dimensional skeletal points of the subject estimated from the camera image are estimated without considering depth, the detection of the subject's body size using the three-dimensional skeletal points estimated from these two-dimensional points may not be performed correctly. For example, the relative positions of body parts differ between adults and children. However, if the 2D skeletal points are estimated without considering depth, then these differences cannot be taken into account, and body size detection may not be performed correctly. Therefore, for example, if the three-dimensional skeletal points of a subject estimated from a camera image are used to determine the subject's physique, then the two-dimensional skeletal points of the subject estimated from the camera image must take depth into account. In other words, the two-dimensional skeletal points of a subject estimated from a camera image need to be corrected to two-dimensional skeletal points that are free from the influence of position parameters. For example, by correcting the two-dimensional skeletal points of a subject estimated from a camera image based on position parameters to the position of a standard two-dimensional skeletal point estimated from a camera image in which the subject was captured when the subject's position relative to camera 2 is a standard position, the two-dimensional skeletal points of a subject estimated from a camera image can be made to be two-dimensional skeletal points free from the influence of position parameters (see, for example, Figure 6C). As a result, the two-dimensional skeletal points estimated from a camera image have a unified depth, or in other words, the influence of depth is removed. Therefore, the skeleton estimation device 1 may estimate the two-dimensional skeletal points of the subject in the camera image, corrected based on the position parameters of the camera 2, from the camera image.

[0053] Specifically, in the skeleton estimation device 1, the corrected two-dimensional skeleton estimation unit 12 estimates the two-dimensional skeleton points of the subject in the camera image, which are corrected based on the environmental parameters of the camera 2, in this case, the position parameters, from the camera image acquired by the image acquisition unit 11. More specifically, in the skeleton estimation device 1, the 2D skeleton estimation unit 121 estimates the subject's 2D skeleton points using a 2D skeleton estimation machine learning model based on the camera image acquired by the image acquisition unit 11. Then, the 2D skeleton correction unit 122 corrects the subject's 2D skeleton points estimated by the 2D skeleton estimation unit 121 based on position parameters.

[0054] An example of a method for correcting 2D skeleton points based on position parameters using the 2D skeleton correction unit 122 will be described. First, the 2D skeleton correction unit 122 calculates the position parameters. For example, the 2D skeleton correction unit 122 sets a representative point to be used for calculating position parameters based on the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121. Which point to use as the representative point is determined in advance by the developer or the like. Preferably, the representative point is a point corresponding to a body part that is estimated to move little, such as a point corresponding to the position of the neck or a point corresponding to the position between the eyebrows. For example, when the point corresponding to the position of the neck is used as the representative point, the 2D skeleton correction unit 122 can set the 2D skeleton point corresponding to the position of the neck as the representative point. Also, for example, when the point corresponding to the position between the eyebrows is used as the representative point, the 2D skeleton correction unit 122 can set the midpoint of the 2D skeleton points corresponding to the positions of both eyes as the representative point.

[0055] Once a representative point is set, the 2D skeletal correction unit 122 then calculates position parameters based on the position of the set representative point on the camera image, at least in the lateral direction. More specifically, the 2D skeletal correction unit 122 calculates position parameters according to, for example, how much the representative point is shifted to the right or left in the lateral direction, i.e., horizontal direction, relative to the image center. In other words, the 2D skeletal correction unit 122 calculates position parameters according to, for example, the amount of lateral displacement of the representative point on the camera image relative to the image center. For example, the developers may pre-set a formula that determines what value to use for the position parameter when the representative point is shifted horizontally to the right or left relative to the image center. The developers may set the formula so that the position parameter is calculated according to the amount of the shift to the right or left relative to the image center, taking into account where the representative point of the subject is estimated to be captured in the camera image, assuming the subject is at a reference position. The 2D skeletal correction unit 122 calculates position parameters based on the lateral displacement of the representative point relative to the image center and the above calculation formula. The 2D skeletal correction unit 122 can calculate the lateral displacement of the representative point relative to the image center by taking the difference between the x-coordinate of the image center and the x-coordinate of the representative point. Then, the 2D skeleton correction unit 122 corrects the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 based on the estimated position parameters. For example, the 2D skeleton correction unit 122 corrects the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 by moving the position of the 2D skeleton points on a straight line passing through the center of the image and the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121, so that they move closer to or further away from the center of the image by an amount corresponding to the position parameter. As a result, the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 are corrected so that they become the 2D skeleton points estimated from the camera image assuming that the subject is at a reference position predetermined by the developer or the like. In other words, the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 are corrected so that the influence of the position parameter is removed.

[0056] Furthermore, for example, the 2D skeleton correction unit 122 may estimate features other than the 2D skeleton points of the subject from the camera image acquired by the image acquisition unit 11, estimate position parameters based on the estimated features of the subject, and correct the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 based on the estimated position parameters. Features other than the 2D skeleton points of the subject are, for example, features obtained from information on facial parts such as the distance between the eyes or the size of the head. In addition, features other than the 2D skeleton points of the subject may include, for example, features other than the subject's biological features. Features other than the subject's biological features are, for example, the width of the seat belt worn by the subject in the camera image when the subject is an occupant of a vehicle. The width of the seat belt in the camera image changes in size in the camera image under the influence of position parameters, similar to the subject's biological features (for example, the distance between the eyes mentioned above). The 2D skeleton correction unit 122 estimates features other than the 2D skeleton points by, for example, performing known image recognition processing on the camera image. More specifically, the 2D skeletal correction unit 122 estimates positional parameters by comparing the subject's features with reference features. Reference features are pre-defined by the developers. The developers pre-define the reference features by assuming the features of a person of standard build at a reference location. In addition, the developers pre-define the formula for calculating the position parameter based on how much the subject's features are relative to the reference features. The developers pre-define the above formula by considering how much the features estimated to be from the camera image of the subject, assuming the subject is at the reference location. The 2D skeleton correction unit 122 corrects the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 based on the calculated position parameters. The 2D skeleton correction unit 122 should correct the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 in the same manner as in the example described above.

[0057] Alternatively, for example, the 2D skeletal correction unit 122 may acquire information related to the depth from the camera 2 to the subject (hereinafter referred to as "depth-related information") and calculate position parameters based on the acquired depth-related information. Depth-related information includes, for example, information indicating the front-to-back position of the seat in which the subject is seated. For example, if the subject is a vehicle occupant, the 2D skeletal correction unit 122 can obtain depth-related information from a seat position sensor (not shown) installed in the vehicle. Since the installation position of camera 2 is known in advance, the 2D skeletal correction unit 122 can calculate the subject's position relative to camera 2, i.e., its depth, based on the depth-related information, from the position of the seat in which the subject is seated. Then, the 2D skeleton correction unit 122 corrects the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 based on the calculated position parameters. The 2D skeleton correction unit 122 should correct the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 in the same manner as in the example described above.

[0058] The operation of the skeleton estimation device 1 in the <modified version (2)> is basically the same as the operation of the skeleton estimation device 1 described using the flowchart in Figure 4. In the modified version (2), the details of the corrected 2D skeleton estimation process in step ST2 are as follows: The 2D skeleton estimation unit 121 estimates the 2D skeleton points of the driver using a machine learning model for 2D skeleton estimation, based on the camera image acquired by the image acquisition unit 11 in step ST1 of Figure 4 (step ST11 in Figure 5). Then, the 2D skeleton correction unit 122 corrects the 2D skeleton points of the driver estimated by the 2D skeleton estimation unit 121 in step ST11 based on the environmental parameters of the camera 2, in this case, the position parameters (step ST12 in Figure 5).

[0059] Thus, the skeleton estimation device 1 may include an image acquisition unit 11 that acquires camera images of a subject captured by the camera 2, and an image acquisition unit 11 that estimates the two-dimensional skeleton points of the subject in the camera images, corrected based on the position parameters of the camera 2. The skeleton estimation device 1 then estimates the 3D skeleton points of the subject in real space using a machine learning model for 3D skeleton estimation, based on 2D skeleton points corrected based on position parameters. Therefore, the skeleton estimation device 1 can estimate the subject's 3D skeleton points while excluding the influence of environmental parameters, thereby preventing a decrease in the estimation accuracy of the 3D skeleton points.

[0060] In the above embodiment 1, examples of environmental parameters were given as distortion parameters and position parameters, but these are merely examples. Environmental parameters include camera parameters other than distortion parameters and position parameters. The skeleton estimation device 1 can correct the 2D skeleton points of a subject in a camera image by considering various environmental parameters as camera parameters and eliminating the influence of those environmental parameters. Based on the corrected 2D skeleton points of the subject, it can estimate the 3D skeleton points of the subject in real space using a machine learning model for 3D skeleton estimation.

[0061] An example of the hardware configuration of the skeleton estimation device 1 according to Embodiment 1 will be described. Figures 7A and 7B show an example of the hardware configuration of the skeleton estimation device 1 according to Embodiment 1. In Embodiment 1, the functions of the image acquisition unit 11, the 2D skeleton estimation unit 121, the 2D skeleton correction unit 122, the 3D skeleton estimation unit 13, and the 3D skeleton information output unit 14 are realized by the processing circuit 1001. That is, the skeleton estimation device 1 includes a processing circuit 1001 for estimating the 2D skeleton points of the subject in the camera image, corrected based on the environmental parameters of the camera 2, and for controlling the estimation of the 3D skeleton points of the subject in real space using a machine learning model for 3D skeleton estimation based on the estimated 2D skeleton points of the subject. The processing circuit 1001 may be dedicated hardware as shown in Figure 7A, or it may be a processor 1004 that executes a program stored in memory 1005 as shown in Figure 7B.

[0062] If the processing circuit 1001 is dedicated hardware, it may be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a combination thereof.

[0063] When the processing circuit is a processor 1004, the functions of the image acquisition unit 11, the 2D skeleton estimation unit 121, the 2D skeleton correction unit 122, the 3D skeleton estimation unit 13, and the 3D skeleton information output unit 14 are realized by software, firmware, or a combination of software and firmware. The software or firmware is written as a program and stored in memory 1005. The processor 1004 executes the functions of the image acquisition unit 11, the 2D skeleton estimation unit 121, the 2D skeleton correction unit 122, the 3D skeleton estimation unit 13, and the 3D skeleton information output unit 14 by reading and executing the program stored in memory 1005. In other words, the skeleton estimation device 1 includes memory 1005 for storing a program that, when executed by the processor 1004, will result in the execution of steps ST1 to ST4 in Figure 4 described above. Furthermore, the program stored in memory 1005 can be said to cause the computer to execute the processing procedures or methods of the image acquisition unit 11, the 2D skeleton estimation unit 121, the 2D skeleton correction unit 122, the 3D skeleton estimation unit 13, and the 3D skeleton information output unit 14. Here, memory 1005 refers to, for example, non-volatile or volatile semiconductor memory such as RAM, ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), magnetic disks, flexible disks, optical disks, compact disks, minidiscs, DVDs (Digital Versatile Discs), etc.

[0064] Furthermore, the functions of the image acquisition unit 11, the 2D skeleton estimation unit 121, the 2D skeleton correction unit 122, the 3D skeleton estimation unit 13, and the 3D skeleton information output unit 14 may be partially implemented by dedicated hardware and partially by software or firmware. For example, the image acquisition unit 11 can be implemented by a processing circuit 1001 as dedicated hardware, while the functions of the 2D skeleton estimation unit 121, the 2D skeleton correction unit 122, the 3D skeleton estimation unit 13, and the 3D skeleton information output unit 14 can be implemented by the processor 1004 reading and executing a program stored in memory 1005. The first storage unit 15 and the second storage unit 16 are composed of, for example, a memory 1005. Furthermore, the skeleton estimation device 1 includes a camera 2 and other devices, as well as an input interface device 1002 and an output interface device 1003 that perform wired or wireless communication.

[0065] As described above, according to Embodiment 1, the skeleton estimation device 1 is configured to include an image acquisition unit 11 that acquires camera images of a subject captured by a camera 2, a corrected two-dimensional skeleton estimation unit 12 that estimates the two-dimensional skeleton points of the subject in the camera image, corrected based on the environmental parameters of the camera 2, from the camera image acquired by the image acquisition unit 11, and a three-dimensional skeleton estimation unit 13 that estimates the three-dimensional skeleton points of the subject in real space using a machine learning model for three-dimensional skeleton estimation, based on the two-dimensional skeleton points of the subject estimated by the corrected two-dimensional skeleton estimation unit 12. The skeleton estimation device 1 can estimate the 3D skeleton points of a subject while excluding the influence of environmental parameters, thereby preventing a decrease in the estimation accuracy of the 3D skeleton points.

[0066] In the above configuration, the machine learning model for estimating the 3D skeleton used by the 3D skeleton estimation unit 13 to estimate the 3D skeleton points of the subject is a machine learning model that takes information indicating 2D skeleton points unaffected by the environmental parameters as input and outputs information indicating 3D skeleton points corresponding to those 2D skeleton points. Therefore, the skeleton estimation device 1 improves the accuracy of estimating 3D skeleton points, and even if the 3D skeleton estimation machine learning model is trained with a small amount of training data, it can still estimate the 3D skeleton points of a subject using the said 3D skeleton estimation machine learning model. Furthermore, even if the 3D skeleton estimation machine learning model is trained using information about 2D skeleton points based on images that do not take the real environment into account, such as computer graphics (CG), the skeleton estimation device 1 can still estimate the 3D skeleton points of a subject using the 3D skeleton estimation machine learning model.

[0067] More specifically, in the above configuration, the skeleton estimation device 1 includes a corrected 2D skeleton estimation unit 12 which estimates the 2D skeleton points of a subject using a 2D skeleton estimation machine learning model that takes a camera image acquired by the image acquisition unit 11 as input and outputs information indicating the 2D skeleton points of a person on the camera image, and a 2D skeleton correction unit 122 which corrects the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 based on the environmental parameters of the camera 2, and a 3D skeleton estimation unit 13 which estimates the 3D skeleton points of the subject in real space using a 3D skeleton estimation machine learning model based on the 2D skeleton points corrected by the 2D skeleton correction unit 122. Therefore, the skeleton estimation device 1 can estimate the subject's 3D skeleton points while excluding the influence of environmental parameters, thereby preventing a decrease in the estimation accuracy of the 3D skeleton points.

[0068] For example, the environmental parameters include distortion parameters related to the distortion of camera 2, and in the skeleton estimation device 1, the two-dimensional skeleton correction unit 122 corrects the two-dimensional skeleton points of the subject estimated by the two-dimensional skeleton estimation unit 121 based on the distortion parameters. The skeletal estimation device 1 can estimate the 3D skeletal points of a subject while excluding the influence of strain parameters. Therefore, it can provide information indicating the 3D skeletal points that reduces the occurrence of false detections due to strain in various applications (e.g., posture detection) performed using the 3D skeletal points.

[0069] More specifically, for example, the 2D skeleton correction unit 122 calculates distortion parameters based on the distortion on the camera image of the object used for parameter calculation, which is captured in the camera image acquired by the image acquisition unit 11, and corrects the 2D skeleton points estimated by the 2D skeleton estimation unit 121 based on the calculated distortion parameters. As a result, the skeletal estimation device 1 can provide estimated 3D skeletal point information that reduces the occurrence of false detections due to distortion, and can also perform automatic calibration of the camera 2 installed in the vehicle, for example, when the subject is a vehicle occupant.

[0070] Furthermore, for example, the environmental parameters include positional parameters relating to the depth of the subject relative to camera 2, and the 2D skeleton correction unit 122 corrects the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 based on the positional parameters. The skeletal estimation device 1 can estimate the 3D skeletal points of a subject while excluding the influence of positional parameters. Therefore, in various applications that use 3D skeletal points (e.g., body size detection), it can reduce false detections due to depth and provide estimated 3D skeletal point information that takes into account differences in body size, etc.

[0071] More specifically, for example, the 2D skeleton correction unit 122 calculates position parameters based on the position of at least the lateral position of a representative point on the camera image, which is set based on the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121, and corrects the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 based on the calculated position parameters. As a result, the skeleton estimation device 1 can provide estimated 3D skeleton point information that reduces the occurrence of false detections due to depth, without using any additional devices other than the camera 2.

[0072] Furthermore, for example, the 2D skeleton correction unit 122 estimates features other than the 2D skeleton points of the subject from the camera image acquired by the image acquisition unit 11, calculates position parameters based on the estimated features of the subject, and corrects the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 based on the calculated position parameters. The skeletal estimation device 1 uses features other than the subject's two-dimensional skeletal points to calculate position parameters. For example, in situations where a representative point is not captured in the camera image, the device can calculate position parameters with higher accuracy compared to calculating them using two-dimensional skeletal points. As a result, the two-dimensional skeletal estimation unit 121 can correct the subject's two-dimensional skeletal points estimated by the device with higher accuracy.

[0073] Furthermore, for example, the 2D skeleton correction unit 122 acquires depth-related information relating to the depth from the camera 2 to the subject, calculates position parameters based on the acquired depth-related information, and corrects the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 based on the calculated position parameters. The skeleton estimation device 1 can calculate position parameters with higher accuracy by using information related to the depth of the subject in calculating position parameters, and as a result, it can correct the 2D skeleton points of the subject estimated by the 2D skeleton estimation unit 121 with higher accuracy.

[0074] It should be noted that any component of the embodiment can be modified, or any component of the embodiment can be omitted. [Explanation of Symbols]

[0075] 1 Skeleton estimation device, 11 Image acquisition unit, 12 Corrected 2D skeleton estimation unit, 121 2D skeleton estimation unit, 122 2D skeleton correction unit, 13 3D skeleton estimation unit, 14 3D skeleton information output unit, 15 First memory unit, 16 Second memory unit, 100 Skeleton estimation system, 1001 Processing circuit, 1002 Input interface device, 1003 Output interface device, 1004 Processor, 1005 Memory.

Claims

1. An image acquisition unit that acquires camera images of a subject captured by a camera, A corrected two-dimensional skeleton estimation unit estimates the two-dimensional skeleton points of the subject in the camera image, corrected based on the camera's environmental parameters, from the camera image acquired by the image acquisition unit. Based on the two-dimensional skeletal points of the subject estimated by the corrected two-dimensional skeletal estimation unit, a three-dimensional skeletal estimation unit estimates the three-dimensional skeletal points of the subject in real space using a machine learning model for three-dimensional skeletal estimation. A skeletal estimation device equipped with the following.

2. The machine learning model for 3D skeleton estimation is This is a machine learning model that takes information indicating the two-dimensional skeleton points that are not affected by the aforementioned environmental parameters as input and outputs information indicating the three-dimensional skeleton points corresponding to those two-dimensional skeleton points. The skeletal estimation device according to feature 1.

3. The corrected two-dimensional skeleton estimation unit is, A two-dimensional skeleton estimation unit estimates the two-dimensional skeleton points of a subject using a two-dimensional skeleton estimation machine learning model that takes the camera image as input and outputs information indicating the two-dimensional skeleton points of a person on the camera image, based on the camera image acquired by the image acquisition unit. The system includes a two-dimensional skeleton correction unit that corrects the two-dimensional skeleton points of the subject estimated by the two-dimensional skeleton estimation unit based on the environmental parameters of the camera, The three-dimensional skeleton estimation unit estimates the three-dimensional skeleton points of the subject in real space using a machine learning model for three-dimensional skeleton estimation, based on the two-dimensional skeleton points corrected by the two-dimensional skeleton correction unit. A skeletal estimation device according to claim 1 or 2, characterized by the above.

4. The aforementioned environmental parameters include distortion parameters related to the camera's distortion, The aforementioned two-dimensional skeleton correction unit is Based on the aforementioned strain parameters, the two-dimensional skeleton points of the subject estimated by the two-dimensional skeleton estimation unit are corrected. The skeletal estimation device according to claim 3.

5. The aforementioned two-dimensional skeleton correction unit is The image acquisition unit calculates the distortion parameters based on the distortion of the object used for parameter calculation captured in the camera image acquired by the image acquisition unit, and corrects the two-dimensional skeleton points estimated by the two-dimensional skeleton estimation unit based on the calculated distortion parameters. The skeletal estimation device according to feature 4.

6. The aforementioned two-dimensional skeleton correction unit is Based on the aforementioned strain parameters, the coordinates of the two-dimensional skeleton points of the subject estimated by the two-dimensional skeleton estimation unit are converted to coordinates that are not affected by the strain parameters, thereby correcting the two-dimensional skeleton points estimated by the two-dimensional skeleton estimation unit. The skeletal estimation device according to feature 4.

7. The aforementioned environmental parameters include positional parameters relating to the depth of the subject relative to the camera, The aforementioned two-dimensional skeleton correction unit is Based on the position parameters, the two-dimensional skeleton points of the subject estimated by the two-dimensional skeleton estimation unit are corrected. The skeletal estimation device according to claim 3.

8. The aforementioned two-dimensional skeleton correction unit is Based on the position of a representative point set based on the two-dimensional skeletal points of the subject estimated by the two-dimensional skeletal estimation unit, at least in the lateral direction on the camera image, the position parameter is calculated, and the two-dimensional skeletal points of the subject estimated by the two-dimensional skeletal estimation unit are corrected based on the calculated position parameter. The skeletal estimation device according to claim 7.

9. The aforementioned two-dimensional skeleton correction unit is The position parameter is calculated according to the amount of lateral displacement of the representative point relative to the center of the image on the camera image. The skeletal estimation device according to feature 8.

10. The aforementioned two-dimensional skeleton correction unit is The image acquisition unit estimates feature quantities other than the two-dimensional skeletal points of the subject from the camera image acquired by the image acquisition unit, calculates the position parameters based on the estimated feature quantities of the subject, and corrects the two-dimensional skeletal points of the subject estimated by the two-dimensional skeletal estimation unit based on the calculated position parameters. The skeletal estimation device according to claim 7.

11. The aforementioned two-dimensional skeleton correction unit is The position parameter is calculated by comparing the aforementioned feature quantities of the subject with the reference feature quantities. The skeletal estimation device according to claim 10, characterized by its features.

12. The aforementioned two-dimensional skeleton correction unit is Depth-related information relating to the depth from the camera to the subject is acquired, position parameters are calculated based on the acquired depth-related information, and the two-dimensional skeleton points of the subject estimated by the two-dimensional skeleton estimation unit are corrected based on the calculated position parameters. The skeletal estimation device according to claim 7.

13. The aforementioned depth-related information is information indicating the front-to-back position of the seat in which the subject is seated. The skeletal estimation device according to claim 12, characterized by its features.

14. Computers, An image acquisition unit that acquires camera images of a subject captured by a camera, A corrected two-dimensional skeleton estimation unit estimates the two-dimensional skeleton points of the subject in the camera image, corrected based on the camera's environmental parameters, from the camera image acquired by the image acquisition unit. Based on the two-dimensional skeletal points of the subject estimated by the corrected two-dimensional skeletal estimation unit, a three-dimensional skeletal estimation unit estimates the three-dimensional skeletal points of the subject in real space using a machine learning model for three-dimensional skeletal estimation. A skeletal estimation program designed to function as such.

Citation Information

Patent Citations

  • Filtering method, filtering program, and filtering apparatus

    JP2022049521A