Learning device, learning method, learning program, behavior detection device, behavior detection method, and behavior detection program
The learning device generates multiple learning models by transforming three-dimensional skeletal coordinates onto two-dimensional planes with different camera parameters, addressing high costs and improving behavior detection accuracy across varying installation heights.
Patent Information
- Application Number
- JP2024576715
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-10-04
- Publication Date
- 2025-10-20
- Estimated Expiration
- 2043-10-04
AI Technical Summary
Conventional devices requiring multiple detection units for different installation heights incur high costs due to the need for separate learning data preparation for each unit.
A learning device that generates multiple learning models by projecting and transforming three-dimensional skeletal coordinates onto two-dimensional planes using different camera parameters, allowing for low-cost behavior inference and detection using selected models.
Enables accurate and cost-effective behavior detection by generating and selecting appropriate learning models for various installation conditions, reducing the need for multiple detection units.
Smart Images

Figure 0007756817000001 
Figure 0007756817000002 
Figure 0007756817000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a learning device, a learning method, a learning program, a behavior detection device, a behavior detection method, and a behavior detection program. [Background technology]
[0002] A device has been proposed that supports monitoring of a subject in bed (see, for example, Patent Document 1). This device has an image acquisition unit that acquires images captured by an imaging device, an information acquisition unit that acquires information related to the installation height of the imaging device, and a plurality of detection units (i.e., a plurality of modules) that correspond to the installation heights of the imaging device, respectively, for detecting the subject or the subject's condition from the acquired images. In this device, of the plurality of detection units, the detection unit that corresponds to the installation height of the imaging device is used to detect the subject or the subject's condition. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent No. 6729510 (see, for example, claim 1, paragraphs 0023-0027, and figures 1-3) Summary of the Invention [Problem to be solved by the invention]
[0004] However, since the above-mentioned conventional device has multiple detection units corresponding to multiple installation heights, it is necessary to prepare learning data for each of the multiple detection units, which results in high costs for designing the device.
[0005] The present disclosure aims to provide a learning device, a learning method, and a learning program that enable the generation of multiple learning models for inferring behavior at low cost, as well as a behavior detection device, a behavior detection method, and a behavior detection program that accurately detect behavior using a learning model selected from multiple learning models. [Means for solving the problem]
[0006] The learning device of the present disclosure is a device that receives three-dimensional learning data including three-dimensional learning skeletal coordinates including three-dimensional coordinates of a skeleton of a target and teacher data consisting of correct answers for actions linked to the skeleton, and generates a learning model, the device comprising: a two-dimensional projection transformation unit that projects and transforms the three-dimensional learning skeletal coordinates onto a first two-dimensional plane according to predetermined first camera parameters to generate first two-dimensional learning data including first two-dimensional learning skeletal coordinates and the teacher data, and projects and transforms the three-dimensional learning skeletal coordinates onto a second two-dimensional plane according to predetermined second camera parameters to generate second two-dimensional learning data including second two-dimensional learning skeletal coordinates and the teacher data; The apparatus includes a first model generation unit that generates a first learning model for inferring an action associated with a first two-dimensional skeletal coordinate for inference from a first two-dimensional skeletal coordinate for inference using two-dimensional data, and a second model generation unit that generates a second learning model for inferring an action associated with a second two-dimensional skeletal coordinate for inference from a second two-dimensional skeletal coordinate for inference using the second two-dimensional skeletal data, wherein the two-dimensional projection transformation unit generates three-dimensional skeletal coordinates for training by interpolation or thinning from original data of a plurality of three-dimensional skeletal coordinates for training at a plurality of times arranged in time series, and changes the number of three-dimensional skeletal coordinates for training to obtain three-dimensional skeletal data for training at a speed different from that of the original data. and further generating the first two-dimensional training data and the second two-dimensional training data based on the three-dimensional training skeleton data at different speeds. It is characterized by:
[0007] The behavior detection device disclosed herein is a device that detects behavior performed by a subject included in a two-dimensional image acquired by a two-dimensional image acquisition unit, and is characterized by having a two-dimensional skeletal coordinate calculation unit that calculates two-dimensional skeletal coordinates for two-dimensional image-derived inference that indicate the coordinates of the subject's skeleton from the two-dimensional image, and a behavior detection unit that selects one or more learning models from the first learning model and the second learning model generated by the learning device, and detects the behavior from the two-dimensional skeletal coordinates for two-dimensional image-derived inference using the selected one or more learning models. [Effects of the Invention]
[0008] By using the learning device, learning method, and learning program of the present disclosure, it is possible to generate multiple learning models for inferring behavior at low cost.
[0009] Furthermore, by using the behavior detection device, behavior detection method, and behavior detection program of the present disclosure, it is possible to accurately detect behavior using a learning model selected from a plurality of learning models. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a functional block diagram schematically illustrating a configuration of a learning device according to a first embodiment. [Figure 2] FIG. 10 is a conceptual diagram showing a process of generating two-dimensional training skeleton coordinates from three-dimensional training skeleton coordinates in accordance with first camera parameters. [Figure 3] FIG. 10 is a conceptual diagram showing a process of generating two-dimensional training skeleton coordinates from three-dimensional training skeleton coordinates in accordance with first camera parameters. [Figure 4] FIG. 10 is a conceptual diagram showing the process of generating two-dimensional training skeleton coordinates from three-dimensional training skeleton coordinates in accordance with second camera parameters. [Figure 5] FIG. 10 is a conceptual diagram showing the process of generating two-dimensional training skeleton coordinates from three-dimensional training skeleton coordinates according to the third camera parameters. [Figure 6] 4 is a flowchart schematically illustrating the operation of the learning device according to the first embodiment. [Figure 7] (A) is a diagram showing the process of rotating the learning two-dimensional skeletal coordinates generated from the learning three-dimensional skeletal coordinates around the y-axis, (B) is a diagram showing the process of rotating the learning two-dimensional skeletal coordinates generated from the learning three-dimensional skeletal coordinates around the z-axis, and (C) is a diagram showing the process of rotating the learning two-dimensional skeletal coordinates generated from the learning three-dimensional skeletal coordinates around the x-axis. [Figure 8]10 is a flowchart showing an outline of the operation of the learning device according to the first embodiment (when rotation around the x-axis and rotation around the y-axis is included, and there is no rotation around the z-axis). [Figure 9] 10(A) and 10(B) are diagrams showing the state in which the training two-dimensional skeletal coordinates generated from the training three-dimensional skeletal coordinates are rotated 90° around the z-axis. [Figure 10] 10(A) and 10(B) are diagrams showing the state in which the training two-dimensional skeleton coordinates generated from the training three-dimensional skeleton coordinates are rotated 90° around the x-axis. [Figure 11] FIG. 10 is an explanatory diagram showing an example of bone extension of a training three-dimensional skeletal coordinate system. [Figure 12] FIG. 10 is an explanatory diagram showing the interpolation process of the learning two-dimensional skeleton coordinates generated from the learning three-dimensional skeleton coordinates. [Figure 13] 1 is a diagram illustrating an example of a hardware configuration of a learning device according to a first embodiment. [Figure 14] FIG. 10 is a functional block diagram illustrating a schematic configuration of a behavior detection device according to a second embodiment. [Figure 15] 10A, 10B, and 10C are conceptual diagrams showing two-dimensional images of a person taken by cameras with different depression angles and installation positions. [Figure 16] This is a conceptual diagram showing a two-dimensional image obtained by camera photography, two-dimensional skeletal coordinates for inference generated from the two-dimensional image, and behavior detection according to a first learning model. [Figure 17] This is a conceptual diagram showing two-dimensional images obtained by camera photography, two-dimensional skeletal coordinates for inference generated from the two-dimensional images, and behavior detection according to a second learning model. [Figure 18] This is a conceptual diagram showing two-dimensional images obtained by camera photography, two-dimensional skeletal coordinates for inference generated from the two-dimensional images, and behavior detection according to a third learning model. [Figure 19] 10 is a flowchart schematically illustrating the operation of the behavior detection device according to the second embodiment. [Figure 20] FIG. 10 is a diagram illustrating an example of a hardware configuration of a behavior detection device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] A learning device, a learning method, a learning program, a behavior detection device, a behavior detection method, and a behavior detection program according to embodiments will be described below with reference to the drawings. The following embodiments are merely examples, and the embodiments can be combined as appropriate and each embodiment can be modified as appropriate.
[0012] First Embodiment FIG. 1 is a functional block diagram showing a schematic configuration of a learning device 10 according to the first embodiment. Learning device 10 is a device capable of implementing the learning method according to the first embodiment. Learning device 10 is, for example, a computer capable of executing the learning program according to the first embodiment. Learning device 10 may be a computer (i.e., a computer system) configured by cloud computing using a computer network.
[0013] The learning device 10 is a device that enables low-cost generation of multiple learning models used to infer behaviors performed by a subject. The subject is, for example, a human subject. The subject may be an animal with a skeleton (including humans and non-human animals). The subject may include a robot with joints (for example, a humanoid robot). The behaviors performed by the subject may include the movement of the subject, the state of the subject, and the posture of the subject. The behaviors include standing, walking, falling, lying, etc.
[0014] The learning device 10 can generate (including "update") multiple learning models from the three-dimensional training data D10 by using projective transformation. Projective transformation will be described below with reference to FIGS. 2 to 5. The learning device 10 includes a two-dimensional projection transformation unit 12 and multiple model generation units. The learning device 10 may also include a three-dimensional data acquisition unit 11. The three-dimensional data acquisition unit 11 may be a part of the learning device 10 or an external device that provides data to the learning device 10. In the first embodiment, the multiple model generation units include two model generation units, i.e., a first model generation unit 13 and a second model generation unit 14. The multiple model generation units may also include three or more model generation units. In other words, the multiple learning models may include a first learning model M1, a second learning model M2, a third learning model M3, etc. The first learning model M1, the second learning model M2, and the third learning model M3 are stored in storage devices 15, 16, and 17, respectively. The storage devices 15, 16, and 17 may be the same or different storage devices, and may be part of the learning device 10 or may be storage devices in external devices.
[0015] The three-dimensional data acquisition unit 11 acquires three-dimensional training data D10. The three-dimensional data acquisition unit 11 is, for example, a detector such as a stereo camera, a time-of-flight (TOF) camera, motion capture, or a light detection and ranging (LiDAR). A stereo camera is an imaging device with a distance measurement function based on the principles of trigonometry from two-eye images. A TOF camera is an imaging device with a function to measure distance using the time it takes for irradiated light to reflect and return. A motion capture device is a device with a function to attach markers to a person or object and digitally record their movements. A LiDAR is a device that emits electromagnetic waves and can measure distance based on the time until the reflected waves are detected and speed by measuring the Doppler shift of the reflected waves. The three-dimensional training data D10 includes three-dimensional training skeletal coordinates D11 including the three-dimensional coordinates of the target's skeleton and training data D12 consisting of correct answers for actions linked to the skeleton. The three-dimensional data acquisition unit 11 may also have a function of calculating three-dimensional skeletal coordinates for training from images captured by a monocular two-dimensional camera using deep learning or the like. The three-dimensional data acquisition unit 11 may also receive the three-dimensional skeletal coordinates for training from a database related to skeletons stored in a storage device. The storage device is shown in FIG. 12, which will be described later.
[0016] The two-dimensional projection transformation unit 12 acquires a plurality of camera parameters of a plurality of virtual cameras, for example, from a plurality of virtual cameras or from a storage device (for example, shown in FIG. 12 described later) that stores a plurality of camera parameters in advance. The plurality of camera parameters include information about the position and orientation (for example, optical axis direction) of each of the plurality of virtual cameras that generate the training three-dimensional skeleton coordinates D11. In the first embodiment, the plurality of camera parameters include a predetermined first camera parameter P1 used to generate the first learning model M1 and a predetermined second camera parameter P2 used to generate the second learning model M2. The plurality of camera parameters may also include three or more predetermined camera parameters that are different from each other, and the two-dimensional projection transformation unit 12 may calculate three or more training two-dimensional skeleton coordinates using the three or more camera parameters.
[0017] When generating the two-dimensional training skeleton coordinates using perspective projection, each of the multiple camera parameters preferably includes coordinates in the world coordinate system of the virtual camera to be set and the focal length of the virtual camera. Furthermore, each of the multiple camera parameters is defined as a rotation angle r about three axes consisting of the x-axis, y-axis, and z-axis with respect to the three-dimensional training skeleton coordinates D11, where the vertical coordinate axis (the direction of gravity on the Earth) is defined as the y-axis, the coordinate axis on the line obtained by projecting a straight line in the projection direction used when generating the multiple two-dimensional training skeleton coordinates (for example, the first two-dimensional training skeleton coordinates D21a and the second two-dimensional training skeleton coordinates D21b) onto a horizontal plane that is perpendicular to the y-axis is defined as the z-axis, and the coordinate axis perpendicular to both the y-axis and the z-axis is defined as the x-axis. x ,r y ,r z It is desirable to include a rotation angle (r y1 ) and the rotation angle around the y-axis according to the second camera parameter P2 (r y2 ) can be chosen to any value.
[0018] Also, the rotation angle (r x1) as a predetermined angle r that is greater than or equal to -90° and less than or equal to 0° x1min from a predetermined angle r that is greater than or equal to 0° and less than or equal to 90° x1max Within the range of r x1min ≦r x1 ≦r x1max Similarly, the rotation angle around the x-axis according to the second camera parameter P2 (r x2 ) as a predetermined angle r that is greater than or equal to -90° and less than or equal to 0° x2min from a predetermined angle r that is greater than or equal to 0° and less than or equal to 90° x2max Within the range of r x2min ≦r x2 ≦r x2max It is possible to select an angle where
[0019] Furthermore, the rotation angle (r z1 ) as a predetermined angle r that is greater than or equal to -90° and less than or equal to 0° zmin1 from a predetermined angle r that is greater than or equal to 0° and less than or equal to 90° z1max Within the range of r z1min ≦r z1 ≦r z1max Similarly, the rotation angle around the z-axis according to the second camera parameter P2 (r z2 ) as a predetermined angle r that is greater than or equal to -90° and less than or equal to 0° z2min from a predetermined angle r that is greater than or equal to 0° and less than or equal to 90° z2max Within the range of r z2min ≦r z2 ≦r z2max It is possible to select an angle where
[0020] FIG. 7(A) shows the rotation angle r of the training 3D skeleton coordinate D11 around the y axis. y 7(B) shows the process of rotating the training 3D skeleton coordinates D11 around the z axis by a rotation angle r z 7(C) shows the process of rotating the training 3D skeleton coordinates D11 around the x-axis by a rotation angle r xFIG. 10 is a diagram showing a process of rotating by [°].
[0021] FIG. 2 is a conceptual diagram showing the process of generating first training two-dimensional skeleton coordinates D21a from training three-dimensional skeleton coordinates D11 by parallel projection according to first camera parameters P1. In parallel projection, the relative movement of each axis does not affect the shape of the projection. In parallel projection, the parameters that affect the projection are the y-axis (vertical axis), the z-axis (the direction of the line projected onto the horizontal plane (xz plane) that is perpendicular to the vertical axis, i.e., the x-axis), and the rotation angle of the x-axis (the axis perpendicular to the y-axis and z-axis). For ease of explanation, the z-axis is positioned so as to coincide with the line obtained by projecting the projection direction line 21 onto the horizontal plane (xz plane).
[0022] FIG. 2 shows an example in which the optical axis direction of the virtual camera 20 (first virtual camera) is the z-axis direction, i.e., r x1 = 0°. The two-dimensional projection transformation unit 12 performs projection transformation of the training three-dimensional skeleton coordinates D11 onto a first two-dimensional plane 22a, which is a projection plane, in accordance with the first camera parameters P1, to generate first training two-dimensional data D21 including the first training two-dimensional skeleton coordinates D21a and teacher data. The first two-dimensional plane 22a is, for example, a virtual plane perpendicular to the projection direction (straight line 21) of the virtual camera 20. The first two-dimensional plane 22a may be a non-flat surface such as a sphere, a cylinder, a paraboloid, a hyperboloid, or a hyperbolic paraboloid, or a part thereof, or a surface of a non-linear coordinate system for correcting geometric distortion.
[0023] Perspective projection can also be used for the projection transformation. Fig. 3 is a conceptual diagram showing the process of generating first training two-dimensional skeleton coordinates D21a' from training three-dimensional skeleton coordinates D11 by perspective projection according to the first camera parameter P1'. Compared to parallel projection, the difference is that the projected image changes depending on the relative positional relationship between the optical axis of the virtual camera and the training three-dimensional skeleton coordinates.
[0024] 4 is a conceptual diagram showing the process of generating second training two-dimensional skeleton coordinates D21b from training three-dimensional skeleton coordinates D11 by parallel projection in accordance with the second camera parameter P2. FIG. 4 shows an example in which the projection direction of the virtual camera 20 (second virtual camera) is inclined with respect to the z-axis, that is, r x2 2. The two-dimensional projection transformation unit 12 performs projection transformation of the training three-dimensional skeleton coordinates D11 onto a second two-dimensional surface 22b, which is a projection surface, in accordance with the second camera parameters P2, to generate second training two-dimensional data D22 including the second training two-dimensional skeleton coordinates D21b and teacher data. The second two-dimensional surface 22b is, for example, a virtual plane perpendicular to the projection direction (straight line 21) of the virtual camera 20. The second two-dimensional surface 22b may be a non-flat surface such as a sphere, a cylinder, a paraboloid, a hyperboloid, or a hyperbolic paraboloid, or a part thereof, or a surface of a non-linear coordinate system for correcting geometric distortion.
[0025] 5 is a conceptual diagram showing the process of generating third training two-dimensional skeleton coordinates D21c from training three-dimensional skeleton coordinates D11 in accordance with the third camera parameter P3. FIG. 5 shows an example in which the projection direction of the virtual camera 20 is inclined with respect to the z-axis, that is, r x3 >r x2 >0° is shown. The two-dimensional projection transformation unit 12 performs projection transformation of the training three-dimensional skeleton coordinates D11 onto a third two-dimensional plane 22c, which is a projection plane, in accordance with the third camera parameters P3, to generate third training two-dimensional data including the third training two-dimensional skeleton coordinates D21c and the teacher data. The third two-dimensional plane 22c is, for example, a virtual plane perpendicular to the projection direction (straight line 21) of the virtual camera 20. The third two-dimensional plane 22c may be a non-flat surface such as a sphere, a cylinder, a paraboloid, a hyperboloid, or a hyperbolic paraboloid, or a part thereof, or a surface of a non-linear coordinate system for correcting geometric distortion.
[0026] 5 is a flowchart schematically illustrating the operation of learning device 10 according to embodiment 1. First, learning device 10 acquires three-dimensional learning data D10 including three-dimensional learning skeletal coordinates D11 including three-dimensional coordinates of the skeleton of a target and teacher data D12 consisting of correct answers for actions associated with the skeleton (step S11).
[0027] Next, the learning device 10 generates first two-dimensional training data D21 including first two-dimensional training coordinates D21a and teacher data by projecting the training three-dimensional skeleton coordinates D11 onto the first two-dimensional plane 22a according to the predetermined first camera parameters P1 (step S12). Also, the learning device 10 generates second two-dimensional training data D22 including second two-dimensional training coordinates D21b and teacher data by projecting the training three-dimensional skeleton coordinates D11 onto the second two-dimensional plane 22b according to the predetermined second camera parameters P2 (step S12).
[0028] Next, the learning device 10 executes a learning process using the first two-dimensional training data D21 to generate (including "update") a first learning model M1 for inferring behavior associated with the first two-dimensional training skeletal coordinates from the first two-dimensional training skeletal coordinates, which are new two-dimensional skeletal coordinates (step S13). The learning device 10 also executes a learning process using the second two-dimensional training data D22 to generate (including "update") a second learning model M2 for inferring behavior associated with the second two-dimensional training skeletal coordinates from the second two-dimensional training skeletal coordinates, which are new two-dimensional skeletal coordinates (step S13). The first two-dimensional training skeletal coordinates and the second two-dimensional training skeletal coordinates are different coordinates, but may also be the same coordinates.
[0029] Next, the learning device 10 stores the generated learning models in a storage device (step S14).
[0030] The learning device 10 rotates the learning three-dimensional skeletal coordinates D11 of Figures 7(A) to (C) in various ways around the three axes of the x-axis, y-axis, and z-axis, and uses the output at that time to update multiple learning models (e.g., a first learning model M1 and a second learning model M2, etc.).
[0031] FIG. 8 is a flowchart outlining the operation of the learning device 10 according to the first embodiment (including rotation around the x-axis and rotation around the y-axis, but no rotation around the z-axis). In FIG. 8, three-dimensional learning data D10 including three-dimensional learning skeletal coordinates D11 is sequentially input to the learning device 10 (steps S101 and S102). The three-dimensional learning skeletal coordinates D11 are then rotated around three axes (steps S103 to S107, S110 to S113, and S116 to S119). The data projected onto a two-dimensional plane (projection plane) is used (steps S108 and S114), and the first learning model M1 (step S109) and the second learning model M2 (step S115) are updated. When parallel projection is used, the coordinates on the two-dimensional plane (projection plane) are the x- and y-coordinate values of the rotated three-dimensional coordinates.
[0032] Rotation angle r around the y-axis y It is desirable that the rotation angle r covers the entire range from 0° to 360°. This is because if the rotation around the y-axis is limited to a narrow range, it may not be possible to detect certain actions (for example, a fall on the right side may be detected, but a fall on the left side may not be detected), which would reduce the action detection performance of the inference device (described in the second embodiment). In the process of Figure 8, the rotation angle r y is changed as a parameter in increments of 10° within the range from 0° to 350° (steps S103, S118, S119).
[0033] In the process of Figure 8, the rotation angle r around the z axis z 9(A) and (B) are diagrams showing the state in which the training 2D skeleton coordinates generated from the training 3D skeleton coordinates are rotated around the z axis. As shown in FIG. 9(A), r z At r = 0°, the training 2D skeleton coordinates are in an upright posture, but z At r = 90°, the training 2D skeleton coordinates are in a lying position. z At r = 0°, the training 2D skeletal coordinates are in a supine position, but z = -90°, the training 2D skeleton coordinates are in an upright position.z It is desirable to limit the rotation angle r about the z-axis to a certain extent (for example, less than 90°, preferably less than 60°). In other words, the two-dimensional projection transformation unit 12 performs projection transformation on the training three-dimensional skeleton coordinates to generate the first training two-dimensional data. z within a first range from a predetermined first angle of -90° or more and 0° or less to a predetermined second angle of 0° or more and 90° or less (for example, the angle from the first angle to the second angle is set within a range of 60°), and a rotation angle r about the z axis when generating second training two-dimensional data by projectively transforming the training three-dimensional skeleton coordinates. z It is desirable to set the angle within a second range from a predetermined third angle that is greater than or equal to -90° and less than or equal to 0° to a predetermined fourth angle that is greater than or equal to 0° and less than or equal to 90° (for example, setting the angle from the third angle to the fourth angle within a range of 60°).
[0034] However, in general (except in a zero-gravity environment such as outer space), people adopt a posture that corresponds to gravity. Therefore, the rotation angle r z It is desirable to limit the rotation angle to a narrow range centered on 0°, ideally 0° from the vertical direction, and more preferably within a narrow range such as -2° to +2° from the vertical direction. z = 0°, but the rotation angle r z may vary.
[0035] In the process of Figure 8, the rotation angle r around the x-axis x 9(A) and (B) are diagrams showing the state in which the training 2D skeleton coordinates generated from the training 3D skeleton coordinates are rotated around the x-axis. As shown in FIG. 9(A), r x At r = 0°, the training 2D skeleton coordinates are in an upright posture, but x At r = 90°, the training 2D skeleton coordinates are in a lying position. z At r = 0°, the training 2D skeletal coordinates are in a supine position, but x= 90°, the training 2D skeleton coordinates are in an upright position. x In other words, the two-dimensional projection transformation unit 12 limits the rotation angle r about the x-axis when generating the first two-dimensional training data by projectively transforming the three-dimensional training skeleton coordinates. x is set to a third range that is greater than a predetermined fifth angle within the range of 0° to 90° and less than a predetermined sixth angle within the range of 0° to 90° (for example, the angle from the fifth angle to the sixth angle is set to a range of 60°), and a rotation angle r about the x-axis when generating second two-dimensional training data by projectively transforming the training three-dimensional skeleton coordinates is set to a third range that is greater than a predetermined fifth angle within the range of 0° to 90° (for example, the angle from the fifth angle to the sixth angle is set to a range of 60°), and x is preferably set within a fourth range that is greater than a seventh predetermined angle within the range of 0° to 90° and less than an eighth predetermined angle within the range of 0° to 90° (e.g., the angle from the sixth angle to the seventh angle is set within a range of 60°).
[0036] However, from the viewpoint of practical convenience, the rotation angle r x It is desirable to make r broad. That is, the situation used in the study, x If the situation at the time of inference, that is, the rotation angle around the x-axis at the time of inference, specifically the depression angle which is the angle at which the camera (camera 50 in the second embodiment) looks down on the ground, differs significantly, the accuracy of the inference will decrease.
[0037] On the other hand, for surveillance cameras that monitor care recipients and surveillance cameras for surveillance purposes, the ideal installation angle for the camera varies depending on the shape of the room or the position of obstructions, making it difficult to determine it solely for the convenience of the fall detector as a behavior detection device. For example, in an installation situation where the ceiling is low and the depth is long, it is desirable to set the angle of view down to be relatively shallow (small rotation angle around the x-axis), and in an installation situation where the ceiling is high and the depth is short, it is desirable to set the angle of view down to be relatively deep (large rotation angle around the x-axis). In order to perform behavior detection robustly under these various conditions, the rotation angle r x It is desirable to widen the scope of
[0038] To address this issue, we prepare multiple model generators (two in Figure 8), each with a rotation angle r around the x-axis. x In FIG. 8, the range of the rotation angle r x1 is set to three values of 20°, 30°, and 40°, and the r for the second training two-dimensional skeleton coordinates is input to the second model generation unit 14. x2 The rotation angle r is set to 40°, 50°, and 60°. x When the range is around 30°, phenomena such as the confusion between upright and lying position mentioned above do not occur, and rather the generalization characteristics resulting from the variation in the data have the effect of increasing the detection rate.
[0039] This process is performed on a large number of training three-dimensional skeletons to construct first and second training models.
[0040] Furthermore, the two-dimensional projection transformation unit 12 may generate the first two-dimensional training data D21 and the second two-dimensional training data D22 from the three-dimensional training skeletal coordinates obtained by adding a random number within a predetermined range to the three-dimensional training skeletal coordinates D11. In this case, training data for the skeletons of a plurality of people with different physiques can be obtained (i.e., the variation in physiques can be increased).
[0041] The two-dimensional projection transformation unit 12 may also have a function (i.e., a bone extension function) for generating first and second training two-dimensional data from three-dimensional skeletal coordinates obtained by multiplying by a scalar value the difference vector between the first and second skeletal points determined by the training three-dimensional skeletal coordinates. Bone extension may be performed to generate first and second training two-dimensional data D21 and D22 from the training three-dimensional skeletal coordinates obtained by extending or shortening the difference vector between skeletal points determined by the training three-dimensional skeletal coordinates D11. In this case, different training two-dimensional data can be obtained with only the cost of calculation processing. FIG. 11 is an explanatory diagram showing an example of bone extension of the training three-dimensional skeletal coordinates D11. This is an example in which, instead of retaking data of a person with relatively long legs, the bones from both hips to both knees and the bones from both knees to both ankles of the training three-dimensional skeleton are each multiplied by 1.1. By calculating the coordinates of both knees and both ankles from both hips, training data of a person with long legs can be obtained without any problems.
[0042] 12 is an explanatory diagram showing the interpolation process of the learning three-dimensional skeletal coordinates D11. The two-dimensional projection transformation unit 12 may generate learning three-dimensional skeletal coordinates interpolated from multiple learning three-dimensional skeletal coordinates at multiple times arranged in chronological order, thereby increasing the number of learning three-dimensional skeletal coordinates.
[0043] The upper part of Figure 12 shows the original data (constant speed) for the learning 3D skeletal coordinates D11. This data is based on learning 3D skeletal coordinates at three times in a chronological order, and the lower part of Figure 12 shows an example of converting this data into 2 / 3 speed, i.e., slower, behavioral data. The data for the first time remains unchanged, but the learning 3D skeletal coordinates for the next time at 2 / 3 speed are obtained by interpolation between the learning 3D skeletal coordinates for the first time at constant speed and the learning 3D skeletal coordinates for the second time. Similarly, the learning 3D skeletal coordinates for the second time at 2 / 3 speed are obtained by interpolation between the learning 3D skeletal coordinates for the second time at constant speed and the learning 3D skeletal coordinates for the third time. The learning 3D skeletal coordinates for the last time at 2 / 3 speed are the same as the learning 3D skeletal data for the last time at constant speed.
[0044] In this way, by using interpolation or thinning, three-dimensional skeletal data for learning at different speeds can be obtained by calculation processing on a computer.
[0045] In addition, in behavior detection using skeletal data, skeletal data from multiple time points may be input. In such cases, data showing the same behavior performed quickly and slowly can be obtained from a single set of 3D training skeletal data.
[0046] 12 is a diagram illustrating an example of the hardware configuration of a learning device 10 according to embodiment 1. The learning device 10 includes a processor 101 such as a CPU (Central Processing Unit), a memory 102 serving as a storage device such as a RAM (Random Access Memory), a storage device 103 serving as a non-volatile storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), and an interface 104. These components may be configured using dedicated processing circuitry.
[0047] The processor 101 can execute the learning program according to the first embodiment. The learning program is provided by being recorded on a recording medium such as an SD (Secure Digital memory card) or a USB (Universal Serial Bus) memory card, or by being downloaded via a network. The hardware configuration shown in Fig. 12 is an example, and various modifications to the hardware configuration are possible.
[0048] As described above, by using the learning device, learning method, and learning program according to the first embodiment, it is possible to generate a plurality of learning models for inferring behavior at low cost.
[0049] Second Embodiment FIG. 14 is a functional block diagram showing a schematic configuration of a behavior detection device 40 serving as an inference device according to the second embodiment. The behavior detection device 40 is a device capable of implementing a behavior detection method serving as an inference method according to the second embodiment. The behavior detection device 40 is, for example, a computer capable of executing a behavior detection program according to the second embodiment. The behavior detection device 40 may be a computer (i.e., a computer system) configured by cloud computing using a computer network.
[0050] The behavior detection device 40 is a device that detects behaviors performed by a target included in a two-dimensional image acquired by a two-dimensional image acquisition unit 41. The target is described in the first embodiment and is, for example, a person. The two-dimensional image acquisition unit 41 includes, for example, a camera 50. The behavior detection device 40 includes a two-dimensional skeletal coordinate calculation unit 42 as a two-dimensional image-derived two-dimensional skeletal coordinate calculation unit that calculates two-dimensional skeletal coordinates for inference D42 (hereinafter also referred to as "two-dimensional skeletal coordinates for inference D42") indicating the coordinates of the target's skeleton from the two-dimensional image, and a behavior detection unit 44 that selects one or more learning models from a plurality of learning models (for example, a first learning model M1 and a second learning model M2) and detects behaviors from the two-dimensional skeletal coordinates for inference D42 using the selected one or more learning models. The plurality of learning models are generated by the learning device 10 according to the first embodiment. The behavior detection device 40 may use three or more learning models. Furthermore, the storage devices 15 and 16 in which the learning models are stored may be part of the behavior detection device 40. In addition, the whole or part of the two-dimensional image acquisition unit 41 may be a part of the behavior detection device 40.
[0051] The behavior detection device 40 has a camera information acquisition unit 43 as a two-dimensional image acquisition unit information acquisition unit that acquires or calculates the depression angle of the optical axis when a two-dimensional image is captured by the two-dimensional image acquisition unit 41. The behavior detection unit 44 selects one or more learning models from among a plurality of learning models based on the depression angle, and uses the selected one or more learning models to detect behavior.
[0052] The camera information acquisition unit 43 can acquire the depression angle of the camera 50 from a storage device that stores the depression angle of the camera 50 (i.e., the depression angle of the shooting direction by the two-dimensional image acquisition unit 41). The camera information acquisition unit 43 may also acquire (calculate) the depression angle of the camera 50 based on the output of an acceleration sensor fixed to the camera 50 in the direction of gravitational acceleration.
[0053] Figures 15(A), (B), and (C) are conceptual diagrams showing two-dimensional images of a person captured by a camera 50 with different depression angles and installation positions. Figure 15(A) shows an example where the depression angle is small, Figure 15(C) shows an example where the depression angle is large, and Figure 15(B) shows an example where the depression angle is an intermediate value between Figures 15(A) and (C).
[0054] Figure 16 is a conceptual diagram showing a two-dimensional image D41 obtained by camera photography when the camera 50 has the depression angle of Figure 15(A), two-dimensional skeletal coordinates D42 for inference generated from the two-dimensional image D41, behavior detection according to the first learning model M1, and behavior (or behavior likelihood) D44 as the inference result.
[0055] Figure 17 is a conceptual diagram showing a two-dimensional image D41 obtained by camera photography when the camera 50 has the depression angle of Figure 15(B), two-dimensional skeletal coordinates D42 for inference generated from the two-dimensional image D41, behavior detection according to the second learning model M2, and behavior (or behavior likelihood) D44 as the inference result.
[0056] Figure 18 is a conceptual diagram showing a two-dimensional image D41 obtained by camera photography when the camera 50 has the depression angle of Figure 15(C), two-dimensional skeletal coordinates D42 for inference generated from the two-dimensional image D41, behavior detection according to the third learning model M3, and behavior (or behavior likelihood) D44 as the inference result.
[0057] It is desirable that the behavior detection unit 44 selects, from among a plurality of learning models (for example, a first learning model, a second learning model, and a third learning model), a learning model whose rotation angle around the x-axis is closer to the depression angle of the camera 50. In other words, the behavior detection unit 44 selects the first learning model when the rotation angle around the x-axis of the first camera parameter is closer to the depression angle than the rotation angle around the x-axis of the second camera parameter, and selects the second learning model when the rotation angle around the x-axis of the second camera parameter is closer to the depression angle than the rotation angle around the x-axis of the first camera parameter.
[0058] Furthermore, the behavior detection device 40 may have the functions of the learning device 10 according to embodiment 1. In this case, the behavior detection device 40 further includes a means for acquiring three-dimensional skeletal coordinates (not shown) (such as a stereo camera, a TOF camera, LiDAR, or a deep learning device that estimates three-dimensional skeletal coordinates from two-dimensional images) and a device for acquiring correct answers to behaviors linked to the three-dimensional skeletal coordinates (for example, a device that accepts manual input by an operator), and the learning device 10 updates multiple learning models (for example, a first learning model, a second learning model, and a third learning model).
[0059] In the behavior detection device 40 serving as an inference device, first, a two-dimensional image acquisition unit 41 acquires a two-dimensional image D41 generated by camera photography, and a two-dimensional skeletal coordinate calculation unit 42 calculates two-dimensional skeletal coordinates for inference D42 from the two-dimensional image D41. Known methods that can be used by the two-dimensional skeletal coordinate calculation unit 42 include, for example, OpenPose, MoveNet, PoseNet, BlazePose, and OpenVino.
[0060] The two-dimensional skeleton coordinates for inference obtained in this way are used as the rotation angle r around the x-axis when the depression angle is used to construct the learning model. xThe data is input to a learning model that is closest to the target (i.e., the selected learning model), and the behavior (or behavior likelihood) D44 is output. Examples of behavior likelihoods are standing, walking, falling, and lying. The two-dimensional skeletal coordinates for inference input to the learning model are those for a single time, but two-dimensional skeletal coordinates for inference for multiple times may be input to the learning model to improve accuracy.
[0061] At this time, the behavior detection unit 44 selects one learning model from multiple learning models based on the depression angle of the camera 50. There are several specific methods for the camera information acquisition unit 43. The simplest method is to manually input the depression angle when installing the camera. Another method is to use an acceleration sensor fixed to the camera to acquire the direction of gravitational acceleration and calculate the depression angle. This is effective as a means of solving the errors and complexity of manual input. Another method is to estimate the depression angle from the image features of the acquired two-dimensional image, specifically, the inclination of a horizontal line on the image. This method is effective in reducing costs because it does not require additional equipment such as an acceleration sensor.
[0062] r when constructing the first and second learning models x Since the boundary was 40°, it is advisable to use different models depending on whether the depression angle is above or below 40°.
[0063] When the behavior detection device 40 according to embodiment 2 is used, the amount of data acquired remains constant even if the number of two-dimensional image acquisition units 41 is increased, and all two-dimensional image acquisition units 41 can be configured simply by performing projection calculations of two-dimensional skeletal coordinates for inference from three-dimensional skeletal coordinates for inference on a computer.
[0064] In the above explanation, an example was shown in which one learning model was selected, but the two-dimensional skeletal coordinates for inference may be input into multiple learning models, and the multiple resulting behavioral likelihoods may be used (for example, by weighted addition of the multiple behavioral likelihoods) to obtain the final behavioral likelihood.
[0065] The learning model may also be updated during inference. Specifically, even if the inference output indicates a fall, if the user cancels the alarm because they believe the fall is a false alarm, the learning model can be updated by labeling the input 2D skeleton coordinates for inference as an action other than a fall, thereby preventing false alarms in similar situations.
[0066] 19 is a flowchart schematically illustrating the operation of the behavior detection device 40 according to Embodiment 2. First, the behavior detection device 40 acquires a two-dimensional image D41 generated by capturing an image of a target with the camera 50 (step S21).
[0067] Next, the behavior detection device 40 calculates two-dimensional skeletal coordinates for inference D42 indicating the coordinates of the skeleton of the target from the two-dimensional image D41 (step S22).
[0068] Next, the behavior detection device 40 acquires (or calculates) the depression angle of the camera 50 (step S23). If the depression angle is equal to or less than a predetermined threshold (YES in step S24), the behavior detection device 40 performs inference using the first learning model M1 (step S25). If the depression angle is greater than the predetermined threshold (NO in step S24), the behavior detection device 40 performs inference using the second learning model M2 (step S26).
[0069] The behavior detection device 40 outputs the result of the inference (behavior or behavior likelihood) (step S27).
[0070] 20 is a diagram illustrating an example of the hardware configuration of the behavior detection device 40 according to embodiment 2. The behavior detection device 40 has a processor 201 such as a CPU, a memory 202 as a storage device such as a RAM, a storage device 203 as a non-volatile storage device such as an HDD or SSD, and an interface 204. These components may be configured using dedicated processing circuits.
[0071] The processor 201 can execute the behavior detection program according to the second embodiment. The behavior detection program is provided by being recorded on a recording medium such as an SD memory card or a USB memory card, or by being downloaded via a network. The hardware configuration shown in Fig. 20 is an example, and various modifications to the hardware configuration are possible.
[0072] As described above, by using the behavior detection device 40, the behavior detection method, and the behavior detection program according to the second embodiment, it is possible to accurately detect a behavior using a learning model selected from a plurality of learning models. [Explanation of symbols]
[0073] 10 Learning device, 11 Three-dimensional data acquisition unit, 12 Two-dimensional projection transformation unit, 13 First model generation unit, 14 Second model generation unit, 20 Virtual camera, 22a First two-dimensional surface, 22b Second two-dimensional surface, 40 Behavior detection device (inference device), 41 Two-dimensional image acquisition unit, 42 Two-dimensional skeletal coordinate calculation unit (two-dimensional image-derived two-dimensional skeletal coordinate calculation unit), 43 Camera information acquisition unit (two-dimensional image acquisition unit information acquisition unit), 44 Behavior detection unit, 50 Camera, D10 Three-dimensional data for learning, D11 Three-dimensional skeletal coordinates for learning, D12 Teacher data, D21 First two-dimensional data for learning, D22 Second two-dimensional data for learning, D41 Two-dimensional image, D42 Two-dimensional skeletal coordinates for inference, P1 First camera parameter, P2 Second camera parameter, M1 First learning model, M2 second learning model, M3 third learning model.
Claims
1. A learning device that receives learning three-dimensional data including learning three-dimensional skeletal coordinates including three-dimensional coordinates of a skeleton of a target and teacher data consisting of correct answers for actions linked to the skeleton, and generates a learning model, a two-dimensional projection transformation unit that projects and transforms the training three-dimensional skeleton coordinates onto a first two-dimensional plane according to predetermined first camera parameters to generate first training two-dimensional data including first training two-dimensional skeleton coordinates and the teacher data, and that projects and transforms the training three-dimensional skeleton coordinates onto a second two-dimensional plane according to predetermined second camera parameters to generate second training two-dimensional data including second training two-dimensional skeleton coordinates and the teacher data; a first model generation unit that generates a first learning model for inferring an action associated with the first two-dimensional skeletal coordinates for inference from the first two-dimensional skeletal coordinates for inference using the first two-dimensional learning data; a second model generation unit that generates a second learning model for inferring an action associated with the second two-dimensional skeletal coordinates for inference from the second two-dimensional skeletal coordinates for inference using the second two-dimensional learning data; and The two-dimensional projection conversion unit generates training three-dimensional skeletal coordinates that have been interpolated or thinned out from original data of the training three-dimensional skeletal coordinates at a plurality of times arranged in a chronological order, changes the number of the training three-dimensional skeletal coordinates, obtains training three-dimensional skeletal data at a speed different from that of the original data, and further generates the first training two-dimensional data and the second training two-dimensional data based on the training three-dimensional skeletal data at the different speed. A learning device characterized by:
2. Each of the first camera parameters and the second camera parameters includes rotation angles around three axes consisting of the x-axis, the y-axis, and the z-axis, where the vertical coordinate axis is the y-axis, the coordinate axis on the line obtained by projecting a line in the projection direction used when generating the first learning two-dimensional skeleton coordinates and the second learning two-dimensional skeleton coordinates onto a horizontal plane that is a plane perpendicular to the y-axis is the z-axis, and the coordinate axis perpendicular to both the y-axis and the z-axis is the x-axis.
2. The learning device according to claim 1 .
3. the first camera parameters include coordinates in a world coordinate system of a first virtual camera that are set when the first training two-dimensional skeleton coordinates are generated using a perspective projection method, and a focal length of the first virtual camera; The second camera parameters include coordinates in a world coordinate system of a second virtual camera that are set when the second training two-dimensional skeleton coordinates are generated using a perspective projection method, and a focal length of the second virtual camera.
2. The learning device according to claim 1 .
4. Arbitrary values are set as the rotation angle around the y-axis according to the first camera parameter and the rotation angle around the y-axis according to the second camera parameter.
3. The learning device according to claim 2.
5. The two-dimensional projection transformation unit a rotation angle around the z-axis when generating the first training two-dimensional data by projectively transforming the training three-dimensional skeleton coordinates is set within a first range from a predetermined first angle of −90° or more and 0° or less to a predetermined second angle of 0° or more and 90° or less; The rotation angle around the z-axis when generating the second two-dimensional training data by projectively transforming the three-dimensional training skeleton coordinates is set within a second range from a predetermined third angle of −90° or more and 0° or less to a predetermined fourth angle of 0° or more and 90° or less.
3. The learning device according to claim 2.
6. The two-dimensional projection transformation unit a rotation angle around the x-axis when generating the first training two-dimensional data by projectively transforming the training three-dimensional skeleton coordinates is set to a third range that is greater than a predetermined fifth angle within a range of 0° to 90° and less than a predetermined sixth angle within a range of 0° to 90°; The rotation angle around the x-axis when generating the second training two-dimensional data by projectively transforming the training three-dimensional skeleton coordinates is set within a fourth range that is greater than a predetermined seventh angle within a range from 0° to 90° and less than a predetermined eighth angle within a range from 0° to 90°.
3. The learning device according to claim 2.
7. The two-dimensional projection transformation unit generates the first two-dimensional training data and the second two-dimensional training data from three-dimensional skeleton coordinates obtained by adding a random number within a predetermined range to the three-dimensional training skeleton coordinates.
2. The learning device according to claim 1 .
8. The two-dimensional projection transformation unit generates the first training two-dimensional data and the second training two-dimensional data from three-dimensional skeleton coordinates obtained by multiplying a difference vector between a first skeleton point and a second skeleton point determined by the training three-dimensional skeleton coordinates by a scalar value.
2. The learning device according to claim 1 .
9. The behavior includes one or more of standing upright, lying down, falling, and moving.
2. The learning device according to claim 1 .
10. The training 3D skeletal coordinates were obtained by motion capture or a stereo camera.
3. The learning device according to claim 2.
11. A learning method performed by a learning device, comprising: receiving three-dimensional learning data including three-dimensional learning skeletal coordinates including three-dimensional coordinates of a skeleton of a target and training data including correct answers for actions associated with the skeleton; a step of projecting and transforming the training three-dimensional skeleton coordinates onto a first two-dimensional plane in accordance with predetermined first camera parameters to generate first training two-dimensional data including first training two-dimensional skeleton coordinates and the teacher data, and projecting and transforming the training three-dimensional skeleton coordinates onto a second two-dimensional plane in accordance with predetermined second camera parameters to generate second training two-dimensional data including second training two-dimensional skeleton coordinates and the teacher data; generating a first learning model for inferring an action associated with the first two-dimensional skeletal coordinates for inference from the first two-dimensional skeletal coordinates for inference using the first two-dimensional learning data; generating a second learning model for inferring an action associated with the second two-dimensional skeletal coordinates for inference from the second two-dimensional skeletal coordinates for inference using the second two-dimensional learning data; generating interpolated or thinned learning three-dimensional skeletal coordinates from original data of the plurality of learning three-dimensional skeletal coordinates at a plurality of times arranged in a time series, changing the number of the learning three-dimensional skeletal coordinates to obtain learning three-dimensional skeletal data at a speed different from that of the original data, and further generating the first learning two-dimensional data and the second learning two-dimensional data based on the learning three-dimensional skeletal data at the different speed; A learning method comprising:
12. receiving three-dimensional learning data including three-dimensional learning skeletal coordinates including three-dimensional coordinates of a skeleton of a target and training data including correct answers for actions associated with the skeleton; a step of projecting and transforming the training three-dimensional skeleton coordinates onto a first two-dimensional plane in accordance with predetermined first camera parameters to generate first training two-dimensional data including first training two-dimensional skeleton coordinates and the teacher data, and projecting and transforming the training three-dimensional skeleton coordinates onto a second two-dimensional plane in accordance with predetermined second camera parameters to generate second training two-dimensional data including second training two-dimensional skeleton coordinates and the teacher data; generating a first learning model for inferring an action associated with the first two-dimensional skeletal coordinates for inference from the first two-dimensional skeletal coordinates for inference using the first two-dimensional learning data; generating a second learning model for inferring an action associated with the second two-dimensional skeletal coordinates for inference from the second two-dimensional skeletal coordinates for inference using the second two-dimensional learning data; generating interpolated or thinned learning three-dimensional skeletal coordinates from original data of the plurality of learning three-dimensional skeletal coordinates at a plurality of times arranged in a time series, changing the number of the learning three-dimensional skeletal coordinates to obtain learning three-dimensional skeletal data at a speed different from that of the original data, and further generating the first learning two-dimensional data and the second learning two-dimensional data based on the learning three-dimensional skeletal data at the different speed; A learning program characterized by causing a computer to execute the above.
13. A behavior detection device that detects behavior of a target included in a two-dimensional image acquired by a two-dimensional image acquisition unit, a two-dimensional skeletal coordinate calculation unit that calculates two-dimensional skeletal coordinates for two-dimensional image-derived inference indicating skeletal coordinates of the object from the two-dimensional image; a behavior detection unit that selects one or more learning models from the first learning model and the second learning model generated by the learning device according to any one of claims 1 to 10, and detects the behavior from the two-dimensional skeletal coordinates for two-dimensional image-derived inference using the one or more selected learning models; A behavior detection device comprising:
14. A behavior detection device that detects behavior of a target included in a two-dimensional image acquired by a two-dimensional image acquisition unit, a two-dimensional skeletal coordinate calculation unit that calculates two-dimensional skeletal coordinates for two-dimensional image-derived inference indicating skeletal coordinates of the object from the two-dimensional image; A learning device that receives three-dimensional learning data including three-dimensional learning skeletal coordinates including three-dimensional coordinates of the skeleton of the target and teacher data consisting of correct answers for actions linked to the skeleton, and generates a learning model, the learning device comprising: a two-dimensional projection transformation unit that projects and transforms the three-dimensional learning skeletal coordinates onto a first two-dimensional plane according to predetermined first camera parameters to generate first two-dimensional learning data including first two-dimensional learning skeletal coordinates and the teacher data, and projects and transforms the three-dimensional learning skeletal coordinates onto a second two-dimensional plane according to predetermined second camera parameters to generate second two-dimensional learning data including second two-dimensional learning skeletal coordinates and the teacher data; a behavior detection unit that selects one or more learning models from the first learning model and the second learning model generated by a learning device having: a first model generation unit that uses training two-dimensional data to generate a first learning model for inferring behavior associated with the first inference two-dimensional skeletal coordinates from the first inference two-dimensional skeletal coordinates; and a second model generation unit that uses the second training two-dimensional data to generate a second learning model for inferring behavior associated with the second inference two-dimensional skeletal coordinates from the second inference two-dimensional skeletal coordinates; and that detects the behavior from the two-dimensional image-derived inference two-dimensional skeletal coordinates using the one or more selected learning models; a two-dimensional image acquisition unit information acquisition unit that acquires or calculates a depression angle of an optical axis when the two-dimensional image is captured by the two-dimensional image acquisition unit; and The behavior detection unit selects the one or more learning models based on the depression angle. A behavior detection device characterized by:
15. The two-dimensional image acquisition unit information acquisition unit acquires the depression angle of the two-dimensional image acquisition unit from a storage device that stores the depression angle of the optical axis when the two-dimensional image is captured. The behavior detection device according to claim 14 .
16. The two-dimensional image acquisition unit information acquisition unit acquires the depression angle based on the output of an acceleration sensor fixed to the two-dimensional image acquisition unit. The behavior detection device according to claim 14 .
17. A behavior detection device that detects behavior of a target included in a two-dimensional image acquired by a two-dimensional image acquisition unit, a two-dimensional image-derived two-dimensional skeletal coordinate calculation unit that calculates two-dimensional skeletal coordinates for two-dimensional image-derived inference indicating the coordinates of the skeleton of the object from the two-dimensional image; a first model generation unit that uses the first two-dimensional data for learning to generate a first learning model; a first model generation unit that uses the first two-dimensional data for learning to generate a first learning model; a second model generation unit that uses the second two-dimensional data for learning to generate a first learning model; a first model generation unit that uses the first two-dimensional data for learning to generate a first learning model for inferring an action associated with the first two-dimensional skeletal coordinates for inference from the first two-dimensional skeletal coordinates for inference; and a second model generation unit that uses the second two-dimensional data for learning to generate a first learning model. a second model generation unit that generates a second learning model for inferring an action associated with the second two-dimensional skeletal coordinates for inference from the second two-dimensional skeletal coordinates for inference using the above-mentioned second model generation unit, wherein each of the first camera parameters and the second camera parameters is defined as follows: a vertical coordinate axis is defined as a y-axis, a coordinate axis on a line obtained by projecting a straight line in a projection direction used when generating the first two-dimensional skeletal coordinates for inference onto a horizontal plane that is a plane perpendicular to the y-axis is defined as a z-axis, and a coordinate axis perpendicular to both the y-axis and the z-axis is defined as an x-axis; and a behavior detection unit that selects one or more learning models from the first learning model and the second learning model generated by a learning device, the learning model including rotation angles around three axes consisting of the x-axis, the y-axis, and the z-axis, and detects the action from the two-dimensional skeletal coordinates for inference using the one or more selected learning models; a two-dimensional image acquisition unit information acquisition unit that acquires or calculates a depression angle of an optical axis when the two-dimensional image is captured by the two-dimensional image acquisition unit; and The behavior detection unit selects the first learning model when the rotation angle around the x-axis of the first camera parameter is closer to the rotation angle around the x-axis of the second camera parameter, and selects the second learning model when the rotation angle around the x-axis of the second camera parameter is closer to the rotation angle around the x-axis of the first camera parameter. A behavior detection device characterized by:
18. The learning device updating the first learning model using the two-dimensional skeleton coordinates for two-dimensional image-derived inference as the first two-dimensional skeleton coordinates for learning; The second learning model is updated using the two-dimensional skeleton coordinates for two-dimensional image-derived inference as the second two-dimensional skeleton coordinates for learning. The behavior detection device according to claim 14 .
19. A behavior detection method implemented by a behavior detection device that detects behavior of a target included in a two-dimensional image acquired by a two-dimensional image acquisition unit, comprising: calculating two-dimensional skeletal coordinates for inference indicating coordinates of the skeleton of the object from the two-dimensional image; receiving three-dimensional learning data including three-dimensional learning skeletal coordinates including three-dimensional coordinates of the skeleton of the subject and teacher data consisting of correct answers to actions linked to the skeleton; projecting and transforming the three-dimensional learning skeletal coordinates onto a first two-dimensional plane according to predetermined first camera parameters to generate first two-dimensional learning data including first two-dimensional learning skeletal coordinates and the teacher data; projecting and transforming the three-dimensional learning skeletal coordinates onto a second two-dimensional plane according to predetermined second camera parameters to generate second two-dimensional learning data including second two-dimensional learning skeletal coordinates and the teacher data; a step of selecting one or more of the first learning model and the second learning model generated by a learning method including a step of using two-dimensional training data to generate a first learning model for inferring an action associated with the first two-dimensional training skeletal coordinates from the first two-dimensional training skeletal coordinates, and a step of using the second two-dimensional training data to generate a second learning model for inferring an action associated with the second two-dimensional training skeletal coordinates from the second two-dimensional training skeletal coordinates, and detecting the action from the two-dimensional training skeletal coordinates using the one or more selected learning models; acquiring or calculating a depression angle of an optical axis when the two-dimensional image is captured by the two-dimensional image acquisition unit; selecting the one or more learning models based on the depression angle; A behavior detection method comprising:
20. a computer that detects an action performed by a target included in a two-dimensional image acquired by the two-dimensional image acquisition unit; calculating two-dimensional skeletal coordinates for inference indicating coordinates of the skeleton of the object from the two-dimensional image; receiving three-dimensional learning data including three-dimensional learning skeletal coordinates including three-dimensional coordinates of the skeleton of the subject and teacher data consisting of correct answers to actions linked to the skeleton; projecting and transforming the three-dimensional learning skeletal coordinates onto a first two-dimensional plane according to predetermined first camera parameters to generate first two-dimensional learning data including first two-dimensional learning skeletal coordinates and the teacher data; projecting and transforming the three-dimensional learning skeletal coordinates onto a second two-dimensional plane according to predetermined second camera parameters to generate second two-dimensional learning data including second two-dimensional learning skeletal coordinates and the teacher data; a step of selecting one or more of the first learning model and the second learning model generated by a learning program that causes a computer to execute the steps of: generating a first learning model for inferring an action associated with the first two-dimensional skeletal coordinates for inference from the first two-dimensional skeletal coordinates for inference using the first two-dimensional skeletal coordinates for inference using the second two-dimensional skeletal data; and generating a second learning model for inferring an action associated with the second two-dimensional skeletal coordinates for inference using the second two-dimensional skeletal coordinates for inference using the second two-dimensional skeletal data; and detecting the action from the two-dimensional skeletal coordinates for inference using the one or more selected learning models. acquiring or calculating a depression angle of an optical axis when the two-dimensional image is captured by the two-dimensional image acquisition unit; selecting the one or more learning models based on the depression angle; A behavior detection program characterized by executing the above.
Citation Information
Patent Citations
Electronic device and control method thereof
CN112823354A
Information processing device, information processing method and program
JP2018120283A
Object recognition device
JP2019191908A
Information processor and information processing method
JP2021082049A
Detection device, detection method, and learning model production method
JP2022139507A