Skeleton estimation device, skeleton estimation method, and program

The system addresses occlusion in human skeleton recognition by using shadow information and occupancy maps to reconstruct hidden body parts, ensuring accurate skeletal estimation despite occlusions.

JP2026013256APending Publication Date: 2026-01-28DENSO CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024113571
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2026-01-28

AI Technical Summary

Technical Problem

Existing human skeleton recognition systems face recognition failures due to occlusion, where parts of the body are hidden behind objects, and long-term accuracy is difficult to achieve with prediction methods alone.

Method used

The system estimates skeletal information by recognizing shadow information and combining it with standard skeleton recognition, using multiple cameras and light sources to reconstruct hidden body parts from shadow shapes and occupancy maps, complementing skeleton recognition results.

Benefits of technology

Accurately estimates skeletal information even when occlusion occurs, maintaining long-term accuracy by utilizing shadow information to track individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026013256000001_ABST
    Figure 2026013256000001_ABST
Patent Text Reader

Abstract

To accurately estimate skeleton information of a moving body even when occlusion occurs.SOLUTION: A skeleton estimation device includes a skeleton recognition unit that recognizes skeleton information of a moving object based on image data captured by a camera, a shadow recognition unit that recognizes shadow information of the moving object based on the image data captured by the camera, and a skeleton estimation unit that estimates the skeleton information of the moving object based on the recognized skeleton information of the moving object and the recognized shadow information of the moving object.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a skeleton estimation device, a skeleton estimation method, and a program. [Background technology]

[0002] An occlusion-adaptive tracking device that tracks the movement of a moving object in a video image of the moving object is known (see, for example, Patent Document 1). The tracking device includes a tracking means that extracts the moving object from each frame of video data in which the moving object is captured and tracks the movement of the moving object, a judgment frame setting means that sets a judgment frame that includes the moving object in each frame, an occlusion occurrence determination means that determines that an occlusion has occurred based on whether the judgment frame of the current frame overlaps with multiple judgment frames of previous frames, and a correspondence means that, upon occlusion release, matches the moving object immediately before the occlusion occurred with the moving object immediately after the occlusion is released based on feature points. During occlusion, which is difficult to analyze, a trajectory connecting the player's positions before and after the occlusion is obtained without performing low-reliability analysis.

[0003] An occlusion analysis system is known that improves the accuracy of a behavior prediction model by generating occlusion parameters that can inform the mathematical model to generate more accurate predictions (Patent Document 2). This occlusion analysis system trains and applies a model to generate occlusion parameters, such as how a person is occluded, the occlusion percentage, and the occlusion type. The behavior prediction system can input the occlusion parameters, as well as other parameters related to human activity, into a second mathematical model for behavior prediction. The second machine learning model is a higher-level model trained to output a prediction of a person's likely future behavior and a confidence level associated with the prediction. The confidence level is determined at least in part based on the occlusion parameters. The behavior prediction system can output the prediction and the confidence level to a control system and other intelligent video analysis systems that generate commands associated with the vehicle. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-66917 [Patent Document 2] Special Publication No. 2023-552105 Summary of the Invention [Problem to be solved by the invention]

[0005] In human skeleton recognition, recognition failures occur due to occlusion, where a part of the body is hidden behind an object and cannot be directly observed. It is difficult to completely prevent occlusion, even by increasing the number of cameras or optimizing their installation positions.

[0006] In the prior art, many prediction methods based on recent movements or prior information have been devised, but it is difficult to achieve long-term accuracy with prediction alone.

[0007] The present disclosure has been made in consideration of the above-mentioned problems, and aims to accurately estimate skeletal information of a moving object even when occlusion occurs. [Means for solving the problem]

[0008] In order to achieve the above object, the skeleton estimation device according to the first aspect includes a skeleton recognition unit that recognizes skeletal information of a moving body based on image data captured by a camera, a shadow recognition unit that recognizes shadow information of the moving body based on image data captured by the camera, and a skeleton estimation unit that estimates skeletal information of the moving body based on the skeletal information of the recognized moving body and the shadow information of the recognized moving body.

[0009] The skeleton estimation method according to the second aspect is a method executed by a computer to recognize skeletal information of a moving body based on image data captured by a camera, recognize shadow information of the moving body based on the image data captured by the camera, and estimate the skeletal information of the moving body based on the recognized skeletal information of the moving body and the shadow information of the recognized moving body.

[0010] Furthermore, a skeleton estimation program according to a third aspect is a program that causes at least one processor to recognize skeletal information of a moving body based on image data captured by a camera, recognize shadow information of the moving body based on image data captured by the camera, and estimate the skeletal information of the moving body based on the recognized skeletal information of the moving body and the recognized shadow information of the moving body. [Effects of the Invention]

[0011] According to the skeleton estimation device, method, and program of the present disclosure, it is possible to accurately estimate skeleton information of a moving object even when occlusion occurs. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram illustrating a configuration of a skeleton estimation system according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a block diagram illustrating a hardware configuration of a skeleton estimation device according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a functional block diagram of a skeleton estimation device according to an embodiment of the present disclosure. [Figure 4] FIG. 10 is a diagram showing an example in which a target area is irradiated with light from multiple angles and an image is acquired in which multiple shadows are observed. [Figure 5] 10 is a timing chart showing a process of acquiring images captured using a plurality of cameras with a plurality of light sources all turned on. [Figure 6] 10 is a timing chart showing a process of acquiring images captured using a plurality of cameras with all of the plurality of light sources turned off. [Figure 7] 10 is a timing chart showing a process of turning on a plurality of light sources one by one and acquiring images captured using a plurality of cameras. [Figure 8] FIG. 10 is a diagram illustrating an example of skeletal information. [Figure 9] FIG. 10 is a diagram illustrating an example of skeleton information when occlusion occurs. [Figure 10] 10A and 10B are diagrams illustrating examples of input and output of a class map generation unit; [Figure 11] FIG. 10 is a diagram illustrating an example of a class map. [Figure 12] FIG. 10A is a diagram showing an example of a recognized region, and FIG. 10B is a diagram showing an example of a processed region. [Figure 13] FIG. 1 shows an example of image data captured by multiple cameras when a single light source is on. [Figure 14] FIG. 10 is a diagram showing how the shape of a shadow is extracted from image data. [Figure 15] FIG. 10 is a diagram illustrating an example of a reliability score. [Figure 16] FIG. 10 is a diagram showing an example of image data captured when all lights are off. [Figure 17] FIG. 10 is a diagram illustrating coordinate transformation from a pixel coordinate system to a spatial coordinate system. [Figure 18] 10A and 10B are diagrams illustrating examples of input and output of an occupancy map generating unit. [Figure 19] FIG. 10 is a diagram illustrating an example of an occupancy map. [Figure 20] FIG. 10 is a diagram showing an example of reliability scores for a plurality of image data. [Figure 21] 10A and 10B are diagrams for explaining a method for calculating an area where an object may exist from a light source position and a shadow. [Figure 22] FIG. 10 is a diagram for explaining a method of adding up the results of calculating an area where an object may exist. [Figure 23] FIG. 10 is a diagram illustrating an example of an occupancy map. [Figure 24] FIG. 10 is a diagram illustrating an example of input and output of a skeleton estimator. [Figure 25]FIG. 1 is a functional block diagram of a skeleton estimation device according to an embodiment of the present disclosure. [Figure 26] This figure explains a method of using a simulator to create a set of image data that can be acquired by each camera when photographing a person in space and true values ​​of skeletal information, and use this as training data. [Figure 27] 10 is a flowchart illustrating an example of a learning process according to an embodiment of the present disclosure. [Figure 28] 10 is a flowchart illustrating an example of a skeleton estimation process according to an embodiment of the present disclosure. [Figure 29] FIG. 2 is a diagram illustrating an example of input and output in each unit of the skeleton estimation device. [Figure 30] 10 is a flowchart illustrating an example of a process for generating an occupancy map. [Figure 31] 10 is a flowchart illustrating an example of a process for calculating an occupancy probability that an object exists. [Figure 32] FIG. 10 is a diagram for explaining a method for calculating an occupancy probability. [Figure 33] 10 is a flowchart illustrating an example of a process for generating an occupancy map from occupancy probabilities. [Figure 34] FIG. 10A shows an example of voxels in an occupancy map, FIG. 10B shows an example of occupancy probabilities, and FIG. 10C shows an example of integrated occupancy probabilities. [Figure 35] 10 is a flowchart illustrating an example of a process for estimating skeletal information. DETAILED DESCRIPTION OF THE INVENTION

[0013] Before describing the details of the embodiments of the present disclosure, an overview of the embodiments of the present disclosure will be described.

[0014] <Summary of Embodiments of the Present Disclosure> In an embodiment of the present disclosure, focusing on the fact that the shadow of a person hidden behind an object can be directly observed, the person (body part) hidden behind the object is estimated from the position and shape of the shadow. As a result, the skeletal information of the person hidden behind the object is reconstructed from the shape of the shadow and tracked.

[0015] Specifically, the position of a person hidden behind an obstacle is estimated from the shape of the shadow cast by multiple light sources, making the method robust against long-term occlusion.

[0016] In addition, from the shape of the shadow, an occupancy map is estimated that represents the space where the object is located for the area hidden by the object. In this process, the system learns in advance what kind of shadow will be created depending on the skeletal information of the person. In addition, the accuracy of the occupancy map estimation is improved by controlling the lighting patterns of multiple light sources to create various shadows.

[0017] Furthermore, the system combines the estimated occupancy map with the results of standard skeleton recognition to estimate human skeleton information so as to complement the results of skeleton recognition. Specifically, the target space is divided into voxels, and the estimated occupancy map and the results of skeleton recognition are stored in each voxel. For voxels with low confidence scores for skeleton recognition, the system complements the skeleton information using the estimated occupancy map.

[0018] In this way, by using shadows to estimate the skeletal information of a person hidden by occlusion, accuracy does not decrease over time as long as the shadow is visible.In addition, shadow information can be used to track people.

[0019] <System configuration> As shown in FIG. 1, the skeleton estimation system 10 according to this embodiment includes a plurality of cameras 20 provided for a target area to be photographed in real space, a plurality of light sources 22 provided to illuminate the target area to be photographed, a light source control device 24 for controlling the plurality of light sources 22, and a skeleton estimation device 100.

[0020] In this embodiment, a case where a person is used as an example of a moving object to be tracked will be described. Although Fig. 1 shows an example where a plurality of cameras 20 and a plurality of light sources 22 are provided, there may be only one camera 20 and one light source 22. Furthermore, the moving object to be tracked may be something other than a person, such as an animal or a robot.

[0021] The multiple cameras 20 are connected to a skeleton estimation device 100. The multiple light sources 22 are connected to a light source control device 24. The cameras 20, the light sources 22, the light source control device 24, and the skeleton estimation device 100 are connected via a network N such as a LAN (Local Area Network) or the Internet.

[0022] The multiple cameras 20 are fixed cameras installed relative to the target area in real space, and are visible light cameras. The cameras 20 have a function of capturing, for example, an RGBD image, which is an image that combines RGB color information with distance information (depth information) as seen from the camera 20. The multiple cameras 20 are installed so as to capture the target area from different directions.

[0023] When a light source that emits infrared rays is used as the light source 22, an infrared camera that can detect the wavelength of infrared rays may be used as the camera 20.

[0024] The multiple light sources 22 are provided to irradiate the area to be photographed in real space with light, and each light source 22 is controlled to be turned on and off by a light source control device 24. The multiple light sources 22 are installed to irradiate the area to be photographed by the camera with light from different directions. The light sources 22 are fixed lights that irradiate light with an intensity sufficient to clearly observe the shadow of a person. The light sources 22 are light sources that irradiate visible light or infrared light.

[0025] The light source control device 24 is a computer having a CPU (Central Processing Unit), ROM (Read Only Memory), and RAM (Random Access Memory), and controls the on / off of each light source 22 according to instructions from the skeleton estimation device 100.

[0026] Fig. 2 is a block diagram showing the hardware configuration of a skeleton estimation device 100 according to this embodiment. As shown in Fig. 2, the skeleton estimation device 100 includes a CPU 11, a ROM 12, a RAM 13, a storage 14, an input unit 15, a display unit 16, and a communication interface (I / F) 17. Each component is connected to each other via a bus 19 so as to be able to communicate with each other.

[0027] The CPU 11 is a central processing unit that executes various programs and controls each part. That is, the CPU 11 reads a program from the ROM 12 or the storage 14 and executes the program using the RAM 13 as a work area. The CPU 11 controls each of the above components and performs various arithmetic processing in accordance with the program stored in the ROM 12 or the storage 14. In this embodiment, a skeleton estimation program is stored in the ROM 12 or the storage 14.

[0028] The ROM 12 stores various programs and various data. The RAM 13 temporarily stores programs or data as a working area. The storage 14 is configured with a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive) and stores various programs including the operating system and various data.

[0029] The input unit 15 includes a pointing device such as a mouse and a keyboard, and is used to perform various inputs.

[0030] The display unit 16 is, for example, a liquid crystal display, and displays various information. The display unit 16 may function as the input unit 15 by adopting a touch panel system.

[0031] The communication interface 17 is an interface for communicating with other devices such as terminals, etc. For this communication, for example, a wired communication standard such as Ethernet (registered trademark) or FDDI, or a wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark) is used.

[0032] Next, a description will be given of each functional configuration of the skeleton estimation device 100. Fig. 3 is a block diagram showing the configuration of the skeleton estimation device 100 of this embodiment. Each functional configuration is realized by the CPU 11 reading out a skeleton estimation program stored in the ROM 12 or storage 14, expanding it in the RAM 13, and executing it.

[0033] As shown in FIG. 3, the skeleton estimation device 100 includes a data acquisition unit 30, a skeleton recognition unit 32, a class map generation unit 34, a shadow recognition unit 36, an occupancy map generation unit 38, and a skeleton estimation unit 40.

[0034] The data acquisition unit 30 acquires image data for each of the plurality of light sources 22 captured by the plurality of cameras 20 when the light source 22 is on, and image data for each of the plurality of light sources 22 captured by the plurality of cameras 20 when the plurality of light sources 22 is off. At this time, as shown in Fig. 4, it is possible to acquire images in which a plurality of shadows are observed by irradiating the imaging target area with light from a plurality of angles.

[0035] Specifically, first, the light source control device 24 turns on all of the light sources 22, and the data acquisition unit 30 acquires RGBD images captured by each of the cameras 20 in synchronization with this timing (see FIG. 5). FIG. 5 shows an example in which M RGBD images are collected by M cameras 20.

[0036] Furthermore, the light source control device 24 turns off all of the light sources 22, leaving only the light sources other than the light sources 22 on, and the data acquisition unit 30 acquires RGBD images captured using the multiple cameras 20 in synchronization with this timing (see FIG. 6). FIG. 6 shows an example in which M RGBD images are collected by M cameras 20.

[0037] Furthermore, a single light source 22 is turned on by the light source control device 24, and the data acquisition unit 30 acquires RGBD images captured using the multiple cameras 20 in synchronization with this timing (see FIG. 7). This process is repeated as many times as the number of light sources 22. FIG. 7 shows an example in which the process of collecting M RGBD images using M cameras 20 is repeated as many times as N, which is the number of light sources 22, to collect N×M RGBD images.

[0038] The skeleton recognition unit 32 recognizes skeleton information of a person based on image data captured by the camera 20, and outputs the results of tracking the recognized person.

[0039] Specifically, the skeleton recognition unit 32 extracts features of each person represented in the RGBD images from M RGBD images captured by the M cameras 20 when all of the light sources 22 are turned on. For example, the skeleton recognition unit 32 estimates the skeleton information of the person using a classifier described in Non-Patent Document 1.

[0040] Non-Patent Document 1: Bashirov, Renat, et al. "Real-time RGBD-based extended body pose estimation." Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision. 2021.

[0041] Here, the skeletal information represents a collection of three-dimensional position information of a person's joints, for example, as shown in FIG. 8. When occlusion occurs, some parts are not visible, and therefore, as shown in FIG. 9, there may be a loss of skeletal information. In addition, the person's ID is assigned to the skeletal information. The same ID is assigned to each person being tracked.

[0042] The class map generating unit 34 generates a class map that represents information about the body parts of a person for each voxel, based on the skeleton information of the person recognized in each of the plurality of image data.

[0043] Specifically, the skeletal information of a person recognized from each of multiple image data is used as input, and by reflecting each skeletal information in voxel information, a class map is generated that shows where in space each person's body part is located and with what probability (see Figure 10).

[0044] Here, a class map is a map that indicates the probability that an object exists in a space, and divides the entire space into voxels of a certain size, and for each voxel, it indicates which body part of which person exists within that voxel, as shown in Figure 11. Figure 11 shows an example in which, for a certain voxel (x, y, z), the probability that it is "empty" (nothing exists) is "50%", the probability that the arm of person with ID "1" exists is "30%", and the probability that the leg of person with ID "2" exists is "20%".

[0045] At this time, the recognized parts may be processed, such as by expanding them by a certain percentage (Fig. 12(a) and (b)). Fig. 12(a) shows an example in which the probability that the hand of a person with ID "1" exists in a certain voxel is "80%". Fig. 12(b) shows an example in which the probability that the hand of a person with ID "1" exists in the surrounding voxels is "10%".

[0046] The shadow recognition unit 36 ​​recognizes shadow information of a person based on image data captured by the camera 20.

[0047] Specifically, the shadow information of the moving object is recognized based on image data (see FIG. 13) captured by the multiple cameras 20 when a single light source 22 for each of the multiple light sources 22 is on, and image data captured by the multiple cameras 20 when the multiple light sources 22 are off. FIG. 13 shows an example in which an image is captured by M cameras 20 when a single light source is on.

[0048] Specifically, the shape of the shadow is extracted from each of the collected M × N RGBD images (see FIG. 14). A known technique may be used to extract the shape of the shadow, for example, the method described in Japanese Patent No. 6260620.

[0049] Then, for each pixel in the M × N RGBD images from which the shadow shape has been extracted, a confidence score is estimated, which indicates how reliable the extracted shadow is (Fig. 15). Fig. 15 shows an example in which the confidence score, Likelihood(u,v), of pixel (u,v) is 0.8.

[0050] The reliability score is estimated using the difference between an image when all lights are off and an image when a single light source 22 is on, as shown in Fig. 16. Pixels with large differences have a higher reliability score, while pixels with small differences have a lower reliability score. Additionally, pixels with high brightness when all lights are off are also considered to be affected by ambient light, and have a lower reliability score.

[0051] For brightness, convert RGBD to HSVD to obtain brightness (V value). For example, for a pixel (u, v), the confidence score Likelihood(u, v) is obtained according to the following formula:

[0052] JPEG2026013256000002.jpg19164

[0053] where VA(u,v) is the normalized brightness value at coordinates (u,v) in the image when all lights are off,

[0054] VB(u,v) is a normalized value of the brightness at the coordinates (u,v) in the image when a single light source 22 is on. Also, α and β are weighting variables that satisfy α+β=1.

[0055] According to the above formula, the reliability score Likelihood(u,v) of each pixel (u,v) is calculated for the M×N RGBD images from which the shadow shape has been extracted.

[0056] When the shadow is brighter, such as at the shadow's outline, the accuracy of shadow detection may be reduced, so the confidence score is lower in brighter areas. On the other hand, the confidence score is higher in areas where the shadow contrast is clear.

[0057] Then, coordinate conversion is performed from the pixel coordinate system to the spatial coordinate system using the following process (Fig. 17). Fig. 17 shows an example in which coordinates (u, y) in the pixel coordinate system are converted to coordinates (x, y, z) in the spatial coordinate system, and the confidence score Likelihood(u, v) is converted to the confidence score Likelihood(x, y, z).

[0058] The above processing is performed in a two-dimensional pixel coordinate system of the camera image (meaning that each piece of data is represented by u pixels in the vertical direction and v pixels in the horizontal direction). Therefore, coordinate conversion is performed to three-dimensional spatial coordinates (x, y, z) using the camera's internal parameters (camera focal length, center pixel of the image), external parameters (camera position), and depth information of the RGBD image. The coordinate conversion method can be the one described in Non-Patent Document 2.

[0059] Non-Patent Document 2: A New Model of RGB-D Camera Calibration Based (Chenyang Zhang et.al.)

[0060] For example, when coordinates (u,v) are given in the pixel coordinate system, coordinates (x,y,z) in the spatial coordinate system are given as follows:

[0061] First, we perform coordinate transformation from pixel coordinates (u,v) to a position (xc,yc,zc) in the camera coordinate system, where f is the focal length, C=(cx,cy) is the center pixel of the image, and d is the depth information.

[0062] JPEG2026013256000003.jpg5470

[0063] Then, coordinate transformation is performed from the position (xc, yc, zc) in the camera coordinate system to the position (x, y, y) in the spatial coordinate system. However, if the rotation matrix from the origin of the spatial coordinate system to the camera is R and the translation matrix is ​​T, the coordinate transformation is performed according to the following formula.

[0064] JPEG2026013256000004.jpg41122

[0065] By performing the above process for all pixels, a shadow reliability score of the shadow information in the spatial coordinate system is obtained.

[0066] In the above coordinate transformation, lens distortion correction and the like may be added.

[0067] The occupancy map generating unit 38 generates an occupancy map representing the space occupied by the person based on the shadow information of the recognized person.

[0068] Specifically, the reliability score of each pixel transformed into a spatial coordinate system is input, and an occupancy map indicating the probability that an object exists in each voxel is generated by the following process (FIG. 18). Here, the occupancy map is data storing the probability that an object exists in each voxel (FIG. 19). FIG. 18 shows an example in which the reliability score of each pixel estimated for M×N RGBD images from which shadow shapes have been extracted is input, and the occupancy map generator 38 generates an occupancy map in which the probability of no object being present is 75% and the probability of an object being present is 25% for some voxels, and the probability of no object being present is 20% and the probability of an object being present is 80% for other voxels. FIG. 19 shows an example in which the probability of no object being present and the probability of an object being present are stored for each voxel in the occupancy map.

[0069] First, based on the reliability score calculated by the shadow recognition unit 36, the probability that an object exists in each voxel is calculated.

[0070] For example, as shown in Fig. 20, the reliability score for each image data is input, and the region where an object may exist is calculated from the light source position and the shadow using the results of weighting the reliability score of the shadow at that position for all shadow regions (Fig. 21). Fig. 20 shows an example in which the reliability score Likelihood(x,y) of pixel (x,y) in the shadow portion represented by image data obtained when the third light source 22 is on is 0.8. Fig. 21 shows an example in which the reliability score Likelihood(x,y) of pixel (x,y) in the shadow portion represented by image data obtained when the third light source 22 is on is multiplied by a weight α, and the resulting value is assigned as an existence probability to each voxel on the line connecting pixel (x,y) and the third light source 22.

[0071] The results of calculating the areas where an object may exist are then added together (FIG. 22), normalized, and the final probability is calculated, and output as an occupancy map as shown in FIG.

[0072] Fig. 22 shows an example in which the existence probability of each voxel calculated from the likelihood(x,y) of each pixel (x,y) of image data obtained when each light source 22 is on is added up for each voxel. Fig. 23 shows an example in which the probability of no object being present is 75% and the probability of an object being present is 25% for some voxels in the occupancy map, and the probability of no object being present is 20% and the probability of an object being present is 80% for some other voxels.

[0073] The skeleton estimation unit 40 estimates skeleton information of a person based on a class map generated from skeleton information of the recognized person and an occupancy map generated from shadow information of the person.

[0074] Specifically, the skeleton estimation unit 40 estimates the skeletal information of a person based on a class map generated from the skeletal information of the person, time-series data of the results of tracking the person, and time-series data of the generated occupancy map, and outputs the results of tracking the recognized person. At this time, the skeletal information of the person is estimated so that missing skeletal information on the class map is interpolated.

[0075] More specifically, as shown in FIG. 24, time-series data of class maps and time-series data of occupancy maps from a certain period of time in the past are input to a pre-trained skeleton estimator, and skeleton information in which missing parts of the class maps are interpolated is obtained as the output of the skeleton estimator. The time-series data of each map can be implemented using a queue that can store a certain amount of each map data. Each class map from a certain period of time in the past is assigned an ID as a result of tracking a person. FIG. 24 shows an example in which a class map queue storing class maps from a certain period of time in the past and an occupancy map queue storing occupancy maps from a certain period of time in the past are input to the skeleton estimator, and the skeleton information of a person assigned an ID as a result of tracking a recognized person is estimated.

[0076] The skeleton estimator is trained in advance by a training device 150 shown in FIG.

[0077] Similar to the skeleton estimation device 100, the learning device 150 includes a CPU 11, a ROM 12, a RAM 13, a storage 14, an input unit 15, a display unit 16, and a communication interface (I / F) 17, as shown in FIG.

[0078] As shown in FIG. 25, the learning device 150 includes a calculation data acquisition unit 50, a skeleton recognition unit 52, a class map generation unit 54, a shadow recognition unit 56, an occupancy map generation unit 58, a skeleton estimation unit 60, and a learning unit 62.

[0079] First, a simulator that simulates a real environment is created, and the light source 22, the shape of the surrounding environment, the characteristics of the camera 20, the shape of a person, and other factors that affect the quality of the RGBD image are simulated.

[0080] As shown in Fig. 26, the calculation data acquisition unit 50 uses this simulator to create a set of RGBD images that can be acquired by each camera 20 when photographing a person existing in space and true values ​​of skeletal information, and uses this data as training data. Fig. 26 shows an example in which the simulator is used to acquire multiple image data using multiple light sources 22 and multiple cameras 20 as imitation targets, and to acquire skeletal information of a person to which an ID has been assigned as a result of tracking the person.

[0081] Similar to the skeleton recognition unit 32, the skeleton recognition unit 52 recognizes skeleton information of a person based on image data of the training data, and outputs the results of tracking the recognized person.

[0082] Similar to the class map generating unit 34, the class map generating unit 54 generates a class map that represents information about the body parts of a person for each voxel based on the skeleton information of the recognized person.

[0083] The shadow recognition unit 56, like the shadow recognition unit 36, recognizes shadow information of a person based on image data of the training data.

[0084] Similar to the occupancy map generator 38, the occupancy map generator 58 generates an occupancy map that represents the space occupied by the moving object based on the shadow information of the recognized moving object.

[0085] Similar to the skeleton estimation unit 40, the skeleton estimation unit 60 uses a skeleton estimator to estimate the skeletal information of a person based on a class map generated from the skeletal information of the recognized person and an occupancy map generated from the shadow information of the person.

[0086] The learning unit 62 performs supervised learning of the skeleton estimator so as to reduce the difference between the person's skeleton information obtained as an estimation result and the true value of the skeleton information in the training data (FIG. 26).

[0087] Next, the operation of the skeleton estimation system 10 will be described.

[0088] 27 is a flowchart showing the flow of the learning process by the learning device 150. The learning process is performed in advance by the CPU 11 reading out a learning program from the ROM 12 or storage 14, expanding it in the RAM 13, and executing it. The learning device 150 receives a learning instruction as input and performs the following processes.

[0089] In step S100, the CPU 11 causes the calculation data acquisition unit 50 to use this simulator to create a pair of RGBD images that can be acquired by each camera 20 when photographing a person existing in space, and true values ​​of skeletal information, as shown in Figure 26, and use this data as training data.

[0090] In step S102, the CPU 11 causes the skeleton recognition unit 52, similar to the skeleton recognition unit 32, to recognize skeleton information of a person based on the image data of the training data, and outputs the results of tracking the recognized person.

[0091] In step S104, the class map generating unit 54 of the CPU 11 generates a class map representing information about the body parts of a person for each voxel, similar to the class map generating unit 34, based on the skeleton information of the recognized person.

[0092] In step S106, the CPU 11 causes the shadow recognition unit 56 to recognize shadow information of a person based on the image data of the training data, similar to the shadow recognition unit 36.

[0093] In step S108, the CPU 11 causes the occupancy map generating unit 58 to generate an occupancy map representing the space occupied by the moving object, similar to the occupancy map generating unit 38, based on the shadow information of the recognized moving object.

[0094] In step S110, the CPU 11 causes the skeleton estimation unit 60, similar to the skeleton estimation unit 40, to estimate the person's skeleton information using a skeleton estimator based on a class map generated from the recognized person's skeleton information and an occupancy map generated from the person's shadow information.

[0095] In step S112, the CPU 11 causes the learning unit 62 to perform supervised learning of the skeleton estimator so as to reduce the difference between the person's skeleton information obtained as the estimation result and the true value of the skeleton information of the training data, and ends the learning process.

[0096] 28 is a flowchart showing the flow of skeleton estimation processing by skeleton estimation device 100. The CPU 11 reads out a skeleton estimation program from ROM 12 or storage 14, expands it in RAM 13, and executes it to perform the skeleton estimation processing. The skeleton estimation processing is an example of a skeleton estimation method. The skeleton estimation device 100 receives a skeleton estimation instruction as input and performs the following processing.

[0097] In step S120, the CPU 11, as the data acquisition unit 30, acquires image data for each of the plurality of light sources 22 captured by the plurality of cameras 20 when the light source 22 is on, and image data for each of the plurality of light sources 22 captured by the plurality of cameras 20 when the plurality of light sources 22 is off.

[0098] In step S122, the CPU 11, functioning as the skeleton recognition unit 32, recognizes skeleton information of a person based on the acquired image data and outputs the results of tracking the recognized person (see FIG. 29). FIG. 29 shows an example in which the skeleton recognition unit 32 obtains a recognition result of the person's skeleton information for each of the RGBD images captured from multiple viewpoints by multiple cameras 20.

[0099] In step S124, the CPU 11 functions as the class map generating unit 34 to generate a class map that represents information about the body parts of a person for each voxel, based on the skeleton information of the recognized person.

[0100] In step S126, the CPU 11, functioning as the shadow recognition unit 36, recognizes shadow information of a person based on the acquired image data (see FIG. 29). FIG. 29 shows an example in which the shadow recognition unit 36 ​​obtains a recognition result of shadow information of a person for each of the RGBD images captured from multiple viewpoints by multiple cameras 20.

[0101] In step S128, the CPU 11 functions as the occupancy map generating unit 38 to generate an occupancy map representing the space occupied by the moving object based on the shadow information of the recognized moving object.

[0102] In step S130, CPU 11 functions as skeleton estimation unit 40 to estimate the person's skeleton information using the skeleton estimator trained by learning device 150 based on the time series data of the class map generated from the recognized person's skeleton information and the time series data of the occupancy map generated from the person's shadow information, and then ends the skeleton estimation process (see FIG. 29). FIG. 29 shows an example in which skeleton estimation unit 40 obtains the skeletal information of each person using the time series data of the class map and the time series data of the occupancy map as input.

[0103] The above steps S108 and S128 are realized by the processing routine shown in FIG.

[0104] In step S140, the CPU 11, functioning as the occupancy map generator 38, 58, calculates the occupancy probability that an object exists in each region based on the reliability scores calculated in steps S106, S126. Specifically, for all shadow regions, the results of weighting the reliability scores of the shadow at that position are used. As an output, the existence probability that an object exists in each voxel is calculated for each combination of the nth light source 22 and the mth camera 20.

[0105] In step S142, the CPU 11, as the occupancy map generation unit 38, 58, adds up the existence probabilities of N light sources 22 and M cameras 20, totaling M x N, for each voxel in step S140, and finally calculates the probability that an object exists in space, thereby constructing an occupancy map.

[0106] The above step S140 is realized by the processing routine shown in Fig. 31. Steps S150 to S154, which will be described later, are repeated for each combination of N light sources 22 and M cameras 20.

[0107] In step S150, the CPU 11, functioning as the occupancy map generator 38, 58, determines an area connecting the installation position (xn, yn, zn) of the light source 22 and the shadow area. Here, an area consisting of pixels whose reliability score is greater than a threshold is determined to be the shadow area. For example, a cone connecting the installation position (xn, yn, zn) of the light source 22 and the outline of the shadow area is determined (see FIG. 32).

[0108] In step S152, the CPU 11, functioning as the occupancy map generator 38, 58, calculates the existence probability P(x, y, z) for the space (x, y, z) within the region determined in step S150. Specifically, the shadow reliability score Likelihood(x, y, z) calculated in steps S106, S126 is used. As a simple example, a weighted value of the average or maximum reliability score of the shadow region is used as the existence probability P(x, y, z) (see FIG. 32). FIG. 32 shows an example in which the average reliability score of the shadow region (0.8 or 0.7) is multiplied by a weight α (0.8α or 0.7α) to calculate the existence probability P(x, y, z) of the space within the region.

[0109] In step S154, the CPU 11, as the occupancy map generation unit 38, 58, assigns the existence probability P(x, y, z) calculated in step S152 above to the space within the cone connecting the installation position (xn, yn, zn) of the light source 22 and the outline of the shadow area (see Figure 32).

[0110] The above step S142 is realized by the processing routine shown in FIG.

[0111] Steps S160 to S162, which will be described later, are repeated for each voxel in the occupancy map as shown in FIG. 34(a).

[0112] Furthermore, step S160, which will be described later, is repeated for each combination of N light sources 22 and M cameras 20.

[0113] In step S160, the CPU 11, functioning as the occupancy map generator 38, 58, determines the existence probability within the target voxel from information on the nth light source 22 and the mth camera 20. For example, the existence probability of the target voxel is calculated from the existence probability of the shadow area that overlaps with the target voxel (see FIG. 34(b)). FIG. 34(b) shows an example in which the existence probability is 0 when the target voxel does not overlap with the shadow area. Also, it shows an example in which the existence probability is 0.1 when the target voxel partially overlaps with the shadow area. It also shows an example in which the existence probability is 0.3 when the target voxel entirely overlaps with the shadow area.

[0114] In step S162, the CPU 11, functioning as the occupancy map generator 38, 58, integrates the presence probabilities in the target voxel calculated from all light sources 22 and cameras 20. For example, the average of the presence probabilities calculated in step S160, a weighted average using preset weighting, a maximum value, a median value, or the like is used as the integrated result of the presence probability (see FIG. 34(c)). FIG. 34(c) shows an example in which 0.12, which is the average of presence probabilities of 0, 0.1, and 0.3, is calculated as the integrated result of the presence probability.

[0115] The above steps S110 and S130 are realized by the processing routine shown in FIG.

[0116] First, in step S170, the CPU 11, functioning as the skeleton estimation unit 40, 60, determines whether the number of data items in the class map queue storing the time-series data of the class maps generated in steps S104, S124 is N. Here, N is the size of the class map queue. If the number of data items in the class map queue is N, the process proceeds to step S172. On the other hand, if the number of data items in the class map queue is not N, the process proceeds to step S174.

[0117] In step S172, the CPU 11, functioning as the skeleton estimation units 40 and 60, extracts one piece of the oldest data from the class map queue, thereby making the number of pieces of data in the class map queue N-1.

[0118] In step S174, the CPU 11, functioning as the skeleton estimation units 40 and 60, stores the latest class maps generated in steps S104 and S124 in the class map queue, thereby creating time-series data of the class maps in the class map queue.

[0119] In step S176, CPU 11, functioning as skeleton estimation unit 40, 60, determines whether the number of data items in the occupation map queue storing the time-series data of the occupation map generated in steps S108, S128 is N. Here, N is the size of the occupation map queue. If the number of data items in the occupation map queue is N, the process proceeds to step S178. On the other hand, if the number of data items in the occupation map queue is not N, the process proceeds to step S180.

[0120] In step S178, the CPU 11, functioning as the skeleton estimation units 40 and 60, extracts one piece of the oldest data from the occupancy map queue, thereby making the number of pieces of data in the occupancy map queue N-1.

[0121] In step S180, the CPU 11, functioning as the skeleton estimation units 40 and 60, stores the latest occupancy maps generated in steps S108 and S128 in the occupancy map queue, thereby creating time-series data of the occupancy map in the occupancy map queue.

[0122] In step S182, the CPU 11, functioning as the skeleton estimation units 40 and 60, inputs the class map queue and the occupancy map queue to the skeleton estimator.

[0123] In step S184, the CPU 11 executes the skeleton estimator as the skeleton estimation units 40 and 60 to estimate skeleton information of a person.

[0124] As described above, the skeletal estimation device according to this embodiment recognizes skeletal information of a person based on image data captured by a camera, recognizes shadow information of the person based on image data captured by the camera, and estimates skeletal information of the person based on the recognized skeletal information of the person and the recognized shadow information of the person. This makes it possible to accurately estimate skeletal information of a person even when occlusion occurs.

[0125] Focusing on the fact that shadows can be directly observed, even for people hidden behind objects, it is possible to estimate the skeletal information of a person hiding behind an object from the position and shape of the shadow.In addition, it is possible to reconstruct the skeletal information of a person hiding behind an object from the shape of the shadow and track it.

[0126] Furthermore, by learning in advance what kind of shadow will be created when a person with certain skeletal information is present, it is possible to estimate the probability that an object exists in an area hidden by an object from the shape of the shadow and thereby estimate the skeletal information.

[0127] Furthermore, by controlling the on / off of multiple light sources and creating various shadows, it is possible to estimate a person's skeletal information with even greater accuracy.

[0128] <Modification> The various processes executed by the CPU after reading software (programs) in the above embodiments may be executed by various processors other than the CPU. Examples of such processors include programmable logic devices (PLDs) (such as field-programmable gate arrays (FPGAs)) whose circuit configuration can be changed after fabrication, and dedicated electrical circuits such as application-specific integrated circuits (ASICs) that are processors with circuit configurations specifically designed to execute specific processes. Furthermore, the various processes may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (e.g., multiple FPGAs, or a combination of a CPU and an FPGA). Furthermore, the hardware structure of these various processors is, more specifically, an electrical circuit that combines circuit elements such as semiconductor devices.

[0129] In the above embodiment, the program is pre-stored (installed) in a ROM, but the present invention is not limited to this. The program may be provided in a form recorded on a non-transitory tangible storage medium such as a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), or a USB (Universal Serial Bus) memory. The program may also be downloaded from an external device via a network.

[0130] The controller and methods described herein may be implemented by a special-purpose computer having a processor programmed to perform one or more functions embodied in a computer program. Alternatively, the apparatus and methods described herein may be implemented by a special-purpose computer having a processor configured with dedicated hardware logic circuitry. Alternatively, the apparatus and methods described herein may be implemented by one or more special-purpose computers configured by a combination of a processor executing a computer program and one or more hardware logic circuits. Furthermore, the computer program may be stored as instructions executed by a computer on a computer-readable non-transitory storage medium.

[0131] <Additional Notes> The features of the present invention are as follows.

[0132] (Appendix 1) a skeleton recognition unit that recognizes skeleton information of a moving object based on image data captured by a camera; a shadow recognition unit that recognizes shadow information of a moving object based on image data captured by the camera; a skeleton estimation unit that estimates skeleton information of the recognized moving object based on skeleton information of the recognized moving object and shadow information of the recognized moving object; A skeletal structure estimation device comprising:

[0133] (Appendix 2) The shadow recognition unit further generates an occupancy map representing a space occupied by the moving object based on the shadow information of the recognized moving object; 2. The skeleton estimation device according to claim 1, wherein the skeleton estimation unit estimates skeletal information of the moving body based on skeletal information of the recognized moving body and the generated occupancy map.

[0134] (Appendix 3) The camera further includes a light source control unit that controls on / off of a light source that irradiates a photographing area of ​​the camera with light, 3. The skeleton estimation device according to claim 1, wherein the shadow recognition unit recognizes shadow information of the moving object based on image data when the light source is on and image data when the light source is off.

[0135] (Appendix 4) the light source includes a plurality of light sources that irradiate light onto a photographing target area of ​​the camera from different directions; the light source control unit controls the on / off of each of the plurality of light sources; 4. The skeleton estimation device according to claim 3, wherein the shadow recognition unit recognizes shadow information of the moving object based on image data for each of the plurality of light sources when the light source is on and image data for each of the plurality of light sources when the light source is off.

[0136] (Appendix 5) The camera includes a plurality of cameras installed to capture images of a target area from different directions; 5. The skeleton estimation device according to any one of Supplementary Note 1 to Supplementary Note 4, wherein the shadow recognition unit recognizes shadow information of the moving object based on image data captured by the plurality of cameras.

[0137] (Appendix 6) The skeleton estimation device according to any one of Supplementary Note 1 to Supplementary Note 5, wherein the skeleton estimation unit estimates the skeletal information of the recognized moving body based on time series data of skeletal information of the recognized moving body and time series data of shadow information of the recognized moving body.

[0138] (Appendix 7) the skeleton recognition unit further generates a class map representing information about a part of the moving object for each voxel based on the skeleton information of the recognized moving object; 3. The skeleton estimation device according to claim 2, wherein the skeleton estimation unit estimates skeleton information of the moving object based on the generated class map and the generated occupancy map.

[0139] (Appendix 8) the skeleton recognition unit recognizes skeleton information of a moving object based on image data captured by the camera, and outputs a result of tracking the recognized moving object; The skeleton estimation device according to any one of Supplementary Note 1 to Supplementary Note 7, wherein the skeleton estimation unit estimates skeletal information of the moving body based on the skeletal information of the recognized moving body, the output tracking result of the moving body, and shadow information of the recognized moving body, and outputs the tracking result of the recognized moving body.

[0140] (Appendix 9) Based on the image data captured by the camera, the skeletal information of the moving object is recognized. Recognizing shadow information of a moving object based on image data captured by the camera; estimating skeletal information of the moving body based on skeletal information of the recognized moving body and shadow information of the recognized moving body; This is a computer-implemented skeletal estimation method.

[0141] (Appendix 10) At least one processor has Based on the image data captured by the camera, the skeletal information of the moving object is recognized, Recognizing shadow information of a moving object based on image data captured by the camera; and estimating skeletal information of the moving body based on skeletal information of the recognized moving body and shadow information of the recognized moving body. Skeletal estimation program. [Explanation of symbols]

[0142] 10 Skeleton estimation system 11 CPU 20 Camera 22 Light source 24 Light source control device 30 Data Acquisition Section 32, 52 Skeleton recognition unit 34, 54 Class map generation unit 36, 56 Shadow recognition section 38, 58 Occupancy map generation unit 40, 60 Skeleton estimation section 50 Calculation data acquisition section 62 Learning Department 100 Skeleton estimation device 150 Learning Device

Claims

1. a skeleton recognition unit (32) that recognizes skeleton information of a moving object based on image data captured by a camera (20); a shadow recognition unit (36) that recognizes shadow information of a moving object based on image data captured by the camera; a skeleton estimation unit (40) that estimates skeleton information of the moving body based on skeleton information of the recognized moving body and shadow information of the recognized moving body; A skeleton estimation device (100) including:

2. The shadow recognition unit further generates an occupancy map representing a space occupied by the moving object based on the shadow information of the recognized moving object; The skeletal structure estimation device according to claim 1 , wherein the skeletal structure estimation unit estimates skeletal structure information of the moving object based on the skeletal structure information of the recognized moving object and the generated occupancy map.

3. The camera further includes a light source control unit (24) that controls on / off of a light source (22) that irradiates light onto a photographing area of ​​the camera, The skeleton estimation device according to claim 1 , wherein the shadow recognition unit recognizes shadow information of the moving object based on image data when the light source is on and image data when the light source is off.

4. the light source includes a plurality of light sources that irradiate light onto a photographing target area of ​​the camera from different directions; the light source control unit controls the on / off of each of the plurality of light sources; 4. The skeleton estimation device according to claim 3, wherein the shadow recognition unit recognizes shadow information of the moving object based on image data for each of the plurality of light sources when the light source is on and image data for each of the plurality of light sources when the light source is off.

5. The camera includes a plurality of cameras installed to capture images of a target area from different directions; The skeletal structure estimation device according to claim 1 , wherein the shadow recognition unit recognizes shadow information of the moving object based on image data captured by the plurality of cameras.

6. The skeletal estimation device according to claim 1 , wherein the skeletal estimation unit estimates the skeletal information of the recognized moving body based on time series data of skeletal information of the recognized moving body and time series data of shadow information of the recognized moving body.

7. the skeleton recognition unit further generates a class map representing information about a part of the moving object for each voxel based on the skeleton information of the recognized moving object; The skeletal structure estimation device according to claim 2 , wherein the skeletal structure estimation unit estimates skeletal structure information of the moving object based on the generated class map and the generated occupancy map.

8. the skeleton recognition unit recognizes skeleton information of a moving object based on image data captured by the camera, and outputs a result of tracking the recognized moving object; 2. The skeletal estimation device according to claim 1, wherein the skeletal estimation unit estimates skeletal information of the moving body based on the skeletal information of the recognized moving body, the output result of tracking the moving body, and shadow information of the recognized moving body, and outputs the result of tracking the recognized moving body.

9. Based on the image data captured by the camera, the skeletal information of the moving object is recognized. Recognizing shadow information of a moving object based on image data captured by the camera; estimating skeletal information of the moving body based on skeletal information of the recognized moving body and shadow information of the recognized moving body; This is a computer-implemented skeletal estimation method.

10. At least one processor (11) Based on the image data captured by the camera, the skeletal information of the moving object is recognized, Recognizing shadow information of a moving object based on image data captured by the camera; and estimating skeletal information of the moving body based on skeletal information of the recognized moving body and shadow information of the recognized moving body. Skeletal estimation program.

Citation Information

Patent Citations

  • Tracking device for occlusion

    JP2022066917A

  • Occlusion-aware prediction of human behavior

    JP2023552105A