3D data acquisition system, 3D data acquisition method, and 3D data acquisition program
The 3D data acquisition system uses synchronized depth and imaging cameras with resolution adjustment and reference data to accurately identify and reconstruct feature points of multiple objects, addressing misidentification issues in existing systems and enhancing posture estimation for diverse human subjects.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KYOTO UNIV
- Filing Date
- 2024-10-29
- Publication Date
- 2026-05-15
AI Technical Summary
Existing automatic pose estimation systems struggle to accurately identify human bodies that deviate from standard models, such as infants, the elderly, or those with limb deficiencies, especially when multiple such subjects are imaged together or in close proximity, leading to misidentification or failure in 3D data acquisition.
A 3D data acquisition system utilizing a depth camera and a separate imaging camera, combined with a processing device, generates 2D and 3D data to accurately identify and reconstruct feature points of multiple objects, including those difficult to identify, by synchronizing and adjusting data resolutions and using reference data for precise identification.
The system enables accurate 3D data acquisition of multiple objects, including difficult-to-identify subjects, by generating high-precision 3D position information and skeletal models, even in complex imaging scenarios, improving identification and posture estimation accuracy.
Smart Images

Figure 2026078811000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a 3D data acquisition system, a 3D data acquisition method, and a 3D data acquisition program.
Background Art
[0002] In recent years, in research on human motion and the like, a technology for automatic pose estimation has been used. In one example, a target human body is identified from data captured by a camera or the like, and a three-dimensional pose of the identified human body is estimated. Here, in the identification of the human body, a dataset such as the "MPII Human Pose Dataset" is used. In the pose tracking by Azure Kinect evaluated in Non-Patent Document 1 below, the use of a visible light camera and a depth camera aims to improve the accuracy of human motion tracking. In another example, OpenPose and the like that identify feature points and skeletal structures of a human body from captured data such as images and videos by deep learning are known.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The aforementioned datasets, such as the "MPII Human Pose Dataset," are based on a model of a human body of a certain size and with all five limbs intact (i.e., an adult). Therefore, in automatic pose estimation based on depth data using the above datasets, human bodies that are far removed from the above model (difficult-to-identify subjects), such as infants (especially young children), the elderly (especially the very elderly), and those with limb deficiencies, tend to be difficult to identify as targets for automatic pose estimation. In addition, when identifying human bodies to be targeted for automatic pose estimation from data in which multiple human bodies, including the aforementioned difficult-to-identify subjects, are imaged simultaneously, the difficult-to-identify subjects are prone to not being identified as targets at all, or being identified as part of another subject. This problem is particularly likely to occur when the difficult-to-identify subjects are in close proximity or in contact with other subjects.
[0005] The aforementioned problem tends to occur frequently when using automatic pose estimation techniques based on depth data, such as those disclosed in Non-Patent Document 1. On the other hand, with automatic pose estimation based on imaging data such as OpenPose, the above problem tends to occur less frequently, even from imaging data where the difficult-to-identify object is in close proximity to or in contact with another object. However, the data obtained using typical OpenPose is 2D data of the human skeletal structure, not 3D data.
[0006] One aspect of this disclosure aims to provide a 3D data acquisition system, a 3D data acquisition method, and a 3D data acquisition program that can acquire 3D data with high accuracy, even when the target of automatic estimation includes multiple targets, such as targets that are difficult to identify. [Means for solving the problem]
[0007] The 3D data acquisition system, 3D data acquisition method, and 3D data acquisition program relating to one aspect of this disclosure are as follows. [1] A depth camera that images multiple objects, including objects that are difficult to identify, This camera is different from the aforementioned depth camera and is a camera that images multiple objects, A processing device that processes depth data transmitted from the depth camera and imaging data transmitted from the camera, Equipped with, The aforementioned processing apparatus is An acquisition unit that acquires the depth data and the imaging data, A 2D data generation unit generates 2D data from the aforementioned imaging data, including 2D position information of each of the multiple objects' feature points. A 3D data acquisition system comprising: a 3D data generation unit that generates 3D data including 3D position information of each of the feature points of the multiple objects based on the depth data and the 2D data. [2] The 3D data acquisition system according to [1], wherein the 3D data generation unit generates the 3D position information by reconstructing the 2D position information in three dimensions using the depth data. [3] The 3D data acquisition system according to [2], wherein the 3D data generation unit reconstructs the 2D position information in three dimensions based on the 2D position information, the focal length of the depth camera, and the optical center distance of the depth camera. [4] The 3D data acquisition system according to any one of [1] to [3], further comprising an adjustment unit for adjusting at least one of the resolution of the depth data and the resolution of the imaging data. [5] The 3D data acquisition system according to any one of [1] to [4], wherein the 2D data generation unit acquires the 2D position information using OpenPose. [6] The 3D data generation unit identifies the difficult-to-identify object and objects different from the difficult-to-identify object from the plurality of objects based on the reference data, The aforementioned reference data is obtained by imaging the multiple objects placed at predetermined locations for a predetermined period of time, according to any one of the 3D data acquisition systems described in [1] to [5]. [7] The 3D data acquisition system according to [6], wherein the reference data includes at least one of the neck, shoulder width, upper arm, lower arm, and torso length of each of the multiple objects placed at a predetermined position. [8] The feature points indicate multiple physical features, including joints of the human body, as described in any of [1] to [7], a 3D data acquisition system. [9] The 3D data acquisition system described in any of [1] to [8], wherein the subject who is difficult to identify is an infant, an elderly person requiring care, and / or a person with limb deficiencies.
[10] The multiple subjects include the infants and young children who are difficult to identify, and the parents of the infants and young children, A 3D data acquisition system according to any one of [1] to [9], further comprising an estimation unit that estimates the posture of the infant and the posture of the parent based on the 3D data.
[11] A method for acquiring 3D data of multiple objects, including objects that are difficult to identify, The steps include acquiring depth data of the multiple objects using a depth camera, The steps include acquiring imaging data of the multiple objects using a camera different from the depth camera, The steps include obtaining 2D data from the aforementioned imaging data, which includes 2D position information of each of the feature points of the multiple objects, A method for acquiring 3D data, comprising the step of generating 3D data including 3D position information of each of the feature points of the multiple objects based on the depth data and the 2D data.
[12] A 3D data acquisition program for acquiring 3D data of multiple objects, including objects that are difficult to identify, On the computer, The steps include acquiring depth data of the multiple objects acquired by the depth camera, The steps include obtaining 2D data, including 2D positional information of each feature point of the multiple objects, from imaging data of the multiple objects acquired by a camera different from the depth camera, A 3D data acquisition program that performs the steps of generating 3D data including 3D position information of each of the feature points of the multiple objects based on the depth data and the 2D data. [Effects of the Invention]
[0008] According to one aspect of the present disclosure, even if the objects to be automatically estimated are a plurality of objects including difficult-to-identify objects, it is possible to provide a 3D data acquisition system, a 3D data acquisition method, and a 3D data acquisition program that can acquire 3D data with each of the plurality of objects accurately identified.
Brief Description of Drawings
[0009] [Figure 1] FIG. 1 is a schematic configuration diagram of a 3D data acquisition system according to an embodiment. [Figure 2] FIG. 2 is a flowchart showing an example of a 3D data acquisition method according to an embodiment. [Figure 3] FIG. 3(a) is a schematic diagram of a main part of imaging data showing a plurality of objects, namely an infant (a difficult-to-identify object) and the mother of the infant (an object different from the difficult-to-identify object), and FIG. 3(b) is a diagram in which a 2D skeleton model connecting the feature points of the infant and a 2D skeleton model connecting the feature points of the mother in FIG. 3(a) are superimposed. [Figure 4] FIG. 4(a) is depth data including a plurality of objects, namely an infant (a difficult-to-identify object) and the mother of the infant (an object different from the difficult-to-identify object), and FIG. 4(b) is the adjusted depth data. [Figure 5] FIG. 5 is imaging data of a plurality of objects, namely an infant (a difficult-to-identify object) and the mother of the infant (an object different from the difficult-to-identify object), with 3D feature points superimposed. [Figure 6] FIG. 6 is a schematic diagram of imaging data of a plurality of objects with 3D feature points obtained by a method according to a comparative example superimposed. [Figure 7] FIGS. 7(a) and (b) are schematic diagrams of imaging data of a plurality of objects with 3D feature points obtained by a method according to a comparative example superimposed.
Embodiments for Carrying Out the Invention
[0010] Embodiments of this disclosure will be described in detail below with reference to the attached drawings. In the following description, the same reference numerals will be used for elements that are identical or have the same function, and redundant descriptions will be omitted. The terms “identical” and similar terms used herein are not limited to “exactly identical.” Furthermore, since the drawings are for conceptual purposes to illustrate embodiments, the dimensions and ratios of the components shown may differ from those of actual components.
[0011] The 3D data acquisition system according to this embodiment is a system that utilizes so-called video motion capture. This 3D data acquisition system estimates feature points such as joints from multiple types of imaging data, performs three-dimensional reconstruction from the estimated feature points, and identifies the object to be captured by the imaging data. When using this 3D data acquisition system, the object does not need to be fitted with markers, sensors, etc. Therefore, motion capture can be performed without restricting the object.
[0012] The 3D data acquisition system according to this embodiment estimates the three-dimensional position (3D position information) of each feature point of multiple objects, including objects that are difficult to identify, and estimates the posture of each of the multiple objects based on these feature points. In addition, the 3D data acquisition system can also estimate the movement of each of the multiple objects by accumulating the estimated postures of each of the multiple objects. In this embodiment, the multiple objects are multiple human bodies, but are not limited to this. Animals, robots, etc., that have a multi-joint structure similar to the human body may also be objects in the 3D data acquisition system. The number of multiple objects is not particularly limited, but from the viewpoint of estimation accuracy, it may be 15 or less, 10 or less, or 8 or less.
[0013] Human bodies that are different from those that are difficult to identify are those of a certain size, with all five limbs intact (i.e., adults), and whose body types are similar to many of the models in the human body estimation database. In other words, human bodies that are different from those that are difficult to identify are those with a skeletal structure similar to or identical to the skeletal structure of the models pre-set in the database. On the other hand, human bodies that are difficult to identify are those that are far removed from the above models, such as infants (especially young children), elderly people requiring care (especially the very elderly), and human bodies that lack joints that are normally present, such as those with limb deficiencies. Human bodies that are difficult to identify and those that are different from those that are difficult to identify are distinguished, for example, based on reference data (details will be described later). In one example, objects that are difficult to identify and objects that are not difficult to identify are distinguished based on at least one length of each of the following in multiple objects placed at predetermined locations, included in the above reference data: neck (e.g., from the base of the neck to the tip of the nose), shoulder width (e.g., from the neck to the left shoulder and / or from the neck to the right shoulder), upper arm (from the left shoulder to the left elbow and / or from the right shoulder to the right elbow), forearm (from the left elbow to the left hand and / or from the right elbow to the right hand), and torso (from the neck to the pelvis). In another example, objects that are difficult to identify and objects that are not difficult to identify are distinguished by whether the maximum length of the line segment between feature points is less than a predetermined threshold. This is because the distance from the elbow to the shoulder of an infant is significantly shorter than the distance from the elbow to the shoulder of the above model. Also, the positional relationship between the pelvis and neck of an elderly person may differ significantly from the positional relationship between the pelvis and neck of the above model. Due to these factors, in cases where multiple objects are closely related, the aforementioned difficult-to-identify objects may be mistakenly identified as part of another object. Keypoints represent multiple physical features, including, for example, the joints of the human body being studied. However, since keypoints are points estimated by the 3D data acquisition system, there may be cases where the keypoints do not coincide with the joints of the human body being studied.
[0014] The information output from the 3D data acquisition system according to this embodiment can be used, for example, to estimate the relationships and interactions between multiple subjects. For example, if the multiple subjects include an infant who is difficult to identify and the infant's parents, the 3D data acquisition system, or a user of the 3D data acquisition system, can estimate the normal communication methods and relationship between the parent and child based on the information output from the system. For example, if the multiple subjects include a very elderly person requiring care who is difficult to identify and the caregiver of the very elderly person, the appropriateness of the care methods can be estimated based on the information output from the 3D data acquisition system.
[0015] Figure 1 is a schematic diagram of the 3D data acquisition system according to this embodiment. As shown in Figure 1, the 3D data acquisition system 1 comprises an imaging device 2, a processing device 3, and a display device 4. Figure 1 shows a block diagram showing part of the hardware configuration of the imaging device 2 and a block diagram showing the functional configuration of the processing device 3. In the 3D data acquisition system 1, the imaging device 2, the processing device 3, and the display device 4 may be integrated with each other or may be separate. In the latter case, the imaging device 2, the processing device 3, and the display device 4 may be interconnected by electronic communication. The imaging device 2 is, for example, located in a space or room for imaging multiple objects. The processing device 3 is, for example, a device located in a different location from the imaging device 2, such as a cloud computer. The display device 4 may be located in a space or room for imaging multiple objects, or it may be located in a different location from said space, or it may be portable.
[0016] The imaging device 2 is a device for imaging multiple objects and is fixed in a predetermined position. By imaging multiple objects, the imaging device 2 generates still image data and / or video data including the multiple objects. Each data generated by the imaging device 2 is transmitted to the processing device 3. The imaging device 2 has a depth camera 21 and a camera 22 that is different from the depth camera 21. The imaging range of the depth camera 21 and the imaging range of the camera 22 coincide with each other. When the imaging device 2 is operating, the depth camera 21 and the camera 22 are synchronized. For example, the depth camera 21 and the camera 22 are synchronized by using an external synchronization signal generator. In this embodiment, the resolution of the depth camera 21 is lower than the resolution of the camera 22.
[0017] The depth camera 21 is a device that generates depth data indicating distance information for multiple objects. The depth camera 21 may be a stereo camera, a ToF camera (Time of Flight), or a structured light camera. A stereo camera generates depth data using the parallax between two cameras. A ToF camera generates depth data using the time it takes for emitted light to reflect and return. A structured light camera generates depth data by analyzing changes in projected dot patterns, grid patterns, etc. In this embodiment, the depth data is a 2D image (depth map) in which distance information is assigned to each pixel in the image data. The distance information displayed on the depth map is shown in grayscale or by a color gradient. Note that the depth data may be RGB-D data, a point cloud, etc. instead of a depth map. The depth data is transmitted directly or indirectly from the depth camera 21 to the processing device 3. The transmission of depth data may be performed in real time or not in real time. In this embodiment, the depth camera 21 is located in one place, but is not limited to this. From the perspective of improving the accuracy of distance information for multiple objects, the imaging device 2 may have multiple depth cameras 21 arranged at different positions.
[0018] Camera 22 is a visible light camera that generates imaging data for multiple objects. The imaging data is, for example, RGB image data. The imaging data is transmitted directly or indirectly from camera 22 to processing unit 3. The transmission of imaging data may be performed in real time or not in real time. For example, a combination of depth data and imaging data captured at the same time may be transmitted to processing unit 3 by a computer in an imaging device 2 (not shown). In this embodiment, the imaging data is 2D image data, but is not limited to this. For example, if the imaging device 2 has multiple cameras 22 arranged at different positions, the imaging data may be 3D image data.
[0019] In one example, a depth map (resolution: 800 x 600) is generated by the depth camera 21 at 60fps, while RGB image data (resolution: 1200 x 800) is generated by the camera 22 at 60fps. In this case, video data of the depth map can be generated by the depth camera 21, and RGB video data can be generated by the camera 22.
[0020] The processing unit 3 is a device that processes depth data transmitted from the depth camera 21 and imaging data transmitted from the camera 22. Physically, it has a computer (not shown) that includes a processor, a CPU (Central Processing Unit), and a recording medium, such as RAM (Random Access Memory) or ROM (Read Only Memory). As shown in Figure 1, the processing unit 3 has, as functional components, an acquisition unit 31, an adjustment unit 32, a 2D data generation unit 33, a 3D data generation unit 34, an estimation unit 35, a storage unit 36, and a communication unit 37. Each functional unit of the processing unit 3 is realized by loading a program (a 3D data acquisition program according to the embodiment) onto the aforementioned computer (not shown), and under the control of the CPU, reading and writing data to the RAM. The CPU included in the processing unit 3 makes the functional components shown in Figure 1 function by executing the above program. This sequentially executes the processing corresponding to the 3D data acquisition method described later. The CPU may be a standalone hardware unit or a soft processor implemented in a programmable logic such as an FPGA. The same applies to RAM and ROM as to the CPU. All data necessary for the execution of the above program, as well as all data generated by the execution of the above program, are stored on the storage medium. The functions of the functional components of the processing unit 3 are described in detail below.
[0021] The acquisition unit 31 acquires depth data and imaging data. The acquisition unit 31 receives depth data transmitted from the depth camera 21 and imaging data transmitted from the camera 22. The acquisition unit 31 transmits the depth data to the adjustment unit 32 and the imaging data to the 2D data generation unit 33.
[0022] The adjustment unit 32 adjusts at least one of the resolution of the depth data and the resolution of the imaging data. The adjustment unit 32 matches the resolution (size) of the depth data with the resolution (size) of the imaging data. For example, it performs a process to interpolate pixels in low-resolution data and a process to decimate pixels in high-resolution data. This prevents or suppresses the occurrence of positional misalignment (pixel misalignment) between multiple objects in the depth data and multiple objects in the imaging data. The adjustment unit 32 transmits the adjusted depth data and / or imaging data to the 3D data generation unit 34. The adjustment unit 32 may also transmit the adjusted depth data and / or imaging data to the display device 4 via the communication unit 37. In this case, the adjusted depth data and / or imaging data can be confirmed on the display device 4. The user of the 3D data acquisition system 1 can then manually adjust the adjusted depth data and / or imaging data. The manually adjusted depth data and / or imaging data generated in this way is transmitted to the adjustment unit 32 and the 3D data generation unit 34 via the communication unit 37. The depth data and / or imaging data after manual adjustment may be transmitted to the adjustment unit 32, thereby providing feedback on the resolution conversion process performed by the adjustment unit 32. In this embodiment, the adjustment unit 32 performs a conversion to match the resolution of the depth data to the resolution of the imaging data. In the following description, it will be assumed that only the adjusted depth data is used.
[0023] The 2D data generation unit 33 generates 2D data containing 2D position information of each feature point of multiple objects from the imaging data acquired by the acquisition unit 31. In one example, the above 2D data is generated using OpenPose, which is open-source software. In OpenPose, 18 feature points are set for each body. Specifically, the 18 feature points consist of 13 joints, the nose, left and right eyes, and left and right ears. OpenPose generates Part Confidence Maps (PCMs) for each feature point on the body in the imaging data offline or in real time by using a trained convolutional neural network (CNN). Subsequently, OpenPose generates one or more 2D skeletal models based on each generated feature point. Then, the 2D data generation unit 33 identifies each of the multiple objects in the imaging data based on these 2D skeletal models. The above-mentioned PCM is one of the indicators that represents the spatial distribution of the likelihood of the 2D position of feature points on the body, including each joint.
[0024] Each generated feature point is represented by its pixel position in the RGB image, which is the imaging data, and can be expressed in XY coordinates (horizontal coordinates). That is, for each feature point generated by the 2D data generation unit 33, both a horizontal position and a vertical position are defined. The 2D position information corresponding to the XY coordinates of each feature point is included in the 2D data. Furthermore, the 2D skeletal model, which represents a stick-figure-like shape by connecting the feature points, is also included in the 2D data. The results of identifying each of the multiple objects in the imaging data are included in the 2D data. Therefore, the 2D data generated by the 2D data generation unit 33 includes the 2D position information of each feature point of the multiple objects. In addition, the 2D data generation unit 33 may generate estimated pose data for each of the multiple objects based on the 2D skeletal model. In this case, the estimated data may be included in the 2D data. In generating the pose estimation data, the position of each feature point, the distance between feature points, and the angles formed by multiple line segments connecting the feature points are used.
[0025] The 2D data generation unit 33 may transmit the 2D data to the display device 4 via the communication unit 37. In this case, the 2D data can be viewed on the display device 4. The 2D data viewed on the display device 4 is, for example, an RGB image in which feature points and a 2D skeletal model are superimposed on the image data. The user of the 3D data acquisition system 1 can then manually adjust the position of the feature points in the 2D data, the results of identifying multiple objects, etc. In this case, the manually adjusted 2D data is transmitted to the 3D data generation unit 34 via the communication unit 37.
[0026] In generating the 2D data in the 2D data generation unit 33, reference data may be used. In this case, a 2D skeletal model is more easily generated based on the generated feature points, making it easier for each of the multiple objects to be identified as a target for automatic pose estimation. Therefore, difficult-to-identify objects become easier to identify as targets for automatic pose estimation, and less likely to be identified as part of another object. The reference data is data for clarifying what the imaging target by the 3D data acquisition system 1 is. The reference data includes the skeletal models of each of the multiple objects estimated by the 3D data acquisition system 1. The reference data is, for example, default 2D data that includes 2D position information of each feature point for multiple objects placed at predetermined positions, and is stored in the storage unit 36 in advance. The default 2D data is acquired, for example, when the 3D data acquisition system 1 is first put into use, and is generated in the same way as the generation of 2D data described later. The default 2D data may include information on the two-dimensional length of each of the multiple objects (2D size information) obtained from 2D imaging data accumulated for each of the multiple objects over a predetermined period. In one example, the default 2D data may include the length of at least one of the following for each of the multiple objects: the neck, shoulder width, upper arm, lower arm, and torso. In this case, since the reference data includes data for identifying each of the multiple objects being imaged, the identification of difficult-to-identify objects and objects other than difficult-to-identify objects can be made more accurate. The neck length mentioned above is the average value of multiple neck lengths obtained from 2D imaging data accumulated over a predetermined period. The same applies to the lengths of the upper arm, lower arm, and torso mentioned above.
[0027] The above reference data may include default 3D data containing 3D position information of each feature point of multiple objects placed at predetermined locations. This default 3D data is acquired, for example, when the 3D data acquisition system 1 is started to be used, and is generated in the same manner as the 3D data generation described later. The default 3D data may also include the three-dimensional size (3D size information) of each of the multiple objects, obtained from 3D imaging data accumulated for each of the multiple objects over a predetermined period. For example, the default 3D data may include the size of at least one of the neck, shoulder width, upper arm, lower arm, and torso of each of the multiple objects as 3D size information. In this case, the identification of difficult-to-identify objects and objects other than difficult-to-identify objects can be made more accurate when generating 3D data as described later. The neck size mentioned above is the average value of the neck size obtained from multiple 3D imaging data accumulated over a predetermined period. The same applies to the upper arm, lower arm, and torso sizes mentioned above.
[0028] The 2D data generation unit 33 may use methods other than OpenPose. Note that OpenPose, a trained convolutional neural network (CNN), etc., are pre-stored in the memory unit 36. The CNN comprises an input layer, an intermediate layer (hidden layer), and an output layer, with the intermediate layer constructed using deep learning with training data.
[0029] The 3D data generation unit 34 generates 3D data, including 3D position information for each feature point of multiple objects, based on depth data and 2D data. The 3D data generation unit 34 generates 3D position information for each of the multiple feature points by using depth data to reconstruct the 2D position information in three dimensions. Therefore, the feature points (3D feature points) generated by the 3D data generation unit 34 have defined horizontal, vertical, and depth positions. The 3D data generation unit 34 also generates one or more 3D skeletal models based on the 3D position information and the 2D skeletal models included in the 2D data. The 3D data generation unit 34 then identifies multiple objects based on these 3D skeletal models. The 3D data generated by the 3D data generation unit 34 includes 3D position information for each feature point of multiple objects, a 3D skeletal model for each object, and so on. Furthermore, information regarding the size of at least one of the following parts of the 3D skeletal model of each subject may also be included in the 3D data: the neck (e.g., from the base of the neck to the tip of the nose), shoulder width (e.g., from the neck to the left shoulder and / or from the neck to the right shoulder), upper arm (from the left shoulder to the left elbow and / or from the right shoulder to the right elbow), forearm (from the left elbow to the left hand and / or from the right elbow to the right hand), and torso (from the neck to the pelvis).
[0030] In one example, the 3D data generation unit 34 reconstructs the 2D position information in three dimensions based on the 2D position information, the focal length of the depth camera 21, and the optical center distance of the depth camera 21. Specifically, the 3D data generation unit 34 reconstructs the 2D position information in three dimensions by using the following formula (matrix). This allows the 3D data generation unit 34 to generate 3D position information for each feature point obtained by the 2D data generation unit 33. In the following formula, x and y are the coordinates of a predetermined feature point included in the 2D position information, and f x ,f y is the focal length of the depth camera 21, and c x ,c yis the optical center distance of the depth camera 21, and X, Y, and Z are the coordinates of a predetermined 3D feature point that has been reconstructed in three dimensions. Specifically, x is the horizontal coordinate of a predetermined feature point included in the 2D position information, and X is the horizontal coordinate of a predetermined 3D feature point that has been reconstructed in three dimensions. y is the vertical coordinate of a predetermined feature point included in the 2D position information, and Y is the vertical coordinate of a predetermined 3D feature point that has been reconstructed in three dimensions. Z is the depth coordinate of a predetermined 3D feature point that has been reconstructed in three dimensions.
[0031]
number
[0032] Similar to the 2D data generation unit 33, the above-mentioned reference data may be used in the generation of the 3D data in the 3D data generation unit 34. In this case, for example, the 3D data generation unit 34 generates one or more 3D skeletal models based on the 3D position information of each of the multiple feature points, the 2D skeletal model included in the 2D data, and the above-mentioned reference data. In addition, the 3D data generation unit 34 may identify difficult-to-identify objects and objects different from difficult-to-identify objects from a group of objects based on the above-mentioned reference data (in particular, the above-mentioned 3D size information). For example, difficult-to-identify objects and objects different from difficult-to-identify objects may be identified from a group of objects based on the 3D size information included in the above-mentioned reference data and information regarding the size of each 3D skeletal model included in the 3D data. The 3D data generation unit 34 may transmit the 3D data to the display device 4 via the communication unit 37. In this case, the 3D data can be viewed on the display device 4. The 3D data viewed by the display device 4 is, for example, an RGB image in which the 3D feature points and the 3D skeletal model are superimposed on the imaging data. The user of the 3D data acquisition system 1 can then manually adjust the position, coordinates, and identified objects of feature points in the 3D data. The manually adjusted 3D data generated in this way is then transmitted to the estimation unit 35 via the communication unit 37.
[0033] The estimation unit 35 estimates the posture of each of the multiple objects based on the 3D data. The estimation unit 35 estimates the three-dimensional posture (3D posture) of each of the multiple objects. In generating the 3D posture estimation data, the positions of each feature point included in the 3D data, the distance between feature points, and the angles formed by multiple line segments connecting the feature points are used. For example, if the multiple objects include an infant, which is a difficult object to identify, and the parent of the infant, the estimation unit 35 estimates the posture of the infant and the posture of the parent.
[0034] The estimation unit 35 estimates the changes in the posture of each of the multiple objects based on multiple 3D data. Based on this estimated change data, the estimation unit 35 generates 3D motion estimation data showing the estimated results of the movement of each of the multiple objects. The 3D motion estimation data for each object is displayed on the display device 4, for example, via the communication unit 37. The 3D motion estimation data can be used, for example, to determine whether the movement of the multiple objects is appropriate, or to estimate the relationships between the multiple objects.
[0035] The memory unit 36 stores various data necessary for executing the 3D data acquisition program described above, various data generated by the execution of the program, and the default 3D data. The various data include, for example, 2D data including depth data, imaging data, 2D position information and estimated data of the 2D skeleton model, 3D data including 3D position information and estimated data of the 3D skeleton model, estimated 3D posture data, and estimated 3D motion data.
[0036] The communication unit 37 exchanges information with the imaging device 2, the display device 4, and other devices via wired and / or wireless connections. Other devices with which the communication unit 37 exchanges information include, for example, an external synchronization signal generator and a printer.
[0037] The display device 4 displays data transmitted from the imaging device 2 (depth data, imaging data) and data transmitted from the processing device 3 (2D data, 3D data, 3D pose estimation data, etc.). The display device 4 may be fixed or portable. In the latter case, the display device 4 may be, for example, a laptop PC, mobile phone, smartphone, or tablet. The display device 4 may also be editable for displaying data.
[0038] Next, with reference to Figure 2, an example of a 3D data acquisition method using a 3D data acquisition system for acquiring 3D data of multiple objects, including difficult-to-identify objects, according to this embodiment will be described. Figure 2 is a flowchart of an example of a 3D data acquisition method.
[0039] As shown in Figure 2, first, preparations for acquiring 3D data of multiple objects are carried out (step ST1). Step ST1 involves setting up the imaging device 2, inputting data to the processing device 3, and acquiring reference data. For example, in acquiring reference data, first, multiple objects placed at predetermined locations are imaged by the imaging device 2 for a predetermined period of time. Then, at least one of the following is acquired as reference data: 2D position information of each feature point of each of the multiple objects, 3D position information of each feature point, 2D skeletal model of each object, 3D skeletal model of each object, length in 2D of each object, and size in 3D of each object. This reference data is stored in the storage unit 36. The method for acquiring reference data in step ST1 is the same as, but not limited to, the method for acquiring 2D data and 3D data described below.
[0040] After step ST1, image data of multiple objects is acquired by camera 22 (step ST2). In step ST2, for example, RGB images of multiple objects are acquired. In one example, although it is a grayscale image, as shown in Figure 3(a), image data is acquired that shows multiple objects, including infant T (object difficult to identify) and infant T's mother M (object different from the object difficult to identify). Subsequently, after step ST2, 2D data including 2D position information of each feature point of the multiple objects is acquired from the image data (step ST3). In step ST3, for example, OpenPose, an open-source software, is used to generate each feature point on the body from the RGB image data, either offline or in real time. Then, a 2D skeletal model is generated based on the generated feature points. Then, each of the multiple objects is identified based on the 2D skeletal model. As a result, 2D data including 2D position information of each feature point of the multiple objects is acquired. In one example, as shown in Figure 3(b), a 2D skeletal model S1 connecting the feature points of infant T and a 2D skeletal model S2 connecting the feature points of mother M are generated. The 2D skeletal models S1 and S2 are shown superimposed on the imaging data. Note that in step ST3, the reference data acquired in step ST1 may be used. In this case, it becomes easier to distinguish between difficult-to-identify objects and other objects.
[0041] After step ST1, depth data of multiple objects is acquired by the depth camera 21 (step ST4). In step ST4, for example, a depth map (heatmap) showing the infant T and mother M is acquired, as shown in Figure 4(a). In this embodiment, steps ST2 and ST4 are performed simultaneously. This makes it possible to acquire imaging data and depth data of multiple objects at the same time. After step ST4, the resolution of the depth data is adjusted (step ST5). In step ST5, the depth data is converted to match the resolution (size) of the imaging data. This results in adjusted depth data as shown in Figure 4(b). Step ST5 may be performed simultaneously with step ST3, or at a different timing than step ST3.
[0042] After step ST3 and after step ST5, 3D data containing 3D position information for each feature point of multiple objects is generated based on depth data and 2D data (step ST6). In step ST6, the 2D position information is reconstructed in three dimensions based on the adjusted depth data and 2D data obtained in step ST5. This generates 3D position information for each of the multiple feature points. By performing step ST6, feature points (3D feature points) with defined horizontal, vertical, and depth positions are obtained. Next, a 3D skeletal model is generated based on the 3D feature points. Then, each of the multiple objects is identified based on the generated 3D skeletal model. For example, based on the 3D size information included in the above reference data and the size information of each 3D skeletal model included in the 3D data, difficult-to-identify objects and objects different from difficult-to-identify objects may be identified from the multiple objects. As a result, 3D data containing 3D position information for each feature point of multiple objects is generated. The generated 3D data is stored.
[0043] In one example, as shown in Figure 5, each 3D feature point is superimposed on the imaging data. In Figure 5, each of the multiple subjects is appropriately identified. Therefore, the feature points of the infant T are shown in white, and the feature points of the mother M are shown in gray. Although not shown, a 3D skeletal model may also be superimposed on the imaging data based on the 3D feature points.
[0044] In this embodiment, steps ST2 to ST6 described above are repeated. This generates data showing the time change of each 3D skeletal model of multiple objects, that is, data showing the time change of each pose of multiple objects (3D motion estimation data for each object). From this data, it becomes possible to analyze the relationships between each object. This analysis may be performed by the 3D data acquisition system 1.
[0045] The effects and advantages of the 3D data acquisition method using the 3D data acquisition system according to the present embodiment described above will be explained with reference to the comparative example. In the comparative example, 3D data is generated without using 2D data generated using OpenPose. In this comparative example, 3D data is generated by using the Azure Kinect SDK. Figures 6 and 7(a) and 7(b) are schematic diagrams of image data of multiple objects in which 3D feature points obtained by the method according to the comparative example are superimposed. In Figures 7(a) and 7(b), multiple objects are imaged at different timings. The position and posture of infant T, which is a difficult-to-identify object shown in Figure 6, and the position and posture of infant T's mother M are the same as the position and posture of infant T and infant T's mother M shown in Figure 3(a). As shown in Figure 6, in the comparative example, unlike in this embodiment, infant T is mistakenly identified as part of mother M. For this reason, in the comparative example, most of the feature points of infant T and the skeletal model are not generated. On the other hand, as shown in Figure 7(a), when the infant T is separated from its father F, 3D feature points are appropriately distributed to both the infant T and the father F. However, as shown in Figure 7(b), when the infant T is in close contact with its father F, similar to Figure 6, the infant T is mistakenly identified as part of the father F.
[0046] In contrast to the above comparative example, the 3D data acquisition method by the 3D data acquisition system 1 according to this embodiment first generates 2D data including 2D position information of each feature point of multiple objects from the imaging data captured by the camera 22. Subsequently, 3D data including 3D position information of each feature point of multiple objects is generated based on the depth data captured by the depth camera 21 and the above 2D data. In generating this 3D data, each of the multiple objects, including difficult-to-identify objects, is identified, and the 2D position information of the feature points corresponding to each identified object is used. As a result, 3D data including 3D position information of each 3D feature point for each of the multiple objects, including difficult-to-identify objects, is generated with high accuracy. Therefore, according to this embodiment, even if the target of automatic estimation is multiple objects including difficult-to-identify objects, it is possible to acquire 3D data that accurately identifies each of the multiple objects. In addition, by accumulating the above 3D data for a predetermined period and analyzing the large amount of accumulated 3D data, even if the target of automatic estimation is multiple objects including difficult-to-identify objects, 3D motion estimation of each object can be performed with high accuracy.
[0047] In one example, the 3D data acquisition system 1 includes an adjustment unit 32 that adjusts at least one of the resolutions of the depth data and the image data. This prevents or suppresses positional misalignment (pixel misalignment) between multiple objects in the depth data and multiple objects in the image data.
[0048] In one example, the 2D data generation unit 33 acquires 2D position information using OpenPose. This makes it possible to accurately identify each of multiple objects, including those that are difficult to identify.
[0049] In one example, the system includes a storage unit 36 that stores default 3D data, which includes 3D position information of the feature points of multiple objects placed at predetermined locations, and 3D data. This makes it possible to obtain 3D data that accurately identifies each of the multiple objects, even if the target of automatic estimation includes multiple objects that are difficult to identify.
[0050] The following describes a specific example of the usefulness of 3D data acquisition system 1. The specific example involves a case where the multiple subjects mentioned above include infants and their parents, who are difficult to identify, and involves the analysis of their everyday communication using the 3D motion estimation results. Generally, unlike adults, infants often exhibit unexpected behavior and spend long periods in close contact with their parents. Therefore, using the comparative examples described above, infants are often identified as part of their parents. Furthermore, using only OpenPose as described above yields only 2D data. Consequently, using either the comparative examples or OpenPose alone, it was not possible to thoroughly analyze the infant's movements or their interactions with their parents (communication methods, the appropriateness of communication methods, etc.). In contrast, using 3D data acquisition system 1 makes it less likely for infants to be identified as part of their parents, thus enabling the accurate acquisition of 3D motion estimation data for both the infant and their parents. Therefore, the use of this 3D motion estimation data allows for a thorough and appropriate analysis of the infant's communication methods, the appropriateness of those methods, and areas for improvement. In addition, the imaging device 2 in the 3D data acquisition system 1 can be placed, for example, in a room where an infant and their parents normally reside. On the other hand, the processing device 3 and display device 4 in the 3D data acquisition system 1 can be placed in a location different from the aforementioned room. Therefore, by using the 3D data acquisition system 1, it is possible to obtain 3D motion estimation data of an infant and their parents in a normal environment where no other people are present (i.e., an environment that is not a special environment such as a laboratory). In other words, by using the 3D data acquisition system 1, it is possible to obtain 3D motion estimation data of each infant and their parents in a natural environment without tension. By using this 3D motion estimation data, it becomes possible to appropriately and thoroughly analyze the communication methods between infants and their parents under normal circumstances, the appropriateness of those communication methods, and areas for improvement.
[0051] Another concrete example of the usefulness of 3D data acquisition system 1 is elderly care. In elderly care, for example, there are many opportunities for close contact between caregivers and those receiving care. Therefore, by using 3D data acquisition system 1, it becomes possible to accurately analyze the appropriateness of the care methods used by caregivers to those receiving care.
[0052] One aspect of this disclosure is not limited to the embodiments described above. One aspect of this disclosure can be further modified without departing from its essence. For example, in the embodiments described above, estimated pose data for each of the multiple objects may be included in the 3D data.
[0053] In the above embodiment, the resolution of the image data captured by the camera and the resolution of the image data captured by the depth camera are different, but are not limited to this. Their resolutions may be the same. In this case, the processing unit does not need to include an adjustment unit, and step ST5 does not need to be performed. Alternatively, the resolution of the image data captured by the camera may be adjusted to match the resolution of the image data captured by the depth camera.
[0054] In the above embodiment, the estimation unit estimates the posture and changes of each of the multiple objects, but it is not limited to this. For example, the 3D data generation unit may estimate the posture and changes of each of the multiple objects and generate posture estimation data, 3D motion estimation data, etc., for each object. Also, in the above embodiment, the posture estimation data and 3D motion estimation data for each object are said to be different from the 3D data, but it is not limited to this. The posture estimation data and 3D motion estimation data for each object may be included in the 3D data. [Explanation of Symbols]
[0055] 1...3D data acquisition system, 2...imaging device, 3...processing device, 4...display device, 21...depth camera, 22...camera, 31...acquisition unit, 32...adjustment unit, 33...2D data generation unit, 34...3D data generation unit, 35...estimation unit, 36...storage unit, 37...communication unit, F...father, M...mother, S1...2D skeletal model, S2...2D skeletal model, T...infant.
Claims
1. A depth camera that images multiple objects, including objects that are difficult to identify, This camera is different from the aforementioned depth camera and is a camera that images multiple objects, A processing device that processes depth data transmitted from the depth camera and imaging data transmitted from the camera, Equipped with, The aforementioned processing apparatus is An acquisition unit that acquires the depth data and the imaging data, A 2D data generation unit generates 2D data from the imaging data, including 2D position information of each of the multiple objects' feature points. The system includes a 3D data generation unit that generates 3D data including 3D position information of each of the feature points of the multiple objects based on the depth data and the 2D data. 3D data acquisition system.
2. The 3D data acquisition system according to claim 1, wherein the 3D data generation unit generates the 3D position information by using the depth data to reconstruct the 2D position information in three dimensions.
3. The 3D data acquisition system according to claim 2, wherein the 3D data generation unit reconstructs the 2D position information in three dimensions based on the 2D position information, the focal length of the depth camera, and the optical center distance of the depth camera.
4. The 3D data acquisition system according to any one of claims 1 to 3, further comprising an adjustment unit for adjusting at least one of the resolution of the depth data and the resolution of the imaging data.
5. The 3D data acquisition system according to any one of claims 1 to 3, wherein the 2D data generation unit acquires the 2D position information using OpenPose.
6. The 3D data generation unit identifies the difficult-to-identify object and objects different from the difficult-to-identify object from the plurality of objects based on the reference data. The 3D data acquisition system according to any one of claims 1 to 3, wherein the reference data is obtained by imaging the plurality of objects placed at predetermined positions for a predetermined period of time.
7. The 3D data acquisition system according to claim 6, wherein the reference data includes at least one of the neck, shoulder width, upper arm, lower arm, and torso length of each of the multiple objects placed at predetermined positions.
8. The 3D data acquisition system according to any one of claims 1 to 3, wherein the aforementioned feature points represent multiple features on the body, including joints of the human body.
9. The 3D data acquisition system according to any one of claims 1 to 3, wherein the subject that is difficult to identify is an infant, an elderly person requiring care, and / or a person with limb deficiency.
10. The aforementioned multiple targets include the infants and young children who are difficult to identify, and the parents of the infants and young children. The 3D data acquisition system according to any one of claims 1 to 3, further comprising an estimation unit that estimates the posture of the infant and the posture of the parent based on the 3D data.
11. A method for acquiring 3D data of multiple objects, including objects that are difficult to identify, The steps include acquiring depth data of the multiple objects using a depth camera, The steps include acquiring imaging data of the multiple objects using a camera different from the depth camera, The steps include obtaining 2D data from the imaging data, which includes 2D position information of each of the feature points of the multiple objects, The method includes the step of generating 3D data including 3D position information of each of the feature points of the plurality of objects based on the depth data and the 2D data. Methods for acquiring 3D data.
12. A 3D data acquisition program for acquiring 3D data of multiple objects, including objects that are difficult to identify, On the computer, The steps include acquiring depth data of the multiple objects acquired by the depth camera, The steps include obtaining 2D data including 2D positional information of each feature point of the multiple objects from imaging data of the multiple objects acquired by a camera different from the depth camera, A 3D data acquisition program that performs the steps of generating 3D data including 3D position information of each of the feature points of the multiple objects based on the depth data and the 2D data.