Data extension device, data extension method, and program
The data augmentation method aligns 3D and 2D joint coordinates to create accurate training data, addressing posture mismatches and enhancing the learning model's detection accuracy.
Patent Information
- Application Number
- JP2023570502
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-10-07
- Estimated Expiration
- 2041-12-27
AI Technical Summary
Existing methods for expanding training data for 3D joint coordinate detection in learning models face inaccuracies due to mismatches between 3D and 2D joint point coordinates, caused by differing postures in real space appearing similar in 2D images, leading to decreased detection accuracy.
A data augmentation method that projects 3D joint coordinates onto a 2D plane, aligns them with corresponding 2D coordinates using camera parameters, calculates similarity, and synthesizes new 2D images to create training data, ensuring similar postures are matched, thereby avoiding the use of images with different real-space postures.
Enhances the accuracy of training data expansion by aligning 3D and 2D joint coordinates, preventing the inclusion of images with differing real-space postures, thus improving the learning model's detection capabilities.
Smart Images

Figure 0007750306000003 
Figure 0007750306000004 
Figure 0007750306000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to a data expansion device and a data expansion method for expanding training data for constructing a learning model for estimating human posture, and also to a program for realizing them. Mu Regarding. [Background technology]
[0002] In recent years, a technology has been developed that estimates a person's posture by detecting the three-dimensional coordinates of each of the person's joints from a two-dimensional image (see, for example, Patent Document 1). Such a technology is expected to be used in fields such as image monitoring systems, sports, and games. Furthermore, in such a technology, a learning model is used to detect the three-dimensional coordinates of each of the person's joints.
[0003] The learning model is constructed by machine learning using, for example, two-dimensional coordinates of joints extracted from a person in an image (hereinafter referred to as "two-dimensional joint point coordinates") and three-dimensional coordinates of the extracted joints (hereinafter referred to as "three-dimensional joint point coordinates") as training data (see, for example, Non-Patent Document 1).
[0004] However, in order to improve the accuracy of detecting 3D joint coordinates using a learning model, it is necessary to prepare a large amount of training data, but preparing a large amount of training data is not easy. For this reason, Non-Patent Document 1 discloses a method for expanding the training data.
[0005] In the method disclosed in Non-Patent Document 1, first, each joint point constituting the 3D joint point coordinates of a specific person is projected onto a 2D plane. Next, among the projected joint points, joint points of a portion of the person are compared with 2D joint point coordinates prepared in advance, and matching 2D joint point coordinates are identified. Next, a portion corresponding to the identified 2D joint point coordinates is cut out from the 2D image corresponding to the identified 2D joint point coordinates. The cut-out portion is pasted into another 2D image to become a 2D image corresponding to the original 3D joint point coordinates. After that, the 2D joint point coordinates extracted from the obtained 2D image and the original 3D joint point coordinates are used as new training data. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Patent Publication No. 2021-47563 [Non-patent literature]
[0007] [Non-Patent Document 1] Gregory Rogez, Cordelia Schmid, “MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild”, arXiv:1607.02046v2 [cs.CV], 28 Oct 2016, [Retrieved November 1, 2021], Internet<URL:http: / / https: / / arxiv.org / pdf / 1607.02046.pdf> Summary of the Invention [Problem to be solved by the invention]
[0008] However, in the method disclosed in Non-Patent Document 1, there are cases where the original 3D joint point coordinates and the 3D joint point coordinates corresponding to the 2D joint point coordinates that match the projected joint points do not match. In other words, there are cases where the posture of the person corresponding to the original 3D joint point coordinates and the posture of the person corresponding to the matched 2D joint point coordinates differ in real space.
[0009] This is because different poses in real space may appear to be the same in a 2D image due to differences in viewpoint. When this occurs, the accuracy of the learning model's detection of 3D joint coordinates decreases.
[0010] An example of the object of the present disclosure is to provide a data augmentation device, a data augmentation method, and a method for augmenting training data in constructing a learning model for detecting three-dimensional joint point coordinates. program The purpose is to provide [Means for solving the problem]
[0011] In order to achieve the above object, a data extension device according to one aspect of the present disclosure includes: a data acquisition unit that acquires data including a set of three-dimensional coordinates of each of the joint points of a particular person; a projection processing unit that projects each of the three-dimensional coordinates included in the acquired data onto a two-dimensional plane to generate projected coordinates of each of the joint points; for each set of data in which a set of three-dimensional coordinates of each joint point of a person, a two-dimensional image of the person, and camera parameters are associated with each other, using the camera parameters to identify two-dimensional coordinates on the two-dimensional image corresponding to each of the three-dimensional coordinates of the set of data; Furthermore, for each group data, the acquired data or a set of three-dimensional coordinates included in the group data is manipulated so that the generated set of projected coordinates overlaps with the identified set of two-dimensional coordinates, and after the manipulation, a similarity between the set of three-dimensional coordinates included in the acquired data and the set of three-dimensional coordinates included in the group data is calculated; a data search unit that identifies the group data corresponding to the acquired data based on the similarity calculated for each group data; an image generating unit that combines a part or all of the two-dimensional image of the identified set of data with another two-dimensional image to generate a new two-dimensional image; Equipped with It is characterized by:
[0012] In order to achieve the above object, a data extension method according to one aspect of the present disclosure includes: a data acquisition step of acquiring data including a set of three-dimensional coordinates of each of the particular person's joint points; a projection processing step of projecting each of the three-dimensional coordinates included in the acquired data onto a two-dimensional plane to generate projected coordinates of each of the joint points; for each set of data in which a set of three-dimensional coordinates of each joint point of a person, a two-dimensional image of the person, and camera parameters are associated with each other, using the camera parameters to identify two-dimensional coordinates on the two-dimensional image corresponding to each of the three-dimensional coordinates of the set of data; Furthermore, for each group data, the acquired data or a set of three-dimensional coordinates included in the group data is manipulated so that the generated set of projected coordinates overlaps with the identified set of two-dimensional coordinates, and after the manipulation, a similarity between the set of three-dimensional coordinates included in the acquired data and the set of three-dimensional coordinates included in the group data is calculated; a data search step of identifying the group data corresponding to the acquired data based on the similarity calculated for each group data; an image generating step of synthesizing a part or all of the two-dimensional image of the identified set of data with another two-dimensional image to generate a new two-dimensional image; The present invention is characterized by having the following:
[0013] Furthermore, in order to achieve the above object, in one aspect of the present disclosure, program teeth, On the computer, a data acquisition step of acquiring data including a set of three-dimensional coordinates of each of the particular person's joint points; a projection processing step of projecting each of the three-dimensional coordinates included in the acquired data onto a two-dimensional plane to generate projected coordinates of each of the joint points; for each set of data in which a set of three-dimensional coordinates of each joint point of a person, a two-dimensional image of the person, and camera parameters are associated with each other, using the camera parameters to identify two-dimensional coordinates on the two-dimensional image corresponding to each of the three-dimensional coordinates of the set of data; Furthermore, for each group data, the acquired data or a set of three-dimensional coordinates included in the group data is manipulated so that the generated set of projected coordinates overlaps with the identified set of two-dimensional coordinates, and after the manipulation, a similarity between the set of three-dimensional coordinates included in the acquired data and the set of three-dimensional coordinates included in the group data is calculated; a data search step of identifying the group data corresponding to the acquired data based on the similarity calculated for each group data; an image generating step of synthesizing a part or all of the two-dimensional image of the identified set of data with another two-dimensional image to generate a new two-dimensional image; Run Ruko It is characterized by the following. [Effects of the Invention]
[0014] As described above, according to the present invention, training data can be expanded in constructing a learning model for detecting three-dimensional joint point coordinates. [Brief explanation of the drawings]
[0015] [Figure 1] FIG. 1 is a diagram showing a schematic configuration of a data extension device according to the first embodiment. [Figure 2] FIG. 2 is a diagram specifically illustrating the configuration of the data expansion device according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of target data used in the first embodiment. [Figure 4] FIG. 4 is an explanatory diagram illustrating the operation process of the 3D pose data set according to the first embodiment. [Figure 5] FIG. 5 is an explanatory diagram illustrating the similarity calculation process according to the first embodiment. [Figure 6] FIG. 6 is a diagram schematically showing a new two-dimensional image created in the first embodiment. [Figure 7] FIG. 7 is a flowchart showing the operation of the data extension device according to the first embodiment. [Figure 8] FIG. 8 is a diagram showing the configuration of a data expansion device according to the second embodiment. [Figure 9] FIG. 9 is an explanatory diagram illustrating the body shape modification process according to the second embodiment. [Figure 10] FIG. 10 is a flowchart showing the operation of the data extension device according to the second embodiment. [Figure 11] FIG. 11 is a block diagram showing an example of a computer that realizes the data extension device according to the first and second embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0016] (Embodiment 1) The data extension device, the data extension method, and the program according to the first embodiment will be described below with reference to FIGS.
[0017] [Device configuration] First, a schematic configuration of the data expansion device according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing a schematic configuration of the data expansion device according to the first embodiment.
[0018] 1, the data extension device 10 in the first embodiment is a device that extends training data, specifically, training data for constructing a learning model that estimates a person's posture. As shown in FIG. 1, the data extension device 10 includes a data acquisition unit 11, a projection processing unit 12, a data search unit 13, and an image generation unit 14.
[0019] The data acquisition unit 11 acquires data (hereinafter referred to as "target data") including a set of three-dimensional coordinates of each joint point of a specific person. The projection processing unit 12 projects each of the three-dimensional coordinates included in the acquired target data onto a two-dimensional plane to generate projected coordinates of each joint point.
[0020] The data search unit 13 executes the following process for each set of data. The set of data is data in which a set of 3D coordinates of each joint point of a person, a 2D image of the person, and camera parameters are associated with each other. First, for each set of data, the data search unit 13 uses the camera parameters to identify the 2D coordinates on the 2D image that correspond to each of the 3D coordinates of the set of data.
[0021] Next, the data search unit 13 manipulates the set of three-dimensional coordinates included in the acquired target data or group data so that the set of generated projected coordinates overlaps with the identified set of two-dimensional coordinates for each group data.
[0022] Here, "overlap" is not limited to cases where all two-dimensional coordinates constituting the set of projected coordinates completely match two-dimensional coordinates constituting the specified set of two-dimensional coordinates. It also includes cases where some two-dimensional coordinates constituting the set of projected coordinates match some two-dimensional coordinates of the specified set of two-dimensional coordinates.
[0023] Furthermore, if the similarity between a set of projected coordinates and a specified set of two-dimensional coordinates is equal to or greater than a set value, it can be determined that the former and the latter "overlap." In this case, the similarity is calculated, for example, by calculating the deviation between each two-dimensional coordinate constituting the set of projected coordinates and each two-dimensional coordinate in the specified set of two-dimensional coordinates, and then based on the total value, average value, etc. of the deviation.
[0024] Then, for each group of data, the data search unit 13 calculates the similarity between a set of three-dimensional coordinates included in the target data after the operation and a set of three-dimensional coordinates of the group of data. After that, the data search unit 13 identifies group of data corresponding to the acquired target data based on the similarity calculated for each group of data.
[0025] The image generator 14 synthesizes a part or all of the two-dimensional image of the identified set of data with another two-dimensional image to generate a new two-dimensional image. The image data of the new two-dimensional image is used as training data.
[0026] In this way, the data extension device 10 calculates the similarity between the set of 3D coordinates of articulation points of the target data and the set of 3D coordinates of articulation points of the data stored in the database, and if the two are similar, creates new training data using the corresponding 2D image.
[0027] Therefore, unlike conventional data augmentation, it is possible to avoid a situation in which training data is augmented using a 2D image in which the posture of a person in the real space is different from the original 2D image, even though the posture of the person in the 2D image is similar to that of the person in the real space.The data augmentation device 10 can augment training data while solving conventional problems in building a learning model for detecting 3D joint point coordinates.
[0028] Next, the configuration and functions of the data expansion device according to the first embodiment will be specifically described with reference to Fig. 2 to Fig. 6. Fig. 2 is a configuration diagram specifically showing the configuration of the data expansion device according to the first embodiment. Fig. 3 is a diagram showing an example of target data used in the first embodiment.
[0029] As shown in FIG. 2, in the first embodiment, the data extension device 10 includes a database 20 in addition to the data acquisition unit 11, the projection processing unit 12, the data search unit 13, and the image generation unit 14 described above.
[0030] In the first embodiment, the data acquisition unit 11 acquires a 3D pose data set shown in Fig. 3 as target data. As shown in Fig. 3, the 3D pose data set 30 is composed of a set of 3D coordinates for each joint point 31 of one person. The 3D pose data set also includes identification data (right wrist, left wrist, neck, etc.) for identifying each joint point 31.
[0031] In the example of FIG. 3, the three-dimensional coordinates of each joint point 31 are expressed in a camera coordinate system, but the coordinate system is not particularly limited. The three-dimensional coordinates of each joint point 31 may be expressed in a world coordinate system. The camera coordinate system is a coordinate system with the camera position as its origin. In the camera coordinate system, the horizontal direction of the camera is set as the x-axis, the vertical direction as the y-axis, and the optical axis direction as the z-axis. The z-coordinate represents the distance from the camera. The world coordinate system is a coordinate system that is arbitrarily set in real space, and the origin is set on the ground at the feet of the camera. In the world coordinate system, the vertical direction is set as the z-axis.
[0032] In the first embodiment, the projection processing unit 12 projects each of the joint points 31 (see FIG. 3) included in all or specific parts of the 3D pose data set 30 onto a two-dimensional plane, i.e., onto an image coordinate system, and generates projected coordinates (two-dimensional coordinates) in the image coordinate system for each of the joint points 31. The image coordinate system is a coordinate system on a two-dimensional image, and the upper left pixel is usually set as the origin.
[0033] The database 20 has registered in advance a plurality of sets of data 21. In the first embodiment, the sets of data 21 are data that associate a three-dimensional pose data set of a person, image data of a two-dimensional image of a person in the same pose as the three-dimensional pose data set, and camera parameters corresponding to these.
[0034] As camera parameters, internal parameters are used when the three-dimensional coordinates of the joint points are expressed in the camera coordinate system, and internal and external parameters are used when the three-dimensional coordinates of the joint points are expressed in the world coordinate system. Note that the internal parameters are expressed as a matrix connecting the camera coordinate system and the image coordinate system, focal length, optical axis deviation, etc. The external parameters are expressed as a matrix connecting the world coordinate system and the camera coordinate system, the camera position relative to the world coordinate system, and the camera tilt.
[0035] In this embodiment, the data search unit 13 uses internal parameters for each group of data to identify the corresponding two-dimensional coordinates in the image coordinate system for the three-dimensional coordinates of each joint point included in all or specific parts of the three-dimensional pose data set of the group of data.
[0036] Next, in the first embodiment, the data search unit 13 manipulates the 3D pose dataset of the target data for each set of group data so that the set of projective coordinates generated from the target data overlaps with the set of identified 2D coordinates. Then, for each set of group data, the data search unit 13 calculates the similarity between the manipulated 3D pose dataset and the 3D pose dataset of the group data. Furthermore, if the projective coordinates and 2D coordinates have been obtained for a specific body part, the data search unit 13 calculates the similarity using the 3D pose dataset of the specific body part.
[0037] Specifically, for each set of data, the data search unit 13 sets a condition that, for example, two or more joint points included in the generated set of projective coordinates match two or more joint points included in the identified set of two-dimensional coordinates. Then, the data search unit 13 performs, as an operation, one or a combination of translation, rotation, enlargement, and reduction on the three-dimensional pose data set (set of three-dimensional coordinates) of the target data or set of data so that the condition is satisfied.
[0038] Furthermore, the data search unit 13 obtains a unit vector pointing from a specific joint point to another joint point in the three-dimensional coordinates after the operation, and a unit vector pointing from a specific joint point to another joint point in the three-dimensional coordinates of the paired data. Then, the data search unit 13 calculates the similarity based on both the obtained unit vectors.
[0039] The operation process of the 3D pose data set and the similarity calculation process by the data search unit 13 will be described in more detail with reference to Figures 4 and 5. Figure 4 is an explanatory diagram illustrating the operation process of the 3D pose data set in embodiment 1. Figure 5 is an explanatory diagram illustrating the similarity calculation process in embodiment 1.
[0040] First, the 3D pose dataset of the target data is defined as p(={p1, p2, p n}), and the 3D pose dataset of the pair data in the database 20 is q(={q1, q2, q n}). p n and q n indicate articulation points.
[0041] As shown in Figure 4, in the target data, two joint points p j and p i Assume that the joint point p j and articulation point p j Let p be the set of joint points connected by bones. AD,j Let the joint point p j and p i is expressed as p in the 3D pose dataset. c j and p c i This joint point p c j and p c i The joint points obtained by projecting onto the image coordinate system are defined as p l j and p l i Also, p l i ∈p lAD,j is p l j is the articulation point furthest from
[0042] In addition, in the pair data, the two corresponding joint points q j and q i Let q be the set of joint points connected to these by bones. AD,j Let the joint point be q j and q i In the 3D pose dataset, q c j and q c i Joint point q j and q i The joint point in the image coordinate system corresponding to q l j and q l i Also, q l i ∈q l AD,j , q l j is the articulation point furthest from
[0043] As shown in FIG. 4, the data search unit 13 searches for a joint point p l j and p l i is the joint point q l j and q l i In the camera coordinate system, the 3D pose dataset q is c q is translated, rotated, scaled up, or scaled down, or a combination of these. l j and q l i The joint points in the image coordinate system, including the q, are also manipulated. l j and q c j are q l’ j and q c’j (See Figure 5). In the example of Fig. 4, rotation is performed only within the xy plane of the camera coordinate system. Enlargement and reduction are performed at the same magnification on the x-axis, y-axis, and z-axis of the camera coordinate system. Furthermore, in accordance with the operation by the data search unit 13, one or a combination of translation, rotation, enlargement, and reduction is performed on the two-dimensional image I constituting the set data. The two-dimensional image after the operation is designated as I'.
[0044] After the operation, the data search unit 13 searches for the joint point p c j From p c k ∈p C AD,j unit vector t pointing to jk In the data set, the joint point q c’ j From q c’ k ∈q C’ AD,j unit vector s pointing to jk Next, the data search unit 13 calculates the joint point p c j The structure centered on the joint point q c’ j The similarity D between the structure centered on j Calculate k. c k ∈p c AD,j are the indices of articulation points that satisfy
[0045]
number
[0046] In the above equation (1), the cosine similarity is used as the similarity. However, the first embodiment is not limited to this, and the similarity may be expressed as p c k ∈p C AD,j and q c’ k ∈qC’ AD,j The Euclidean distance between
[0047] The data search unit 13 calculates the similarity D for all the paired data stored in the database 20. j Among these, the similarity D j In addition, when the projected coordinates and the two-dimensional coordinates are obtained for a specific part, a set of data in which only the specific part is similar is identified.
[0048] When a set of data in which a specific part is similar is specified, the image generation unit 14 generates a patch image by cutting out the specific part (for example, the left leg, the right leg, the right arm, etc.) from the two-dimensional image I' after the above-mentioned operation. In addition, the image generation unit 14 generates a patch image by cutting out the joint point q in the image coordinate system after the operation. l’ j and the joint point q of the 3D pose dataset after the manipulation. c’ j Using the above, a part of the corresponding 3D pose data set is assigned to the generated patch image. Then, the image generation unit 14 generates a new 2D image by combining the generated patch image with another 2D image (such as an image of a person with a specific part obscured). The new 2D image obtained in this way is used as training data for building a learning model that estimates a person's posture.
[0049] Furthermore, in the first embodiment, the data search unit 13 can identify a set of data with the highest similarity for each different body part. In this case, the image generation unit 14 generates a patch image for each body part, and then pastes the patch image of each body part onto a background image to generate a new image of one person (a new two-dimensional image). At this time, the image generation unit 14 also synthesizes a three-dimensional pose data set corresponding to each patch image. The new two-dimensional image obtained in this way and the three-dimensional pose data set after synthesis also serve as training data for constructing a learning model that estimates a person's posture.
[0050] Fig. 6 is a diagram schematically illustrating a new two-dimensional image created in embodiment 1. In the example of Fig. 6, the new two-dimensional image is created by combining patch image 32, patch image 33, patch image 34, patch image 35, and background image 36, which are different in location.
[0051] [Device operation] Next, the operation of the data expansion device 10 in the first embodiment will be described with reference to Fig. 7. Fig. 7 is a flow diagram showing the operation of the data expansion device in the first embodiment. In the following description, Figs. 1 to 6 will be referred to as appropriate. In addition, in the first embodiment, the data expansion method is implemented by operating the data expansion device 10. Therefore, the description of the data expansion method in the first embodiment will be replaced by the following description of the operation of the data expansion device 10.
[0052] As shown in FIG. 7, first, the data acquisition unit 11 acquires a three-dimensional pose data set of a specific person as target data (step A1).
[0053] Next, the projection processing unit 12 projects each of the joint points 31 (see Figure 3) included in a specific part of the 3D pose data set 30 acquired in step A1 onto the image coordinate system, and generates projected coordinates (two-dimensional coordinates) of each of the joint points 31 in the image coordinate system (step A2).
[0054] Next, the data search unit 13 reads out group data from the database 20, and for each group data, uses the internal parameters to identify the corresponding two-dimensional coordinates in the image coordinate system for the three-dimensional coordinates of each joint point included in a specific part of the three-dimensional pose data set of the group data (step A3).
[0055] Next, the data search unit 13 manipulates the 3D pose data set or the 3D pose data set of the group data acquired in step A1 so that the set of projected coordinates generated in step A2 overlaps with the set of 2D coordinates identified in step A3 for each group data (step A4).
[0056] Specifically, in step A4, the data search unit 13 sets a condition that, for each set of data, two or more joint points included in the set of projective coordinates generated in step A2 match two or more joint points included in the set of 2D coordinates identified in step A3. Then, the data search unit 13 performs one or a combination of translation, rotation, enlargement, and reduction on the 3D pose data set or the 3D pose data set of the set of data acquired in step A1 so that the condition is satisfied.
[0057] Next, the data search unit 13 calculates, for each set of data, the similarity between the three-dimensional pose data set of the target data and the three-dimensional pose data set of the set of data after the operation of step A4 (step A5).
[0058] Specifically, in step A5, the data search unit 13 obtains a unit vector pointing from a specific joint point to another joint point in the three-dimensional coordinates after the operation and a unit vector pointing from the specific joint point to another joint point in the three-dimensional coordinates of the paired data. Then, the data search unit 13 calculates the similarity based on both the obtained unit vectors.
[0059] Next, the data search unit 13 identifies the set of data with the highest similarity based on the similarity calculated for each set of data in step A5 (step A6).
[0060] Next, the image generating unit 14 cuts out a specific part (for example, the left leg, the right leg, the right arm, etc.) of the two-dimensional image of the group data identified in step A6, and generates a patch image (step A7).
[0061] Thereafter, the image generation unit 14 generates a new two-dimensional image using the patch image generated in step A7, and further generates new training data using this (step A8). Specifically, the image generation unit 14 generates a new two-dimensional image using the patch image generated in step A7, a patch image already generated for another body part, and a background image.
[0062] In this way, the data expansion device 10 calculates the similarity between the 3D pose dataset as the target data and the 3D pose dataset stored in the database, and if the two datasets are similar, creates new training data using patch images generated from the corresponding 2D images. This prevents the training data from being expanded using 2D images of a person with a different pose in real space. According to the first embodiment, it is possible to expand the training data while solving conventional problems in building a learning model for detecting 3D joint point coordinates.
[0063] [program] The program in the first embodiment may be any program that causes a computer to execute steps A1 to A8 shown in Fig. 7. By installing and executing this program in a computer, the data extension device 10 and the data extension method in the first embodiment can be realized. In this case, the processor of the computer functions as a data acquisition unit 11, a projection processing unit 12, a data search unit 13, and an image generation unit 14 to perform processing. Examples of the computer include a general-purpose PC, a smartphone, and a tablet terminal device.
[0064] In the first embodiment, the database 20 may be realized by storing the data files that constitute it in a storage device such as a hard disk provided in the computer, or may be realized by a storage device of another computer.
[0065] The program in the first embodiment may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as any one of the data acquisition unit 11, the projection processing unit 12, the data search unit 13, and the image generation unit 14.
[0066] (Embodiment 2) Next, a data extension device, a data extension method, and a program according to a second embodiment will be described with reference to FIGS.
[0067] [Device configuration] First, the configuration of the data expansion device according to the second embodiment will be described with reference to Fig. 8 and Fig. 9. Fig. 8 is a diagram showing the configuration of the data expansion device according to the second embodiment.
[0068] 8, a data extension device 40 according to the second embodiment is a device that extends training data for constructing a learning model that estimates a person's posture, similar to the data extension device 10 according to the first embodiment. Also, as shown in FIG. 8, the data extension device 40, like the data extension device 10, includes a data acquisition unit 11, a projection processing unit 12, a data search unit 13, and an image generation unit 14.
[0069] However, in the second embodiment, the data expansion device 40 includes a body shape modification unit 41 in addition to the above-mentioned configuration. In this respect, the data expansion device 40 in the second embodiment differs from the data expansion device 10 in the first embodiment. The following mainly describes the differences.
[0070] The body shape modification unit 41 modifies the 3D coordinates in the target data (3D pose data set) acquired by the data acquisition unit 11 so as to modify the body shape of a specific person. In the second embodiment, the body shape of a person in the target data can be modified to expand the data. This solves the problem of over-learning a specific body shape in the construction of a learning model for detecting 3D joint point coordinates, resulting in variations in detection accuracy depending on the body shape.
[0071] Specifically, body shape modification unit 41 modifies the three-dimensional coordinates in the acquired three-dimensional pose data set so that the vertical and horizontal change rates of the specific person satisfy the set conditions. Then, in the second embodiment, projection processing unit 12 projects each of the modified three-dimensional coordinates onto a two-dimensional plane.
[0072] The body shape modification process by the body shape modification unit 41 will be described with reference to Figure 9. Figure 9 is an explanatory diagram illustrating the body shape modification process in the second embodiment. Figure 9 shows an example in which the area between joint point 1 and joint point 2 is enlarged (or reduced). In the example of Figure 9, the body shape modification unit 41 determines the rate of change a in the vertical direction and the rate of change b in the horizontal direction so that the setting condition shown in the following equation 2 is satisfied. For example, a = (3 / 2) × α 1 / 2 , b=(2 / 3)×α 1 / 2 is set to
[0073]
number
[0074] In the above equation 2, "α" is set appropriately based on, for example, publicly available statistical information about people's body shapes. Also, "α" may be set appropriately through experiments so as to improve the detection accuracy of the learning model. Note that in the second embodiment, the setting conditions only need to be set so that the changed body shape does not appear unnatural, and are not limited to the example of equation 2 below.
[0075] [Device operation] Next, the operation of the data expansion device 40 in the second embodiment will be described with reference to Fig. 10. Fig. 10 is a flow diagram showing the operation of the data expansion device in the second embodiment. In the following description, Figs. 8 and 9 will be referred to as appropriate. Also, in the second embodiment, the data expansion method is implemented by operating the data expansion device 40. Therefore, the description of the data expansion method in the second embodiment will be replaced by the following description of the operation of the data expansion device 40.
[0076] As shown in FIG. 10, first, the data acquisition unit 11 acquires a three-dimensional pose data set of a specific person as target data (step B1).
[0077] Next, body shape modification unit 41 modifies the three-dimensional coordinates in the three-dimensional pose data set acquired in step B1 so that the vertical and horizontal change rates of the specific person satisfy the set conditions (step B2).
[0078] Next, the projection processing unit 12 projects each of the joint points 31 (see FIG. 3) included in a specific part of the 3D pose data set 30 after the change in step B2 onto the image coordinate system, and generates projected coordinates (two-dimensional coordinates) of each of the joint points 31 in the image coordinate system (step B3). Step B3 is the same as step A2 shown in FIG.
[0079] Next, the data search unit 13 reads out group data from the database 20, and for each group data, uses the internal parameters to identify the corresponding two-dimensional coordinates in the image coordinate system for the three-dimensional coordinates of each joint point included in a specific part of the three-dimensional pose data set of the group data (step B4). Step B4 is the same step as step A3 shown in Figure 7.
[0080] Next, the data search unit 13 manipulates the 3D pose data set acquired in step B1 or the 3D pose data set of the group data so that the set of projective coordinates generated in step B3 overlaps with the set of 2D coordinates identified in step B4 (step B5). Step B5 is the same as step A4 shown in FIG. 7.
[0081] Next, the data search unit 13 calculates the similarity between the 3D pose data set of the target data and the 3D pose data set of the group data after the operation of step B5 for each group data (step B6). Step B6 is a step similar to step A5 shown in FIG.
[0082] Next, the data search unit 13 identifies the set of data with the highest similarity based on the similarity calculated for each set of data in step B6 (step B7). Step B7 is the same as step A6 shown in FIG.
[0083] Next, the image generating unit 14 extracts a specific portion (e.g., left leg, right leg, right arm, etc.) from the two-dimensional image of the group data identified in step B7, and generates a patch image (step B8). Step B8 is the same as step A7 shown in FIG.
[0084] Thereafter, the image generator 14 generates a new two-dimensional image using the patch image generated in step B8, and further generates new training data using this image (step B9). Step B9 is the same as step A8 shown in FIG.
[0085] In this way, in the second embodiment, it is possible to change the body shape represented by the 3D pose data set in the target data. The second embodiment is useful for preventing a situation in which a specific body shape is over-learned in a learning model. Also, in the second embodiment, as in the first embodiment, it is possible to avoid a situation in which training data is expanded using 2D images of people with different postures in real space. In the second embodiment, it is also possible to expand training data while solving the conventional problems in building a learning model for detecting 3D joint point coordinates.
[0086] [program] The program in the second embodiment may be any program that causes a computer to execute steps B1 to B9 shown in Fig. 10. By installing and executing this program in a computer, the data extension device 40 and the data extension method in the second embodiment can be realized. In this case, the processor of the computer functions as a data acquisition unit 11, a projection processing unit 12, a data search unit 13, an image generation unit 14, and a body shape modification unit 41 to perform processing. Examples of the computer include a general-purpose PC, a smartphone, and a tablet terminal device.
[0087] In the second embodiment, the database 20 may be realized by storing the data files that constitute it in a storage device such as a hard disk provided in the computer, or may be realized by a storage device of another computer.
[0088] The program in the second embodiment may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as one of the data acquisition unit 11, the projection processing unit 12, the data search unit 13, the image generation unit 14, and the body shape modification unit 41.
[0089] [Physical configuration] Here, a computer that realizes the data extension device by executing the program according to the first and second embodiments will be described with reference to Fig. 11. Fig. 11 is a block diagram showing an example of a computer that realizes the data extension device according to the first and second embodiments.
[0090] 11, a computer 110 includes a CPU (Central Processing Unit) 111, a main memory 112, a storage device 113, an input interface 114, a display controller 115, a data reader / writer 116, and a communication interface 117. These components are connected to each other via a bus 121 so as to be able to communicate data with each other.
[0091] Furthermore, the computer 110 may include a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array) in addition to or instead of the CPU 111. In this aspect, the GPU or FPGA can execute the programs in the embodiments.
[0092] The CPU 111 loads a program in the embodiment, which is composed of a group of codes and stored in the storage device 113, into the main memory 112 and executes each code in a predetermined order to perform various calculations. The main memory 112 is typically a volatile storage device such as a DRAM (Dynamic Random Access Memory).
[0093] The program in the embodiment is provided in a state stored in a computer-readable recording medium 120. The program in the embodiment may be distributed over the Internet connected via the communication interface 117.
[0094] Specific examples of the storage device 113 include a hard disk drive and a semiconductor storage device such as a flash memory. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and a mouse. The display controller 115 is connected to a display device 119 and controls the display on the display device 119.
[0095] The data reader / writer 116 mediates data transmission between the CPU 111 and the recording medium 120, reads programs from the recording medium 120, and writes processing results from the computer 110 to the recording medium 120. The communication interface 117 mediates data transmission between the CPU 111 and other computers.
[0096] Specific examples of the recording medium 120 include general-purpose semiconductor storage devices such as CF (Compact Flash (registered trademark)) and SD (Secure Digital), magnetic recording media such as flexible disks, or optical recording media such as CD-ROMs (Compact Disk Read Only Memory).
[0097] The data extension device in this embodiment can be realized by using hardware (for example, electronic circuits) corresponding to each part, instead of a computer with a program installed. Furthermore, the data extension device may be partially realized by a program and the remaining part by hardware.
[0098] Some or all of the above-described embodiments can be expressed by (Supplementary Note 1) to (Supplementary Note 18) described below, but are not limited to the following descriptions.
[0099] (Appendix 1) a data acquisition unit that acquires data including a set of three-dimensional coordinates of each of the joint points of a particular person; a projection processing unit that projects each of the three-dimensional coordinates included in the acquired data onto a two-dimensional plane to generate projected coordinates of each of the joint points; for each set of data in which a set of three-dimensional coordinates of each joint point of a person, a two-dimensional image of the person, and camera parameters are associated with each other, using the camera parameters to identify two-dimensional coordinates on the two-dimensional image corresponding to each of the three-dimensional coordinates of the set of data; Furthermore, for each group data, the acquired data or a set of three-dimensional coordinates included in the group data is manipulated so that the generated set of projected coordinates overlaps with the identified set of two-dimensional coordinates, and after the manipulation, a similarity between the set of three-dimensional coordinates included in the acquired data and the set of three-dimensional coordinates included in the group data is calculated; a data search unit that identifies the group data corresponding to the acquired data based on the similarity calculated for each group data; an image generating unit that combines a part or all of the two-dimensional image of the identified set of data with another two-dimensional image to generate a new two-dimensional image; Equipped with A data expansion device characterized by:
[0100] (Appendix 2) 10. A data extension device according to claim 1, The data search unit For each set of data, one or a combination of translation, rotation, enlargement, and reduction is performed on the acquired data or the set of three-dimensional coordinates included in the set of data set so that two or more joint points included in the generated set of projection coordinates coincide with two or more joint points included in the identified set of two-dimensional coordinates. A data expansion device characterized by:
[0101] (Appendix 3) 3. A data extension device according to claim 1 or 2, a body shape modification unit that modifies the set of three-dimensional coordinates in the acquired data so that the body shape of the specific person is modified; the projection processing unit projects each of the changed three-dimensional coordinates onto the two-dimensional plane. A data expansion device characterized by:
[0102] (Appendix 4) 4. A data extension device according to claim 3, the body shape modification unit modifies the set of three-dimensional coordinates in the acquired data so that a vertical change rate and a horizontal change rate of the specific person satisfy a set condition. A data expansion device characterized by:
[0103] (Appendix 5) A data extension device according to any one of Supplementary Notes 1 to 4, the projection processing unit generates the projection coordinates from the three-dimensional coordinates of a specific part in the acquired data, the data search unit identifies the two-dimensional coordinates for the specific portion of the group data; the image generation unit cuts out an image of the specific part as a patch image from the two-dimensional image of the identified set of data, and synthesizes the cut-out patch image with another two-dimensional image to generate a new two-dimensional image. A data expansion device characterized by:
[0104] (Appendix 6) A data extension device according to any one of Supplementary Notes 1 to 5, the data search unit obtains a unit vector directed from a specific joint point to another joint point in the three-dimensional coordinates after the operation and a unit vector directed from a specific joint point to another joint point in the three-dimensional coordinates of the paired data, and calculates the similarity based on both the obtained unit vectors. A data expansion device characterized by:
[0105] (Appendix 7) a data acquisition step of acquiring data including a set of three-dimensional coordinates of each of the particular person's joint points; a projection processing step of projecting each of the three-dimensional coordinates included in the acquired data onto a two-dimensional plane to generate projected coordinates of each of the joint points; for each set of data in which a set of three-dimensional coordinates of each joint point of a person, a two-dimensional image of the person, and camera parameters are associated with each other, using the camera parameters to identify two-dimensional coordinates on the two-dimensional image corresponding to each of the three-dimensional coordinates of the set of data; Furthermore, for each group data, the acquired data or a set of three-dimensional coordinates included in the group data is manipulated so that the generated set of projected coordinates overlaps with the identified set of two-dimensional coordinates, and after the manipulation, a similarity between the set of three-dimensional coordinates included in the acquired data and the set of three-dimensional coordinates included in the group data is calculated; a data search step of identifying the group data corresponding to the acquired data based on the similarity calculated for each group data; an image generating step of synthesizing a part or all of the two-dimensional image of the identified set of data with another two-dimensional image to generate a new two-dimensional image; having A data augmentation method comprising:
[0106] (Appendix 8) 8. The data extension method of claim 7, further comprising: In the data searching step, For each set of data, one or a combination of translation, rotation, enlargement, and reduction is performed on the acquired data or the set of three-dimensional coordinates included in the set of data set so that two or more joint points included in the generated set of projection coordinates coincide with two or more joint points included in the identified set of two-dimensional coordinates. A data augmentation method comprising:
[0107] (Appendix 9) 9. The data extension method according to claim 7 or 8, a body shape modification step of modifying the set of three-dimensional coordinates in the acquired data so that the body shape of the specific person is modified; In the projection processing step, each of the three-dimensional coordinates after the change is projected onto the two-dimensional plane. A data augmentation method comprising:
[0108] (Appendix 10) 10. The data extension method of claim 9, further comprising: In the body shape modification step, the set of three-dimensional coordinates in the acquired data is modified so that a vertical change rate and a horizontal change rate of the specific person satisfy a set condition. A data augmentation method comprising:
[0109] (Appendix 11) A data extension method according to any one of Supplementary Notes 7 to 10, In the projection processing step, the projection coordinates are generated from the three-dimensional coordinates of a specific part in the acquired data; In the data search step, the two-dimensional coordinates of the specific portion of the group data are identified; In the image generating step, an image of the specific part is cut out as a patch image from the two-dimensional image of the identified set of data, and the cut-out patch image is synthesized with another two-dimensional image to generate a new two-dimensional image. A data augmentation method comprising:
[0110] (Appendix 12) A data extension method according to any one of Supplementary Notes 7 to 11, in the data searching step, a unit vector directed from a specific joint point to another joint point in the three-dimensional coordinates after the operation and a unit vector directed from a specific joint point to another joint point in the three-dimensional coordinates of the paired data are obtained, and the similarity is calculated based on both the obtained unit vectors. A data augmentation method comprising:
[0111] (Appendix 13) On the computer, a data acquisition step of acquiring data including a set of three-dimensional coordinates of each of the particular person's joint points; a projection processing step of projecting each of the three-dimensional coordinates included in the acquired data onto a two-dimensional plane to generate projected coordinates of each of the joint points; for each set of data in which a set of three-dimensional coordinates of each joint point of a person, a two-dimensional image of the person, and camera parameters are associated with each other, using the camera parameters to identify two-dimensional coordinates on the two-dimensional image corresponding to each of the three-dimensional coordinates of the set of data; Furthermore, for each group data, the acquired data or a set of three-dimensional coordinates included in the group data is manipulated so that the generated set of projected coordinates overlaps with the identified set of two-dimensional coordinates, and after the manipulation, a similarity between the set of three-dimensional coordinates included in the acquired data and the set of three-dimensional coordinates included in the group data is calculated; a data search step of identifying the group data corresponding to the acquired data based on the similarity calculated for each group data; an image generating step of synthesizing a part or all of the two-dimensional image of the identified set of data with another two-dimensional image to generate a new two-dimensional image; Run Ru, Program Hmm.
[0112] (Appendix 14) As described in Appendix 13 program And, In the data searching step, For each set of data, one or a combination of translation, rotation, enlargement, and reduction is performed on the acquired data or the set of three-dimensional coordinates included in the set of data set so that two or more joint points included in the generated set of projection coordinates coincide with two or more joint points included in the identified set of two-dimensional coordinates. Characterized by program .
[0113] (Appendix 15) Supplementary Note 13 or 14 program And, The computer, a body shape modification step of modifying the set of three-dimensional coordinates in the acquired data so that the body shape of the specific person is modified; Furthermore Executed height, In the projection processing step, each of the three-dimensional coordinates after the change is projected onto the two-dimensional plane. Characterized by program .
[0114] (Appendix 16) As described in Appendix 15 program And, In the body shape modification step, the set of three-dimensional coordinates in the acquired data is modified so that a vertical change rate and a horizontal change rate of the specific person satisfy a set condition. Characterized by program .
[0115] (Appendix 17) Any of Supplementary Notes 13 to 16 program And, In the projection processing step, the projection coordinates are generated from the three-dimensional coordinates of a specific part in the acquired data; In the data search step, the two-dimensional coordinates of the specific portion of the group data are identified; In the image generating step, an image of the specific part is cut out as a patch image from the two-dimensional image of the identified set of data, and the cut-out patch image is synthesized with another two-dimensional image to generate a new two-dimensional image. Characterized by program .
[0116] (Appendix 18) Any of Supplementary Notes 13 to 17 program And, in the data searching step, a unit vector directed from a specific joint point to another joint point in the three-dimensional coordinates after the operation and a unit vector directed from a specific joint point to another joint point in the three-dimensional coordinates of the paired data are obtained, and the similarity is calculated based on both the obtained unit vectors. Characterized by program .
[0117] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. [Industrial Applicability]
[0118] As described above, the present invention makes it possible to expand the training data when constructing a learning model for detecting three-dimensional joint point coordinates. The present invention is useful for various systems that estimate human poses from images. [Explanation of symbols]
[0119] 10 Data expansion device (first embodiment) 11 Data Acquisition Section 12 Projection processing section 13 Data Search Section 14 Image generation unit 20 databases 30 3D pose datasets 31 Articulation Points 32, 33, 34, 35 patch images 36 background images 40 Data expansion device (embodiment 2) 41 Body Shape Change Section 110 Computer 111 CPU 112 main memory 113 Storage device 114 Input Interface 115 Display Controller 116 Data Reader / Writer 117 Communication Interface 118 Input Devices 119 Display Device 120 Recording Media 121 Bus
Claims
1. a data acquisition unit that acquires data including a set of three-dimensional coordinates of each of the specific person's joint points; a projection processing unit that projects each of the three-dimensional coordinates included in the acquired data onto a two-dimensional plane to generate projected coordinates of each of the joint points; for each set of data in which a set of three-dimensional coordinates of each joint point of the person, a two-dimensional image of the person, and camera parameters are associated with each other, using the camera parameters to identify two-dimensional coordinates on the two-dimensional image corresponding to each of the three-dimensional coordinates of the set of data; Furthermore, for each group data, the acquired data or a set of three-dimensional coordinates included in the group data is manipulated so that the generated set of projected coordinates overlaps with the identified set of two-dimensional coordinates, and after the manipulation, a similarity between the set of three-dimensional coordinates included in the acquired data and the set of three-dimensional coordinates included in the group data is calculated; a data search unit that identifies the group data corresponding to the acquired data based on the similarity calculated for each group data; an image generating unit that combines a part or all of the two-dimensional image of the identified set of data with another two-dimensional image to generate a new two-dimensional image; Equipped with A data expansion device characterized by:
2. 2. The data expansion device according to claim 1, The data search unit performing, as the operation, one or a combination of translation, rotation, enlargement, and reduction on the acquired data or the set of three-dimensional coordinates included in the set of group data, so that two or more joint points included in the generated set of projected coordinates coincide with two or more joint points included in the identified set of two-dimensional coordinates; A data expansion device characterized by:
3. 2. The data expansion device according to claim 1, a body shape modification unit that modifies the set of three-dimensional coordinates in the acquired data so that the body shape of the specific person is modified; the projection processing unit projects each of the changed three-dimensional coordinates onto the two-dimensional plane. A data expansion device characterized by:
4. 4. The data expansion device according to claim 3, the body shape modification unit modifies the set of three-dimensional coordinates in the acquired data so that a vertical change rate and a horizontal change rate of the specific person satisfy a set condition. A data expansion device characterized by:
5. 2. The data expansion device according to claim 1, the projection processing unit generates the projection coordinates from the three-dimensional coordinates of a specific part in the acquired data, the data search unit identifies the two-dimensional coordinates for the specific portion of the group data; the image generation unit cuts out an image of the specific part as a patch image from the two-dimensional image of the identified set of data, and combines the cut-out patch image with another two-dimensional image to generate a new two-dimensional image. A data expansion device characterized by:
6. 2. The data expansion device according to claim 1, the data search unit obtains a unit vector directed from a specific joint point to another joint point in the three-dimensional coordinates after the operation and a unit vector directed from a specific joint point to another joint point in the three-dimensional coordinates of the paired data, and calculates the similarity based on both the obtained unit vectors. A data expansion device characterized by:
7. obtaining data including a set of three-dimensional coordinates of each of the joint points of a particular person; projecting each of the three-dimensional coordinates included in the acquired data onto a two-dimensional plane to generate projected coordinates of each of the joint points; for each set of data in which a set of three-dimensional coordinates of each joint point of the person, a two-dimensional image of the person, and camera parameters are associated with each other, using the camera parameters to identify two-dimensional coordinates on the two-dimensional image corresponding to each of the three-dimensional coordinates of the set of data; Furthermore, for each group data, the acquired data or a set of three-dimensional coordinates included in the group data is manipulated so that the generated set of projected coordinates overlaps with the identified set of two-dimensional coordinates, and after the manipulation, a similarity between the set of three-dimensional coordinates included in the acquired data and the set of three-dimensional coordinates included in the group data is calculated; Identifying the group data corresponding to the acquired data based on the similarity calculated for each group data; synthesizing a part or all of the two-dimensional image of the identified set of data with another two-dimensional image to generate a new two-dimensional image; A data augmentation method comprising:
8. On the computer, obtaining data including a set of three-dimensional coordinates of each of the joint points of a particular person; projecting each of the three-dimensional coordinates included in the acquired data onto a two-dimensional plane to generate projected coordinates of each of the joint points; for each set of data in which a set of three-dimensional coordinates of each joint point of a person, a two-dimensional image of the person, and camera parameters are associated with each other, using the camera parameters to identify two-dimensional coordinates on the two-dimensional image corresponding to each of the three-dimensional coordinates of the set of data; Furthermore, for each group data, the acquired data or the group data set of three-dimensional coordinates included in the group data is manipulated so that the generated group of projected coordinates overlaps with the identified group of two-dimensional coordinates, and after the manipulation, a similarity between the group of three-dimensional coordinates included in the acquired data and the group data set of three-dimensional coordinates is calculated; Identifying the group data corresponding to the acquired data based on the degree of similarity calculated for each group data; a part or all of the two-dimensional image of the identified set of data is synthesized with another two-dimensional image to generate a new two-dimensional image; program.
Citation Information
Patent Citations
Posture correction network study device and program thereof and posture estimation device and program thereof
JP2021047563A
Machine learning systems and methods of estimating body shape from images
US10679046B1