Joints detection device, joints detection method, and program
The joint point detection device and method enhance the accuracy of estimating human posture from images by using a machine learning model to learn the positional relationships between joint points, even when some joint points are absent from the image.
Patent Information
- Application Number
- JP2023502225
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-26
- Filing Date
- 2022-02-01
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-02-01
AI Technical Summary
Existing systems for estimating human posture from images, such as those described in Non-Patent Document 1 and Patent Document 1, face a decrease in estimation accuracy due to the absence of certain joint points in images, leading to incorrect positioning of visible joint points.
A joint point detection device and method that utilize a machine learning model to output feature amounts representing joint points from image data, and generate training feature amounts by simulating the absence of specific joint point feature amounts, allowing the model to learn the positional relationships between other joint points in such scenarios.
This approach significantly improves the estimation accuracy of joint point positions, even when certain joint points are not visible in the image, by refining the positioning of visible joint points based on learned relationships.
Smart Images

Figure 0007687382000001 
Figure 0007687382000002 
Figure 0007687382000003
Abstract
Description
Technical Field
[0001] The present invention relates to a joint point detection device and a joint point detection method for detecting joint points of a living body from an image, and further to a program for realizing these. to the mu The present invention also relates to a learning model generation device and a learning model generation method for generating a learning model for detecting joint points of a living body from an image, and further to a program for realizing these. to the mu It relates to.
Background Art
[0002] In recent years, systems for estimating human postures from images have been proposed. Such systems are expected to be used in fields such as video surveillance and user interfaces. For example, in an image surveillance system, if the human posture can be estimated, it is possible to estimate what the person captured by the camera is doing, thus improving the surveillance accuracy. Also, in a user interface, if the human posture can be estimated, input by gestures becomes possible.
[0003] For example, Non-Patent Document 1 discloses a system for estimating a human posture, particularly the posture of a human hand, from an image. The system disclosed in Non-Patent Document 1 first acquires image data including an image of a hand, and then inputs the acquired image data into a neural network that has machine-learned image features for each joint point, and outputs a heat map that expresses the probability of the existence of a joint point for each joint point in terms of color and density.
[0004] Subsequently, the system disclosed in Non-Patent Document 1 inputs the output heat map into a neural network that has machine-learned the relationship between the heat map corresponding to the joint point. Also, a plurality of such neural networks are prepared, and the output result from one neural network is input into another neural network. As a result, the position of the joint point on the heat map is refined.
[0005] In addition, Patent Document 1 also discloses a system for estimating hand postures from images. Similar to the system disclosed in Non-Patent Document 1, the system disclosed in Patent Document 1 also uses a neural network to estimate the coordinates of joint points.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Non-Patent Documents
[0007]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] If the systems disclosed in Non-Patent Document 1 or Patent Document 1 are used, as described above, the coordinates of the joint points of a person's hand can be estimated from an image. However, these systems have the following problem that the estimation accuracy decreases.
[0009] First, a living body has many joint points, and in an image, some joint points may not be shown. In such a case, in the systems disclosed in Non-Patent Document 1 and Patent Document 1, the position of a joint point not shown in the image in the heat map may be an incorrect position. As a result, when the position of each joint point is refined by a neural network, it is dragged to the incorrect position of the joint point not shown in the image, and even the position of the joint point shown in the image becomes an incorrect position.
[0010] An example of the object of the present invention is to provide a joint point detection device, a learning model generation device, a joint point detection method, a learning model generation method, and a program that can improve the estimation accuracy of the position of joint points.
Means for Solving the Problems
[0011] To achieve the above object, a joint point detection device according to an aspect of the present invention includes a total feature amount output unit that outputs a first feature amount representing the joint point for each of the target joint points from the target image data, and a partial feature amount output unit that outputs a second feature amount representing the joint point for each of the target joint points by using a machine learning model that machine-learns the positional relationship between other joint points when the feature amount of a specific joint point does not exist, with the first feature amount for each of the target joint points as an input. It is characterized by comprising the above.
[0012] To achieve the above object, a learning model generation device according to an aspect of the present invention includes a total feature amount output unit that outputs a feature amount representing the joint point for each of the target joint points from the target image data, and a feature amount generation unit that generates, as a training feature amount, a feature amount when the feature amount of a specific joint point does not exist from the feature amounts for each of the target joint points. A learning model generation unit that generates a machine learning model by performing machine learning on the positional relationship between other joint points when the feature amount of the specific joint point does not exist, using the training data including the generated training feature amounts. It is characterized by comprising the following.
[0013] To achieve the above object, a joint point detection method according to one aspect of the present invention A full feature amount output step of outputting, for each of the target joint points, a first feature amount representing the joint point from the target image data. A partial feature amount output step of outputting, for each of the target joint points, a second feature amount representing the joint point, using a machine learning model that performs machine learning on the positional relationship between other joint points when the feature amount of a specific joint point does not exist, with the first feature amount for each of the target joint points as an input. It is characterized by having the following.
[0014] To achieve the above object, a learning model generation method according to one aspect of the present invention A full feature amount output step of outputting, for each of the target joint points, a feature amount representing the joint point from the target image data. A feature amount generation step of generating, as training feature amounts, feature amounts when the feature amount of a specific joint point does not exist, from the feature amounts for each of the target joint points. A learning model generation step of generating a machine learning model by performing machine learning on the positional relationship between other joint points when the feature amount of the specific joint point does not exist, using the training data including the generated training feature amounts. It is characterized by having the following.
[0015] To achieve the above object, the first one in one aspect of the present invention program is Causing a computer to A full feature amount output step of outputting, for each of the target joint points, a first feature amount representing the joint point from the target image data. Using, as input, the first feature amount for each of the target joint points, a machine learning model that machine-learns the positional relationship between other joint points when the feature amount of a specific joint point does not exist, outputs, for each of the target joint points, a second feature amount representing the joint point, a partial feature amount output step; causing to execute 、 characterized in that.
[0016] To achieve the above object, a second program in one aspect of the present invention causes a computer to a total feature amount output step of outputting, for each of the target joint points, a feature amount representing the joint point from the target image data; a feature amount generation step of generating, as training feature amounts, feature amounts when the feature amount of a specific joint point does not exist from the feature amounts for each of the target joint points; a learning model generation step of generating a machine learning model by machine-learning the positional relationship between other joint points when the feature amount of the specific joint point does not exist using training data including the generated training feature amounts; causing to execute 、 characterized in that.
Advantages of the Invention
[0017] As described above, according to the present invention, it is possible to improve the estimation accuracy of the position of the joint point.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
BEST MODE FOR CARRYING OUT THE INVENTION
[0019] (Embodiment 1) First, in Embodiment 1, the learning model generation apparatus, the learning model generation method, and the program for generating the learning model will be described with reference to FIGS. 1 to 5.
[0020] [Apparatus Configuration] First, the schematic configuration of the learning model generation apparatus in Embodiment 1 will be described with reference to FIG. 1. FIG. 1 is a configuration diagram showing the schematic configuration of the learning model generation apparatus in Embodiment 1.
[0021] The learning model generation apparatus 10 in Embodiment 1 shown in FIG. 1 is an apparatus that generates a machine learning model for detecting joint points. As shown in FIG. 1, the learning model generation apparatus 10 includes an all feature quantity output unit 11, a feature quantity generation unit 12, and a learning model generation unit 13.
[0022] The full feature amount output unit 11 outputs, for each target joint point, a feature amount representing the joint point from the target image data. The feature amount generation unit 12 generates, as training feature amounts, feature amounts in the case where the feature amount of a specific joint point does not exist, from the feature amounts for each of the target joint points.
[0023] The learning model generation unit 13 generates a machine learning model by performing machine learning on the positional relationship between other joint points in the case where the feature amount of a specific joint point does not exist, using the training data including the generated training feature amounts.
[0024] As described above, in the first embodiment, the training feature amounts used as the training data are the feature amounts in the case where it is set that the feature amount of a specific joint point does not exist. Therefore, if joint points are detected using the generated machine learning model, it is possible to accurately estimate the target joint points even when the specific target joint points are not shown in the image.
[0025] Subsequently, with reference to FIGS. 2 to 4, the configuration and functions of the learning model generation apparatus 10 in the first embodiment will be specifically described. FIG. 2 is a block diagram specifically showing the configuration of the learning model generation apparatus in the first embodiment. FIG. 3 is a diagram for explaining the function of the full feature amount output unit in the first embodiment. FIG. 4 is a diagram for explaining the functions of the feature amount generation unit and the learning model generation unit in the first embodiment.
[0026] As shown in FIG. 2, in the embodiment, the learning model generation apparatus 10 includes, in addition to the above-described full feature amount output unit 11, feature amount generation unit 12, and learning model generation unit 13, a random number generation unit 14 and a storage unit 15.
[0027] The random number generation unit 14 generates a random number within a set range and inputs the generated random number to the feature amount generation unit 12. The processing in the feature amount generation unit 12 using the random number will be described later. The storage unit 15 stores the machine learning model 16 generated by the learning model generation unit 13.
[0028] Also, in the embodiment, the machine learning model 16 is constructed by a convolutional neural network (CNN: Convolutional Neural Network). In the embodiment, the generation of the machine learning model by the learning model generation unit 13 is performed by updating the initial values of the parameters of the CNN by learning. Hereinafter, the machine learning model is also referred to as "CNN".
[0029] Also, hereinafter, the case where the object is a human hand will be described as an example. Note that in the first embodiment, the object is not limited to a human hand, and may be the entire human body or other parts. The object may be anything having joint points, and may be something other than a human, for example, a robot. Further, in the first embodiment, in addition to the joint points, parts other than the joint points, for example, characteristic parts such as fingertips, may also be the detection targets.
[0030] In addition, in the first embodiment, it is assumed that a heatmap is used as a feature amount. The heatmap is a map that represents the possibility of the existence of joint points on an image, and for example, the possibility of the existence of joint points can be represented by the shade of color. Note that other things than the heatmap, for example, coordinate values, may be used as the feature amount.
[0031] The all feature amount output unit 11 first acquires the image data 20 of the object in the first embodiment. Then, as shown in FIG. 3, the all feature amount output unit 11 outputs a heatmap 21 as a feature amount representing the joint points from the image data 20. In the example of FIG. 3, a plurality of heatmaps 21 are output for each joint point on the image data 20.
[0032] Specifically, the all feature amount output unit 11For example, by using a machine learning model that learns the relationship between joint points and heatmaps on an image, and inputting image data into this machine learning model, a heatmap 21 is output. As the machine learning model in this case, a CNN can also be mentioned. Further, in this CNN machine learning, the image data of the joint points and the correct heatmap serve as training data. And the CNN machine learning is performed by updating the parameters so that the difference between the output result (heatmap) of the image data serving as training data and the correct heatmap becomes small.
[0033] In the embodiment, the feature quantity generation unit 12 generates, as a training feature quantity set 22, a set of feature quantities in which only the feature quantities of specific joint points are set to not exist, for each of the plurality of specific joint points, from the heatmaps 21 for each of the target joint points.
[0034] Specifically, as shown in FIG. 4, the feature quantity generation unit 12 first receives a random number from the random number generation unit 14. Then, among the plurality of heatmaps 21 generated for each joint point on the image, the feature quantity of the j-th joint point indicated by the random number is set to not exist by setting the data on the heatmap of the j-th joint point to zero or one. Thereby, a set of feature quantities (training feature quantity set) 22 in which only the heatmap of the j-th joint point among the plurality of heatmaps 21 generated for each joint point on the image is set to not exist is generated. It is assumed that numbers are assigned to each joint point in advance.
[0035] Also, in the example of FIG. 4, the feature quantities of each of the plurality of joint points are set to not exist according to the generated random number, but it is not limited to this, and the joint points set to not have feature quantities may be set in advance. Further, the feature quantity generation unit 12 may set, for each of all the joint points in order, that the feature quantity does not exist, and generate the training feature quantity set 22 as many times as the number of joint points. In the example of FIG. 4, since the training feature quantity is also a heatmap, the training feature quantity set 22 is denoted as the "training heatmap set 22".
[0036] In the embodiment, the learning model generation unit 13 uses training data including a corresponding set of training heatmaps for each of a plurality of specific joint points to perform machine learning on the positional relationship between other joint points when the heatmap of a specific joint point does not exist, and generates a machine learning model.
[0037] Specifically, as shown in FIG. 4, the learning model generation unit 13 acquires the CNN 16 from the storage unit 15, inputs the selected set of training heatmaps 22 into the CNN 16, and calculates the difference between each heatmap set as the output result and the correct heatmap corresponding thereto. Note that the correct heatmap is prepared in advance. Also, for a heatmap for which it is assumed that no feature amount exists, the difference may not be calculated, or a heatmap having no feature amount may be used as the correct heatmap.
[0038] Then, the learning model generation unit 13 updates the parameters of the CNN 16 so that the calculated difference becomes minimum, and stores the CNN 16 with updated parameters in the storage unit 15. Also, the learning model generation unit 13 executes this process until there is no unselected set of training heatmaps 22. As a result, a CNN that can be used for detecting joint points is generated.
[0039] [Device Operation] Next, the operation of the learning model generation device 10 in Embodiment 1 will be described with reference to FIG. 5. FIG. 5 is a flowchart showing the operation of the learning model generation device in Embodiment 1. In the following description, FIGS. 1 to 4 will be referred to as appropriate. Also, in Embodiment 1, by operating the learning model generation device 10, the learning model generation method is implemented. Therefore, the description of the learning model generation method in Embodiment 1 will be replaced with the following description of the operation of the learning model generation device 10.
[0040] As shown in FIG. 5, first, the all-feature output unit 11 acquires the target image data 20, and outputs a heatmap 21 as a feature amount representing a joint point from the acquired image data 20 (step A1).
[0041] Next, the feature quantity generation unit 12 acquires the random number j generated by the random number generation unit 14 (step A2). Subsequently, the feature quantity generation unit 12 sets, for each joint point output in step A1, that only the feature quantity of the j-th joint point does not exist, that is, generates a set of heatmaps in which only the heatmap of the j-th joint point is set to zero (or 1) as the training heatmap set 22 (step A3).
[0042] Next, the feature quantity generation unit 12 determines whether a predetermined number of training heatmap sets 22 have been generated (step A4). And as a result of the determination in step A4, if a predetermined number of training heatmap sets 22 have not been generated (step A4: No), the feature quantity generation unit 12 again executes step A2.
[0043] On the other hand, as a result of the determination in step A4, if a predetermined number of training heatmap sets 22 have been generated (step A4: Yes), the feature quantity generation unit 12 notifies the learning model generation unit 13 that the generation of the training heatmap set 22 has been completed.
[0044] Upon receiving the notification, the learning model generation unit 13 updates the parameters of the CNN 16 using the predetermined number of training heatmap sets 22 generated in step A3 (step A5). Thereby, the positional relationship between other joint points when the heatmap of a specific joint point does not exist is machine-learned, and a machine learning model is generated. After the execution of step A5, the process for generating the learning model ends.
[0045] As described above, in the first embodiment, the training heatmap set used as training data represents the feature quantity when the feature quantity of a specific joint point does not exist. Therefore, if the joint points are detected as described below using the generated CNN 16, it is possible to accurately estimate the target joint points even when the target specific joint points are not shown in the image.
[0046] [Program] The program for generating the learning model in Embodiment 1 may be any program that causes a computer to execute steps A1 to A5 shown in FIG. 5. By installing and executing this program on a computer, the learning model generation device and the learning model generation method in Embodiment 1 can be realized. In this case, the processor of the computer functions as the all-feature quantity output unit 11, the feature quantity generation unit 12, the learning model generation unit 13, and the random number generation unit 14, and performs processing.
[0047] Also, in Embodiment 1, the storage unit 15 may be realized by storing the data files constituting these in a storage device such as a hard disk provided in the computer, or may be realized by a storage device of another computer. Examples of the computer include smartphones and tablet terminal devices in addition to general-purpose PCs.
[0048] The program for generating the learning model in Embodiment 1 may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as any one of the all-feature quantity output unit 11, the feature quantity generation unit 12, the learning model generation unit 13, and the random number generation unit 14.
[0049] (Embodiment 2) Subsequently, in Embodiment 2, the joint point detection device, the joint point detection method, and the program for joint point detection will be described with reference to FIGS. 6 to 9.
[0050] [Device Configuration] First, the schematic configuration of the joint point detection device in Embodiment 2 will be described with reference to FIG. 6. FIG. 6 is a configuration diagram showing the schematic configuration of the joint point detection device in Embodiment 2.
[0051] The joint point detection device 30 in Embodiment 2 shown in FIG. 6 is a device for detecting joint points of an object, for example, a living body, a robot, or the like. As shown in FIG. 6, the joint point detection device 30 includes a full feature amount output unit 31 and a partial feature amount output unit 32.
[0052] The full feature amount output unit 31 outputs a first feature amount representing a joint point for each joint point of the object from the image data of the object. The partial feature amount output unit 32 uses a machine learning model with the first feature amount for each joint point of the object as an input, and outputs a second feature amount representing a joint point for each joint point of the object. The machine learning model is a machine learning model that has learned the positional relationship between other joint points when there is no feature amount of a specific joint point.
[0053] As described above, in Embodiment 2, the second feature amount is output from the first feature amount representing each joint point using the machine learning model. Since the machine learning model has learned the positional relationship between joint points other than a specific joint point, the second feature amount can appropriately indicate the positions of other joint points when the specific joint point cannot be seen. Therefore, according to Embodiment 2, it is possible to improve the estimation accuracy of the position of each joint point.
[0054] Subsequently, with reference to FIGS. 7 and 8, the configuration and functions of the joint point detection device 30 in Embodiment 2 will be specifically described. FIG. 7 is a diagram more specifically showing the configuration of the joint point detection device in Embodiment 2. FIG. 8 is a diagram for explaining the functions of the full feature amount output unit and the partial feature amount output unit in Embodiment 2.
[0055] As shown in FIG. 7, in Embodiment 2, the joint point detection device 30 includes, in addition to the above-described full feature amount output unit 31 and partial feature amount output unit 32, a joint point detection unit 33 and a storage unit 34. The storage unit 34 stores the CNN16 shown in FIG. 2 in Embodiment 1.
[0056] In the second embodiment as well, the case where the object is a human hand will be described as an example. Note that in the second embodiment as well, the detection target of the joint points is not limited to a human hand, and may be the entire human body or other parts. Further, the detection target of the joint points may be anything having joint points, and may be something other than a human, for example, a robot. Furthermore, in the second embodiment as well, in addition to the joint points, parts other than the joint points, for example, characteristic parts such as fingertips, may also be detection targets.
[0057] In addition, assume that in the second embodiment, a heatmap is used as a feature amount. Note that in the second embodiment, something other than a heatmap, for example, coordinate values, may be used as the feature amount.
[0058] The all-feature amount output unit 31 has the same function as in the first embodiment. First, it acquires the image data 40 of the object. And the all-feature amount output unit 3 1 outputs, as a first feature amount representing the joint points, a first heatmap 41 from the image data 8 as shown in the figure. Also, in the example of FIG. 8, a plurality of first heatmaps 41 are output for each joint point on the image data 40. 4 Specifically, the all-feature amount output unit 31 also uses, in the same manner as the all-feature amount output unit 11 shown in the first embodiment, for example, a machine learning model that has learned the relationship between the joint points on the image and the heatmap, and inputs the image data into this machine learning model to output the first heatmap 41. The CNN can also be mentioned as the machine learning model in this case.
[0059]
[0060] In the second embodiment, the partial feature amount output unit 32 inputs the first heatmap 41 for each of the joint points of the object output from the all-feature amount output unit 31 into the CNN 16, and causes the CNN 16 to output a second heatmap 42 for each of the joint points of the object.
[0061] As described in Embodiment 1, CNN16 is a machine learning model that learns the positional relationship between other joints when there is no feature amount of a specific joint point. Therefore, in the second heatmap 42, the second feature amount as appropriately shows the positions of other joint points when a specific joint point cannot be seen.
[0062] The joint point detection unit 33 acquires the second heatmap 42 for each joint point of the target hand. Then, the joint point detection unit 33 uses the second heatmap 42 for each joint point to detect the coordinates of each target joint point.
[0063] Specifically, for each joint point, the joint point detection unit 33 specifies the location with the highest density in the second heatmap 42 and detects the two-dimensional coordinates of the specified location on the image. Also, when there are multiple second heatmaps 42 for each joint point, the joint point detection unit 33 42 specifies the two-dimensional coordinates of the location with the highest density for each second heatmap, further obtains the average of the specified two-dimensional coordinates, and sets the obtained average coordinates as the final coordinates.
[0064] [Device Operation] Next, the operation of the joint point detection device 30 in Embodiment 2 will be described with reference to FIG. 9. FIG. 9 is a flowchart showing the operation of the joint point detection device in Embodiment 2. In the following description, FIGS. 6 to 8 will be referred to as appropriate. Also, in Embodiment 2, by operating the joint point detection device 30, the joint point detection method is implemented. Therefore, the description of the joint point detection method in Embodiment 2 will be replaced with the following description of the operation of the joint point detection device 30.
[0065] As shown in FIG. 9, first, the all feature amount output unit 31 acquires the target image data 40 and outputs the first heatmap 41 as the feature amount representing the joint point from the acquired image data 40 (step B1).
[0066] Next, the partial feature amount output unit 32 inputs the step to CNN16 BInput the first heatmap 41 output at 1, and output a second heatmap 42 representing the joint points (step B2).
[0067] Next, the joint point detection unit 33 detects the coordinates of each target joint point from the second heatmap 42 of each joint point output in step B2 (step B3).
[0068] As described above, in Embodiment 2, the first heatmap 41 obtained from the image data is input to the CNN 16. Since the CNN 16 is learns the positional relationship between joint points other than specific joint points, the second heatmap 42 can appropriately indicate the positions of other joint points when the specific joint points are not visible. Therefore, according to Embodiment 2, the estimation accuracy of the position of the target joint point can be improved.
[0069] [Program] The program for detecting joint points in Embodiment 2 may be a program that causes a computer to execute steps B1 to B3 shown in FIG. 9. By installing and executing this program on a computer, the joint point detection device and the joint point detection method in Embodiment 2 can be realized. In this case, the processor of the computer functions as the full feature amount output unit 31, the partial feature amount output unit 32, and the joint point detection unit 33, and performs processing.
[0070] Also, in this embodiment, the storage unit 34 may be realized by storing data files constituting these in a storage device such as a hard disk provided in the computer, or may be realized by a storage device of another computer. Examples of the computer include general-purpose PCs, smartphones, and tablet terminal devices in addition to general-purpose PCs.
[0071] The program for joint point detection in Embodiment 2 may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as any one of the all feature quantity output unit 31, the partial feature quantity output unit 32, and the joint point detection unit 33, respectively.
[0072] (Physical Configuration) Here, a computer that realizes the learning model generation device 10 by executing the program in Embodiment 1 and a computer that realizes the joint point detection device 30 by executing the program in Embodiment 2 will be described with reference to FIG. 10. FIG. 10 is a block diagram showing an example of a computer that realizes the learning model generation device in Embodiment 1 and the joint point detection device in Embodiment 2.
[0073] As shown in FIG. 10, the computer 110 includes a CPU (Central Processing Unit) 111, a main memory 112, a storage device 113, an input interface 114, a display controller 115, a data reader / writer 116, and a communication interface 117. These units are connected to each other via a bus 121 so as to be capable of data communication.
[0074] In addition to or instead of the CPU 111, the computer 110 may include a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array). In this mode, the GPU or FPGA can execute the program in the embodiment.
[0075] The CPU 111 expands the program in the embodiment composed of a code group stored in the storage device 113 into the main memory 112 and executes each code in a predetermined order to perform various operations. The main memory 112 is typically a volatile storage device such as a DRAM (Dynamic Random Access Memory).
[0076] In addition, the program in the embodiment is provided in a state stored in a computer-readable recording medium 120. Note that the program in the present embodiment may also be distributed on the Internet connected via the communication interface 117.
[0077] As a specific example of the storage device 113, in addition to a hard disk drive, a semiconductor storage device such as a flash memory can be mentioned. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and a mouse. The display controller 115 is connected to the display device 119 and controls the display on the display device 119.
[0078] The data reader / writer 116 mediates data transmission between the CPU 111 and the recording medium 120, and executes reading of a program from the recording medium 120 and writing of a processing result in the computer 110 to the recording medium 120. The communication interface 117 mediates data transmission between the CPU 111 and other computers.
[0079] As specific examples of the recording medium 120, general-purpose semiconductor memory devices such as CF (Compact Flash (registered trademark)) and SD (Secure Digital), magnetic recording media such as a flexible disk, or optical recording media such as a CD-ROM (Compact Disk Read Only Memory) can be mentioned.
[0080] Note that the learning model generation device 10 and the joint detection device 30 can also be realized by using hardware corresponding to each part, for example, a circuit, instead of a computer in which a program is installed. Further, the learning model generation device 10 and the joint detection device 30 may be partially realized by a program and the remaining part may be realized by hardware.
[0081] Some or all of the above-described embodiments can be expressed by (Supplementary Note 1) to (Supplementary Note 21) described below, but are not limited to the following description.
[0082] (Supplementary Note 1) A full feature amount output unit that outputs a first feature amount representing each of the target joint points from the target image data, and A partial feature amount output unit that uses the first feature amount of each of the target joint points as an input and outputs a second feature amount representing each of the target joint points using a machine learning model that machine-learns the positional relationship between other joint points when the feature amount of a specific joint point does not exist. A joint point detection device characterized by comprising:
[0083] (Supplementary Note 2) The joint point detection device according to Supplementary Note 1, wherein further comprising a joint point detection unit that detects the coordinates of the target joint points using the second feature amount of each of the target joint points. A joint point detection device characterized by:
[0084] (Supplementary Note 3) The joint point detection device according to Supplementary Note 1 or 2, wherein the partial feature amount output unit uses the first feature amount of each of the target joint points as an input and outputs a second feature amount representing each of the target joint points using a machine learning model that machine-learns the positional relationship between other joint points when the feature amount of each of the plurality of specific joint points does not exist. A joint point detection device characterized by:
[0085] (Supplementary Note 4) The joint point detection device according to any one of Supplementary Notes 1 to 3, wherein the machine learning model is constructed by a convolutional neural network, The first feature amount and the second feature amount each include a heat map representing the possibility of the presence of joint points on the image. A joint point detection device characterized by this.
[0086] (Appendix 5) A full feature amount output unit that outputs a feature amount representing each of the target joint points from the target image data; A feature amount generation unit that generates, as training feature amounts, feature amounts when the feature amount of a specific joint point does not exist from the feature amounts for each of the target joint points; A learning model generation unit that generates a machine learning model by performing machine learning on the positional relationship between other joint points when the feature amount of the specific joint point does not exist, using the training data including the generated training feature amounts; A learning model generation device characterized by comprising these.
[0087] (Appendix 6) The learning model generation device according to Appendix 5, The feature amount generation unit generates, as a training feature amount set, a set of feature amounts when only the feature amount of a specific joint point does not exist, for each of a plurality of specific joint points, from the feature amounts for each of the target joint points; The learning model generation unit generates a machine learning model by performing machine learning on the positional relationship between other joint points when the feature amount of the specific joint point does not exist, using the training data including the corresponding training feature amount set for each of the plurality of specific joint points; A learning model generation device characterized by this.
[0088] (Appendix 7) The learning model generation device according to Appendix 5 or 6, The machine learning model is constructed by a convolutional neural network, The feature amount includes a heat map representing the possibility of the presence of joint points on the image, The feature quantity generation unit sets the data on the heat map to zero or 1, thereby setting that there is no feature quantity. A learning model generation device characterized by the above.
[0089] (Appendix 8) From the target image data, for each of the target joint points, an all-feature quantity output step of outputting a first feature quantity representing the joint point; Using the first feature quantity for each of the target joint points as an input, and using a machine learning model that machine-learns the positional relationship between other joint points when the feature quantity of a specific joint point does not exist, for each of the target joint points, a partial feature quantity output step of outputting a second feature quantity representing the joint point; A joint point detection method characterized by including the above.
[0090] (Appendix 9) The joint point detection method according to Appendix 8, Further including a joint point detection step of detecting the coordinates of the target joint points using the second feature quantity for each of the target joint points. A joint point detection method characterized by the above.
[0091] (Appendix 10) The joint point detection method according to Appendix 8 or 9, In the partial feature quantity output step, using the first feature quantity for each of the target joint points as an input, and using a machine learning model that machine-learns the positional relationship between other joint points when the feature quantity of each of a plurality of the specific joint points does not exist, for each of the target joint points, outputting a second feature quantity representing the joint point. A joint point detection method characterized by the above.
[0092] (Appendix 11) The joint point detection method according to any one of Appendices 8 to 10, The machine learning model is constructed by a convolutional neural network. wherein each of the first feature amount and the second feature amount includes a heat map representing the possibility of the presence of joint points on the image A joint point detection method characterized by the above
[0093] (Appendix 12) An all feature amount output step of outputting, for each of the joint points of the target, a feature amount representing the joint point from the target image data, A feature amount generation step of generating, from the feature amounts for each of the joint points of the target, a feature amount when the feature amount of a specific joint point does not exist as a training feature amount, A learning model generation step of generating a machine learning model by machine learning the positional relationship between other joint points when the feature amount of the specific joint point is zero, using the training data including the generated training feature amount, A learning model generation method characterized by having the above
[0094] (Appendix 13) The learning model generation method according to Appendix 12, wherein in the feature amount generation step, from the feature amounts for each of the joint points of the target, for each of a plurality of specific joint points, a set of feature amounts when only the feature amount of the specific joint point does not exist is generated as a training feature amount set, in the learning model generation step, for each of the plurality of specific joint points, a machine learning model is generated by machine learning the positional relationship between other joint points when the feature amount of the specific joint point is zero, using the training data including the corresponding training feature amount set, A learning model generation method characterized by the above
[0095] (Appendix 14) The learning model generation method according to Appendix 12 or 13, wherein the machine learning model is constructed by a convolutional neural network, the feature amount includes a heat map representing the possibility of the presence of joint points on the image, The feature quantity generation step sets the feature quantity to non - existent by setting the data on the heat map to zero or 1. A learning model generation method characterized by the above.
[0096] (Appendix 15) On a computer, From the target image data, for each of the target joint points, an all - feature quantity output step of outputting a first feature quantity representing the joint point, Using a machine learning model that machine - learns the positional relationship between other joint points when the feature quantity of a specific joint point does not exist, with the first feature quantity for each of the target joint points as input, for each of the target joint points, a partial - feature quantity output step of outputting a second feature quantity representing the joint point, Cause to execute there is Progra mu
[0097] (Appendix 16) As described in Appendix 15 program And On the computer, Furthermore Cause to execute a joint - point detection step of detecting the coordinates of the target joint points using the second feature quantity for each of the target joint points. there is Characterized by program .
[0098] (Appendix 17) As described in Appendix 15 or 16 program And In the partial - feature quantity output step, using a machine learning model that machine - learns the positional relationship between other joint points when the feature quantity of a specific joint point does not exist, with the first feature quantity for each of the target joint points as input, for each of the target joint points, output a second feature quantity representing the joint point. Characterized by program .
[0099] (Appendix 18) as described in any one of Appendices 15 to 17 program wherein the machine learning model is constructed by a convolutional neural network each of the first feature amount and the second feature amount includes a heat map representing the possibility of the existence of a joint point on the image characterized by program .
[0100] (Appendix 19) causing a computer to perform a full feature amount output step of outputting, for each of the joint points of the target, a feature amount representing the joint point from the target image data; a feature amount generation step of generating, as a training feature amount, a feature amount when the feature amount of a specific joint point is set to zero from the feature amounts of each of the joint points of the target; a learning model generation step of generating a machine learning model by machine learning the positional relationship between other joint points when the feature amount of the specific joint point does not exist, using training data including the generated training feature amount; and causing it to execute there is program mu
[0101] (Appendix 20) as described in Appendix 19 program wherein in the feature amount generation step, for each of the plurality of specific joint points, a set of feature amounts when only the feature amount of the specific joint point does not exist is generated as a training feature amount set from the feature amounts of each of the joint points of the target; in the learning model generation step, for each of the plurality of specific joint points, a machine learning model is generated by machine learning the positional relationship between other joint points when the feature amount of the specific joint point does not exist, using training data including the corresponding training feature amount set; characterized by program .
[0102] (Appendix 21) as described in Supplementary Note 19 or 20 program and the machine learning model is constructed by a convolutional neural network the feature amount includes a heatmap representing the possibility of the presence of joint points on the image the feature amount generation step in sets the feature amount to non - existent by setting the data on the heatmap to zero or 1 characterized by program
[0103] The present invention has been described with reference to the embodiments above, but the present invention is not limited to the above embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.
[0104] This application claims priority based on Japanese Patent Application No. 2021 - 029411 filed on February 26, 2021, and incorporates all of its disclosures herein.
Industrial Applicability
[0105] As described above, according to the present invention, it is possible to improve the estimation accuracy of the position of joint points. The present invention is useful in fields where posture detection of objects having joint points such as humans and robots is required. Specific fields include video surveillance, user interfaces, and the like.
Explanation of Signs
[0106] 10 Learning model generation device (Embodiment 1) 11 All - feature amount output unit 12 Feature amount generation unit 13 Learning model generation unit 14 Random number generation unit 15 Storage unit 16 Machine learning model (CNN) 20 Image data (for learning) 21 Heatmap 22 Training Heatmap Set 30 Joint Detection Device (Embodiment 2) 31 Full Feature Output Unit 32 Partial Feature Output Unit 33 Joint Detection Unit 34 Memory Unit 40 Image Data (Detection Target) 41 First Heatmap 42 Second Heatmap 110 Computer 111 CPU 112 Main Memory 113 Storage Device 114 Input Interface 115 Display Controller 116 Data Reader / Writer 117 Communication Interface 118 Input Device 119 Display Device 120 Recording Medium 121 Bus
Claims
1. All feature quantity output means for outputting a first feature quantity representing each joint point of the object from the image data of the object; Using, as an input, the first feature quantity for each of the joint points of the object, a machine learning model that has learned the positional relationship between other joint points when the feature quantity of a specific joint point does not exist, and outputting a second feature quantity representing each of the joint points of the object when the specific joint point is not visible in the image, partial feature quantity output means; Comprising: The machine learning model: From the feature quantities for each of the joint points of a specific object, for each of the plurality of specific joint points, generating a set of feature quantities when only the feature quantity of the specific joint point does not exist as a training feature quantity set, and for each of the plurality of specific joint points, using training data including the corresponding training feature quantity set to machine-learn the positional relationship between other joint points when the feature quantity of the specific joint point does not exist, and thereby being generated. A joint point detection device characterized by the above.
2. The joint point detection device according to Claim 1, Further comprising joint point detection means for detecting the coordinates of the joint points of the object using the second feature quantity for each of the joint points of the object. A joint point detection device characterized by the above.
3. The joint point detection device according to Claim 1, The partial feature quantity output means uses, as an input, the first feature quantity for each of the joint points of the object, and uses a machine learning model that has learned the positional relationship between other joint points when the feature quantity of a specific joint point does not exist for each of the plurality of specific joint points, and outputs a second feature quantity representing each of the joint points of the object when the specific joint point is not visible in the image. A joint point detection device characterized by the above.
4. The joint point detection device according to Claim 1, The machine learning model is constructed by a convolutional neural network, Each of the first feature quantity and the second feature quantity includes a heat map representing the possibility of the existence of joint points in the image. A joint point detection device characterized by the above.
5. A method executed by a computer, Outputting, from the image data of the object, a first feature quantity representing each joint point of the object. Using, as input, the first feature quantity for each of the target joint points, a machine learning model that machine-learns the positional relationship between other joint points when the feature quantity of a specific joint point does not exist, outputs a second feature quantity representing each of the target joint points when the specific joint point is not visible in the image, The machine learning model is From the feature quantity for each of the joint points of a specific target, for each of the plurality of said specific joint points, a set of feature quantities when only the feature quantity of the specific joint point does not exist is generated as a training feature quantity set, and for each of the plurality of said specific joint points, using training data including the corresponding training feature quantity set, the positional relationship between other joint points when the feature quantity of the specific joint point does not exist is machine-learned, and is thus generated. A joint point detection method characterized by this.
6. The joint point detection method according to claim 5, Using the second feature quantity for each of the target joint points, detecting the coordinates of the target joint points. A joint point detection method characterized by this.
7. On a computer, Output, for each of the target joint points, from the target image data, a first feature quantity representing the joint point, A program that, using, as input, the first feature quantity for each of the target joint points, outputs a second feature quantity representing each of the target joint points when the specific joint point is not visible in the image, using a machine learning model that machine-learns the positional relationship between other joint points when the feature quantity of a specific joint point does not exist, The machine learning model is From the feature quantity for each of the joint points of a specific target, for each of the plurality of said specific joint points, a set of feature quantities when only the feature quantity of the specific joint point does not exist is generated as a training feature quantity set, and for each of the plurality of said specific joint points, using training data including the corresponding training feature quantity set, the positional relationship between other joint points when the feature quantity of the specific joint point does not exist is machine-learned, and is thus generated, a program characterized by this.
8. The program according to claim 7, On the computer, further, Using the second feature quantity for each of the target joint points, causing the coordinates of the target joint points to be detected. A program characterized by this.
Citation Information
Patent Citations
Image generation device and method
JP2007004732A
Information processor, control method information processor and program
JP2017191576A
Human pose analysis system and method
WO2020000096A1