Learning model generation device, articulation point detection device, learning model generation method, articulation point detection method, and program

JPWO2024143048A5Active Publication Date: 2025-08-21NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024567647
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-08-21
Estimated Expiration
2043-12-18

AI Technical Summary

Technical Problem

Conventional learning models face challenges in achieving high detection accuracy for three-dimensional joint point coordinates due to noise in two-dimensional joint point coordinates extracted from actual images, making it difficult to determine noise and construct accurate models.

Method used

A learning model generation device and method that input first and second sample data into a basic and updated learning model, respectively, calculate the difference between intermediate features, and update parameters to reduce noise, improving detection accuracy.

Benefits of technology

The solution enhances the detection accuracy of three-dimensional joint point coordinates by updating model parameters based on feature differences, effectively addressing noise in two-dimensional data inputs.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A learning model generation device 10 is provided with: a first data input unit 11 that inputs first sample data to a basic learning model that has been trained through machine learning on the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; a second data input unit 12 that inputs second sample data having differences from the first sample data to a learning model to be updated; a difference calculation unit 13 that calculates the difference between an intermediate feature quantity of the basic learning model when the first sample data is input thereto and an intermediate feature quantity of the learning model to be updated when the second sample data is input thereto; and a parameter update unit 14 that uses the calculated difference to update the parameters of the learning model to be updated.
Need to check novelty before this filing date? Find Prior Art

Description

Learning model generation device, articulation point detection device, learning model generation method, articulation point detection method, and computer-readable recording medium

[0001] The present disclosure relates to a learning model generation device and a learning model generation method that generate a learning model for detecting human joint points, and further relates to a computer-readable recording medium on which a program for realizing these is recorded.The present disclosure also relates to a joint point detection device and a joint point detection method that detect human joint points, and further relates to a computer-readable recording medium on which a program for realizing these is recorded.

[0002] In recent years, a technology has been developed that estimates a person's posture by estimating the three-dimensional coordinates of each joint of the person from a two-dimensional image (see, for example, Patent Document 1). Such a technology is expected to be used in fields such as image monitoring systems, sports, and games. In addition, in such a technology, a learning model is used to detect the three-dimensional coordinates of each joint of the person.

[0003] The learning model is constructed by machine learning using, for example, two-dimensional coordinates of each joint (hereinafter referred to as "two-dimensional joint point coordinates") extracted from a person in an image and three-dimensional coordinates of each joint (hereinafter referred to as "three-dimensional joint point coordinates") as training data. In the training data, the three-dimensional joint point coordinates correspond to teacher data.

[0004] Furthermore, machine learning is performed by inputting two-dimensional joint point coordinates, which serve as training data, into a learning model, and updating the parameters of the learning model so that the difference between the output three-dimensional joint point coordinates and the three-dimensional joint point coordinates, which serve as teacher data, becomes smaller.

[0005] Japanese Patent Application Laid-Open No. 2021-47563

[0006] In conventional machine learning of learning models, 2D joint point coordinates with little noise, such as errors in the positions of each joint point, are used as training data. This is because if noisy 2D joint point coordinates are used as training data, the learning model cannot determine whether the data is noise, making it impossible to build a learning model with high detection accuracy.

[0007] However, because 2D joint point coordinates extracted from actual 2D images contain noise, it is difficult for conventional learning models to accurately detect 3D joint point coordinates when 2D joint point coordinates extracted from actual 2D images are used as input.

[0008] An example of an objective of the present disclosure is to improve the detection accuracy when estimating the three-dimensional coordinates of each joint point from the two-dimensional coordinates of each joint.

[0009] In order to achieve the above object, a learning model generation device according to one aspect of the present disclosure is characterized by comprising: a first data input unit that inputs first sample data into a basic learning model that has undergone machine learning to learn the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; a second data input unit that inputs second sample data that has differences from the first sample data into a learning model to be updated; a difference calculation unit that calculates the difference between intermediate features of the basic learning model when the first sample data is input and intermediate features of the learning model to be updated when the second sample data is input; and a parameter update unit that updates parameters of the learning model to be updated using the calculated difference.

[0010] In order to achieve the above object, a joint point detection device according to one aspect of the present disclosure includes a joint point detection unit that inputs two-dimensional joint point coordinates of a person to a learning model and detects three-dimensional joint point coordinates of the person, and parameters of the learning model are updated using a difference between intermediate feature values ​​when first sample data is input to a basic learning model that performs machine learning to learn the relationship between the two-dimensional joint point coordinates and the three-dimensional joint point coordinates, and intermediate feature values ​​when second sample data that differs from the first sample data is input to the learning model.

[0011] In order to achieve the above object, a learning model generation method according to one aspect of the present disclosure is characterized by comprising: a first data input step of inputting first sample data into a basic learning model that has undergone machine learning to learn the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; a second data input step of inputting second sample data that has differences from the first sample data into a learning model to be updated; a difference calculation step of calculating a difference between intermediate features of the basic learning model when the first sample data is input and intermediate features of the learning model to be updated when the second sample data is input; and a parameter update step of updating parameters of the learning model to be updated using the calculated difference.

[0012] In order to achieve the above object, a joint point detection method according to one aspect of the present disclosure includes a joint point detection step of inputting two-dimensional joint point coordinates of a person to a learning model and detecting three-dimensional joint point coordinates of the person, wherein parameters of the learning model are updated using a difference between intermediate feature values ​​when first sample data is input to a basic learning model that performs machine learning to learn the relationship between the two-dimensional joint point coordinates and the three-dimensional joint point coordinates, and intermediate feature values ​​when second sample data that differs from the first sample data is input to the learning model.

[0013] Furthermore, in order to achieve the above object, a first computer-readable recording medium according to one aspect of the present disclosure is characterized in that it records a program including instructions for causing a computer to execute: a first data input step of inputting first sample data into a basic learning model that has learned the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates by machine learning; a second data input step of inputting second sample data that has differences from the first sample data into a learning model to be updated; a difference calculation step of calculating a difference between intermediate features of the basic learning model when the first sample data is input and intermediate features of the learning model to be updated when the second sample data is input; and a parameter update step of updating parameters of the learning model to be updated using the calculated difference.

[0014] Furthermore, to achieve the above object, a second computer-readable recording medium according to one aspect of the present disclosure has recorded thereon a program including instructions for causing a computer to execute a joint point detection step of inputting two-dimensional joint point coordinates of a person into a learning model and detecting three-dimensional joint point coordinates of the person, wherein parameters of the learning model are updated using a difference between intermediate feature values ​​obtained when first sample data is input into a basic learning model that has machine-learned the relationship between the two-dimensional joint point coordinates and the three-dimensional joint point coordinates, and intermediate feature values ​​obtained when second sample data that differs from the first sample data is input into the learning model.

[0015] As described above, according to the present disclosure, it is possible to improve the detection accuracy when estimating the three-dimensional coordinates of each joint point from the two-dimensional coordinates of each joint.

[0016] FIG. 1 is a configuration diagram showing a schematic configuration of an example of a learning model generation device. FIG. 2 is a diagram explaining the function of a basic learning model. FIG. 3 is a configuration diagram specifically showing the configuration of an example of a learning model generation device. FIG. 4 is an explanatory diagram explaining first sample data and second sample data. FIG. 5 is a configuration diagram showing an example of the configuration of a basic learning model and an updated model. FIG. 6 is a flow diagram showing an example of the operation of the learning model generation device. FIG. 7 is a configuration diagram showing the configuration of another example of a learning model generation device. FIG. 8 is a flow diagram showing another example of the operation of the learning model generation device. FIG. 9 is a configuration diagram showing the configuration of an example of an articulation point detection device. FIG. 10 is a flow diagram showing an example of the operation of the articulation point detection device 50. FIG. 11 is a block diagram showing an example of a computer that realizes the learning model generation device and the articulation point detection device.

[0017] First Embodiment Hereinafter, an example of a learning model generation device, a learning model generation method, and a program will be described with reference to FIGS.

[0018] [Device Configuration] First, the schematic configuration of an example of a learning model generation device will be described with reference to Figures 1 and 2. Figure 1 is a diagram showing the schematic configuration of an example of a learning model generation device. Figure 2 is a diagram explaining the function of a basic learning model.

[0019] 1, the learning model generation device 10 is a device for generating a learning model that receives two-dimensional human joint point coordinates as input and outputs corresponding three-dimensional human joint point coordinates. As shown in FIG. 1, the learning model generation device 10 includes a first data input unit 11, a second data input unit 12, a difference calculation unit 13, and a parameter update unit 14.

[0020] The first data input unit 11 inputs first sample data to a basic learning model. The basic learning model is a learning model that performs machine learning to learn the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates. The second data input unit 12 inputs second sample data that differs from the first sample data to a learning model to be updated (hereinafter referred to as an "updated model").

[0021] The basic learning model will be described with reference to Fig. 2. As shown in Fig. 2, when two-dimensional joint point coordinates of a person's joint points are input to the basic learning model, the basic learning model outputs three-dimensional joint point coordinates of the corresponding joint points of the person. Note that the two-dimensional joint point coordinates have been extracted in advance from image data of the person using, for example, a machine learning model that has learned the relationship between image features and joint points by machine learning.

[0022] The difference calculation unit 13 calculates the difference between the intermediate feature values ​​of the basic learning model when the first sample data is input and the intermediate feature values ​​of the learning model to be updated when the second sample data is input. The parameter update unit 14 updates the parameters of the learning model to be updated using the difference calculated by the difference calculation unit 13.

[0023] In this way, in the first embodiment, the parameters of the updated model are updated using the difference between the intermediate feature amounts of the basic learning model and the intermediate feature amounts of the updated model, so that machine learning that takes noise into consideration is performed in the updated model. Therefore, the updated model obtained in the first embodiment can improve the detection accuracy when estimating the three-dimensional coordinates of each joint point from the two-dimensional coordinates of each joint.

[0024] Next, the configuration and functions of an example of a learning model generation device will be specifically described using Figures 3 to 5 in addition to Figures 1 and 2. Figure 3 is a configuration diagram specifically showing the configuration of an example of a learning model generation device. Figure 4 is an explanatory diagram for explaining first sample data and second sample data. Figure 5 is a configuration diagram showing an example of the configuration of a basic learning model and an updated model.

[0025] As shown in FIG. 3 , a basic learning model 20 and an updated model 30 are implemented in the learning model generation device 10. In the first embodiment, the basic learning model 20 and the updated model 30 are, for example, neural networks. In addition, the basic learning model 20 and the updated model 30 are actually implemented by a machine learning program executed on a computer. Note that the basic learning model 20 and the updated model 30 may also be constructed outside the learning model generation device 10.

[0026] As shown in Fig. 4, the first sample data is two-dimensional joint point data including two-dimensional joint point coordinates detected from a specific person image. In contrast, the second sample data is also two-dimensional joint point data including two-dimensional joint point coordinates detected from a specific person image. However, as shown in Fig. 4, the second sample data differs from the first sample data in all or some of the values ​​of the joint point coordinates. In other words, the second sample data is data in which noise has been added to the first sample data.

[0027] 2, the second data input unit 12 acquires second sample data prepared in advance from an external device, but the first embodiment is not limited to this. In the first embodiment, the second data input unit 12 can also acquire first sample data and generate second sample data by adding noise to the acquired first sample data.

[0028] Specifically, the second data input unit 12 randomly selects one of the articulation points of the first sample data, and randomly moves the position of the selected articulation point to another position to generate the second sample data (see FIG. 4).

[0029] 5, in the first embodiment, the basic learning model 20 and the update model 30 are neural networks as described above. The basic learning model 20 includes, for example, an input layer 21, an intermediate layer (hidden layer) 22, and an output layer 23. Similarly, the update model 30 includes an input layer 31, an intermediate layer (hidden layer) 32, and an output layer 33.

[0030] In the first embodiment, the difference calculation unit 13 acquires intermediate features from the intermediate layer 22 of the basic learning model 20, and also acquires intermediate features from the intermediate layer 32 of the updated model 30. Then, the difference calculation unit 13 calculates the difference between the intermediate features acquired from the intermediate layer 22 and the intermediate features acquired from the intermediate layer 32.

[0031] In the first embodiment, when the difference is calculated by the difference calculation unit 13, the parameter update unit 14 updates the parameters of the updated model 30, that is, the weights of each node, so as to reduce the calculated difference.

[0032] [Device Operation] Next, an example of the operation of the learning model generation device 10 will be described with reference to FIG. 6. FIG. 6 is a flow diagram showing an example of the operation of the learning model generation device. In the following description, reference will be made to FIGS. 1 to 5 as appropriate. In addition, in the first embodiment, the learning model generation method is implemented by operating the learning model generation device. Therefore, the description of the learning model generation method in the first embodiment will be replaced by the following description of the operation of the learning model generation device 10.

[0033] As shown in FIG. 6, first, the first data input unit 11 acquires first sample data and inputs the acquired first sample data into the basic learning model 20 (step A1).

[0034] Next, the second data input unit 12 acquires second sample data that has differences from the first sample data, and inputs the acquired second sample data into the updated model (step A2).

[0035] In addition, in embodiment 1, in step A2, the second data input unit 12 may acquire the first sample data instead of the second sample data, and generate the second sample data from the acquired first sample data.

[0036] Next, the difference calculation unit 13 acquires intermediate features from the intermediate layer 22 of the basic learning model 20 (step A3). The difference calculation unit 13 also acquires intermediate features from the intermediate layer 32 of the updated model 30 (step A4).

[0037] Next, the difference calculation unit 13 calculates the difference between the intermediate feature amount acquired from the intermediate layer 22 in step A3 and the intermediate feature amount acquired from the intermediate layer 32 in step A4 (step A5).

[0038] Thereafter, the parameter update unit 14 obtains the difference calculated by the difference calculation unit 13 in step A5, and updates the parameters of the updated model 30 so as to reduce this difference (step A6). Note that in the first embodiment, the updated model is completed by updating the parameters of the updated model. In other words, the updated model is generated by updating the parameters of the updated model.

[0039] Steps A1 to A6 are repeated the same number of times as the number of first sample data and second sample data. It is preferable that the input first sample data and second sample data have no time difference between them. The smaller the time difference, the more noticeable the noise becomes, making it easier to correct.

[0040] As described above, in the first embodiment, a learning model that has been machine-learned using training data with little noise is used as the basic learning model 20. Furthermore, the parameters of the updated model 30 are updated using the difference between the intermediate feature values ​​of the basic learning model 20 and the intermediate feature values ​​of the updated model 30. As a result, the parameters of the updated model 30 are updated taking into account the noise contained in the second sample data. Therefore, the updated model 30 whose parameters have been updated in the first embodiment can improve the detection accuracy when estimating the three-dimensional coordinates of each joint point from the two-dimensional coordinates of each joint.

[0041] [Modification] Here, a modification of the first embodiment will be described. In this modification, images of successive frames of video data capturing a person are used as the specific person images. In this case, the first sample data and the second sample data are each created for each frame of the video data.

[0042] The first data input unit 11 then inputs N pieces of first sample data to the basic learning model 20 in the time series of frames. The second data input unit 12 inputs N pieces of second sample data to the updated model 30 in synchronization with the first data input unit 11, where N is an arbitrary integer.

[0043] In this case, the basic learning model 20 and the update model 30 may output N pieces of output data corresponding to each of the N pieces of sample data, or may output only one piece of output data (for example, the output data corresponding to the central frame).

[0044] The difference calculation unit 13 acquires intermediate features from the intermediate layer 22 of the basic learning model 20 when N pieces of first sample data are input. The difference calculation unit 13 also acquires intermediate features from the intermediate layer 32 of the updated model 30 when N pieces of second sample data are input. Then, the difference calculation unit 13 calculates the difference between the intermediate features acquired from the intermediate layer 22 and the intermediate features acquired from the intermediate layer 32.

[0045] Thereafter, the parameter update unit 14 updates the parameters of the updated model 30 so that the difference calculated by the difference calculation unit 13 becomes smaller, in the same way as when no video data is used.

[0046] In this modified example, a series of multiple frames is input to each learning model as a single input. The second sample data may contain a mixture of noise-containing and noise-free frames. In this case, the updated model 30 compensates for noise between frames, resulting in more pronounced noise correction through parameter updates.

[0047] [Program] The program in the first embodiment may be a program that causes a computer to execute steps A1 to A6 shown in Fig. 6. By installing and executing this program on a computer, the learning model generation device 10 and the learning model generation method in the first embodiment can be realized. In this case, the processor of the computer functions as the first data input unit 11, the second data input unit 12, the difference calculation unit 13, and the parameter update unit 14 and performs processing. Examples of the computer include a general-purpose PC, a smartphone, and a tablet terminal device.

[0048] The program in the first embodiment may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as one of the first data input unit 11, the second data input unit 12, the difference calculation unit 13, and the parameter update unit 14.

[0049] Second Embodiment Hereinafter, another example of a learning model generation device, a learning model generation method, and a program will be described with reference to FIGS.

[0050] [Device Configuration] First, the configuration of another example of a learning model generation device will be described with reference to Fig. 7. Fig. 7 is a diagram showing the configuration of another example of a learning model generation device.

[0051] Similar to the first embodiment, the learning model generation device 40 shown in FIG. 7 is a device for generating a learning model that receives as input the joint point coordinates of a two-dimensional person and outputs the corresponding joint point coordinates of a three-dimensional person.

[0052] As shown in Figure 7, the learning model generation device 40, like the learning model generation device 10 shown in embodiment 1, also includes a first data input unit 11, a second data input unit 12, a difference calculation unit 13, and a parameter update unit 14.

[0053] However, unlike the learning model generation device 10 shown in the first embodiment, the learning model generation device 40 includes, in addition to the above-mentioned configuration, a second difference calculation unit 41 and a statistical processing unit 42. The following mainly describes the differences from the first embodiment.

[0054] The second difference calculation unit 41 calculates, as the second difference, the difference between the correct answer data and the output data of the updated model 30 when the second sample data is input. The correct answer data is data including three-dimensional joint point coordinates that are correct for the first sample data. In other words, the correct answer data is data composed of the actual three-dimensional coordinates of the joint points of the person in the image from which the first and second sample data were extracted.

[0055] The statistical processing unit 42 performs statistical processing using the difference calculated by the difference calculation unit 13 and the second difference calculated by the second difference calculation unit 41. The statistical processing in this case is not particularly limited. A specific example of the statistical processing is weighted averaging.

[0056] In the second embodiment, the parameter update unit 14 updates the parameters of the updated model 30 based on the results of the statistical processing by the statistical processing unit 42. Specifically, the parameter update unit 14 updates the parameters, i.e., the weights, of the updated model 30 so that the value (e.g., weighted average value) obtained by the statistical processing by the statistical processing unit 42 becomes smaller.

[0057] Also in the second embodiment, the first sample data and the second sample data may be input for each frame. In this case, the calculation of the difference by the difference calculation unit 13, the calculation of the second difference by the second difference calculation unit 41, the statistical processing by the statistical processing unit 42, and the updating by the parameter update unit 14 are performed for each frame.

[0058] [Device Operation] Next, another example of the operation of the learning model generation device 40 will be described with reference to FIG. 8. FIG. 8 is a flow diagram showing another example of the operation of the learning model generation device. In the following description, reference will be made to FIG. 7 as appropriate. In addition, in the second embodiment, the learning model generation method is implemented by operating the learning model generation device. Therefore, the description of the learning model generation method in the second embodiment will be replaced by the description of the operation of the learning model generation device 4 below.

[0059] As shown in FIG. 8, first, the first data input unit 11 acquires first sample data and inputs the acquired first sample data into the basic learning model 20 (step B1).

[0060] Next, the second data input unit 12 acquires second sample data that has differences from the first sample data, and inputs the acquired second sample data into the updated model (step B2).

[0061] Next, the difference calculation unit 13 acquires intermediate features from the intermediate layer 22 of the basic learning model 20 (step B3). The difference calculation unit 13 also acquires intermediate features from the intermediate layer 32 of the updated model 30 (step B4).

[0062] Next, the difference calculation unit 13 calculates the difference between the intermediate feature acquired from the intermediate layer 22 in step A3 and the intermediate feature acquired from the intermediate layer 32 in step A4 (step B5). Note that steps B1 to B5 above are substantially the same as steps A1 to A5 shown in FIG. 6 in the first embodiment.

[0063] Next, in the second embodiment, the second difference calculation unit 41 acquires correct answer data from the outside, acquires the output data when the second sample data is input from the updated model 30, and calculates the difference between the correct answer data and the output data of the updated model 30 as the second difference (step B6).

[0064] Next, the statistical processing unit 42 executes statistical processing using the difference calculated in step B5 and the second difference calculated in step B6 (step B7).

[0065] Thereafter, the parameter update unit 14 updates the parameters of the updated model 30 based on the results of the statistical processing in step B7 (step B8).

[0066] Steps B1 to B8 are repeatedly executed the same number of times as the number of first sample data and second sample data, as in embodiment 1. Also in embodiment 2, if the first sample data and second sample data are obtained for each frame of video data, steps B1 to B8 are repeatedly executed for each frame.

[0067] As described above, in the second embodiment, the parameters of the updated model 30 are updated using the difference (second difference) between the output data of the updated model and the ground truth data. As a result, in the second embodiment, noise contained in the second sample data is taken into consideration to an even greater extent. According to the second embodiment, it is possible to further improve the detection accuracy when estimating the three-dimensional coordinates of each joint point from the two-dimensional coordinates of each joint.

[0068] Also in the second embodiment, as described in the modified example of the first embodiment, images of successive frames of video data capturing a person may be used as the specific person image. In this case, also in the second embodiment, the first data input unit 11 inputs N pieces of first sample data to the basic learning model 20 in chronological order of the frames. Furthermore, the second data input unit 12 inputs N pieces of second sample data to the updated model 30 in synchronization with the first data input unit 11.

[0069] In this case, the basic learning model 20 and the update model 30 may output N pieces of output data corresponding to each of the N pieces of sample data, or may output only one piece of output data (for example, the output data corresponding to the central frame).

[0070] Then, in the former case where N pieces of output data are output, unlike step B6 described above, the second difference calculation unit 41 calculates, for each frame, N differences between the correct answer data and the output data of the updated model 30. Thereafter, the second difference calculation unit 41 uses each calculated difference to perform statistical processing such as arithmetic averaging or weighted averaging to calculate one (scalar) loss, which is set as the second difference.

[0071] In the latter case where only one output data is output, the second difference calculation unit 41 calculates the difference between the correct data and the output data of the updated model 30, as in step B6 above, and sets the calculated difference as the second difference.

[0072] [Program] The program in the second embodiment may be any program that causes a computer to execute steps B1 to B8 shown in Fig. 8. By installing and executing this program on a computer, the learning model generation device 40 and the learning model generation method in the second embodiment can be realized. In this case, the processor of the computer functions as the first data input unit 11, the second data input unit 12, the difference calculation unit 13, the parameter update unit 14, the second difference calculation unit 41, and the statistical processing unit 42 to perform processing. Examples of the computer include a general-purpose PC, a smartphone, and a tablet terminal device.

[0073] The program in the second embodiment may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as any one of the first data input unit 11, the second data input unit 12, the difference calculation unit 13, the parameter update unit 14, the second difference calculation unit 41, and the statistical processing unit 42.

[0074] Third Embodiment Next, an example of a joint point detection device, a joint point detection method, and a program will be described with reference to FIGS.

[0075] [Device Configuration] First, the configuration of an example of a joint point detection device will be described with reference to Fig. 9. Fig. 9 is a diagram showing the configuration of an example of a joint point detection device.

[0076] The joint point detection device 50 shown in Fig. 9 is a device that detects, from the two-dimensional joint point coordinates of a person, the corresponding three-dimensional joint point coordinates of a person. As shown in Fig. 9, the joint point detection device 50 includes a joint point detection unit 51. The joint point detection unit 51 inputs the two-dimensional joint point coordinates of the person to a learning model 60 and detects the three-dimensional joint point coordinates of the person.

[0077] The learning model 60 is a learning model that performs machine learning to learn the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates. The parameters of the learning model 60 are updated using the difference between the first intermediate feature amount and the second intermediate feature amount. The first intermediate feature amount is an intermediate feature amount obtained when first sample data is input to the basic learning model that performs machine learning to learn the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates. The second intermediate feature amount is an intermediate feature amount obtained when second sample data that differs from the first sample data is input to the learning model 60.

[0078] Specifically, in the third embodiment, the learning model 60 is a neural network, and is the learning model generated by the learning model generation device in the first or second embodiment, i.e., the updated model 30 (see FIGS. 1, 3, and 7). The learning model 60 is also implemented by a machine learning program executed on a computer.

[0079] In this way, the joint point detection device 50 detects three-dimensional joint point coordinates of a person using the learning model generated by the learning model generation device according to the first or second embodiment.

[0080] [Device Operation] Next, an example of the operation of the joint point detection device 50 will be described with reference to Fig. 10. Fig. 10 is a flow diagram showing an example of the operation of the joint point detection device 50. In the following description, Fig. 9 will be referred to as appropriate. Furthermore, in the third embodiment, the joint point detection method is carried out by operating the joint point detection device 50. Therefore, the description of the joint point detection method in the third embodiment will be replaced by the following description of the operation of the joint point detection device 50.

[0081] 10, first, the joint point detection unit 51 acquires two-dimensional coordinates of the joint points of the person, which serve as input data (step C1). Next, the joint point detection unit 51 inputs the two-dimensional coordinates acquired in step C1 into the learning model 60 (step C2).

[0082] Next, the joint point detection unit 51 acquires the three-dimensional coordinates of the joint points output by the learning model 60 (step C3). In the third embodiment, the three-dimensional coordinates of the joint points of the person are detected by step C3.

[0083] Next, the joint point detection unit 51 outputs the three-dimensional coordinates of the joint points acquired in step C3 to the outside (step C4).

[0084] As described above, according to the third embodiment, the three-dimensional coordinates of the joint points can be detected using the learning model generated according to the first or second embodiment. Therefore, the three-dimensional coordinates of the joint points are detected taking into consideration noise in the two-dimensional coordinates that serve as input data, and are highly accurate values.

[0085] [Program] The program in the third embodiment may be a program that causes a computer to execute steps C1 to C4 shown in Fig. 10. By installing and executing this program on a computer, the joint point detection device 50 and the joint point detection method in this embodiment can be realized. In this case, the processor of the computer functions as and performs processing. Examples of the computer include a general-purpose PC, a smartphone, and a tablet terminal device.

[0086] The program in the third embodiment may be executed by a computer system constructed by a plurality of computers. In this case, the plurality of computers function as the joint point detection unit 51.

[0087] [Physical Configuration] A computer that realizes the learning model generation device and the joint point detection device by executing the programs in the first to third embodiments will now be described with reference to Fig. 11. Fig. 11 is a block diagram showing an example of a computer that realizes the learning model generation device and the joint point detection device.

[0088] 11, the computer 110 includes a CPU (Central Processing Unit) 111, a main memory 112, a storage device 113, an input interface 114, a display controller 115, a data reader / writer 116, and a communication interface 117. These components are connected to each other via a bus 121 so as to be able to communicate data with each other.

[0089] Furthermore, the computer 110 may include a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array) in addition to or instead of the CPU 111. In this aspect, the GPU or FPGA can execute the programs in the embodiments.

[0090] The CPU 111 loads a program in the embodiment, which is composed of a group of codes and stored in the storage device 113, into the main memory 112 and executes each code in a predetermined order to perform various calculations. The main memory 112 is typically a volatile storage device such as a DRAM (Dynamic Random Access Memory).

[0091] The programs in the first to third embodiments are provided in a state stored in a computer-readable recording medium 120. The programs in the present embodiments may be distributed over the Internet connected via the communication interface 117.

[0092] Specific examples of the storage device 113 include a hard disk drive and a semiconductor storage device such as a flash memory. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and a mouse. The display controller 115 is connected to a display device 119 and controls the display on the display device 119.

[0093] The data reader / writer 116 mediates data transmission between the CPU 111 and the recording medium 120, reads programs from the recording medium 120, and writes processing results from the computer 110 to the recording medium 120. The communication interface 117 mediates data transmission between the CPU 111 and other computers.

[0094] Specific examples of the recording medium 120 include general-purpose semiconductor storage devices such as CF (Compact Flash (registered trademark)) and SD (Secure Digital), magnetic recording media such as flexible disks, or optical recording media such as CD-ROMs (Compact Disk Read Only Memory).

[0095] Note that the learning model generation device and the joint point detection device in the embodiments can be realized not by a computer on which a program is installed, but by hardware corresponding to each unit, such as an electronic circuit. Furthermore, the learning model generation device and the joint point detection device in the embodiments may be realized in part by a program and in part by hardware. In the embodiments, the computer is not limited to the computer shown in FIG. 11.

[0096] Some or all of the above-described embodiments can be expressed by (Supplementary Note 1) to (Supplementary Note 18) described below, but are not limited to the following descriptions.

[0097] (Supplementary Note 1) A learning model generation device comprising: a first data input unit that inputs first sample data into a basic learning model that is machine learning the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; a second data input unit that inputs second sample data that has differences from the first sample data into a learning model to be updated; a difference calculation unit that calculates a difference between intermediate features of the basic learning model when the first sample data is input and intermediate features of the learning model to be updated when the second sample data is input; and a parameter update unit that updates parameters of the learning model to be updated using the calculated difference.

[0098] (Supplementary Note 2) The learning model generation device according to Supplementary Note 1, wherein the first sample data includes two-dimensional joint point coordinates detected from a specific person image, and the second sample data includes two-dimensional joint point coordinates detected from the specific person image, and all or some of the values ​​of the joint point coordinates differ from those of the first sample data.

[0099] (Supplementary Note 3) The learning model generation device according to Supplementary Note 2, comprising: a second difference calculation unit that calculates, as a second difference, the difference between correct answer data including three-dimensional joint point coordinates that are correct for the first sample data and output data of the learning model to be updated when the second sample data is input; and a statistical processing unit that performs statistical processing using the difference calculated by the difference calculation unit and the second difference calculated by the second difference calculation unit, wherein the parameter update unit updates parameters of the learning model to be updated based on the results of the statistical processing.

[0100] (Supplementary Note 4) The learning model generation device according to Supplementary Note 2 or 3, wherein the specific person images are images of successive frames of video data capturing a person, the first sample data and the second sample data each consist of a plurality of the frames, the first data input unit inputs the plurality of first sample data in chronological order of the frames, and the second data input unit inputs the plurality of second sample data in synchronization with the first data input unit.

[0101] (Supplementary Note 5) The learning model generation device according to Supplementary Note 1, wherein the second data input unit generates the second sample data by adding noise to the first sample data, and inputs the generated second sample data into the basic learning model.

[0102] (Supplementary Note 6) A joint point detection device comprising: a joint point detection unit that inputs two-dimensional joint point coordinates of a person to a learning model and detects three-dimensional joint point coordinates of the person; and parameters of the learning model are updated using a difference between intermediate feature values ​​when first sample data is input to a basic learning model that machine-learns the relationship between the two-dimensional joint point coordinates and the three-dimensional joint point coordinates, and intermediate feature values ​​when second sample data that differs from the first sample data is input to the learning model.

[0103] (Supplementary Note 7) A learning model generation method comprising: a first data input step of inputting first sample data into a basic learning model that has machine-learned the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; a second data input step of inputting second sample data that has differences from the first sample data into a learning model to be updated; a difference calculation step of calculating a difference between intermediate features of the basic learning model when the first sample data is input and intermediate features of the learning model to be updated when the second sample data is input; and a parameter update step of updating parameters of the learning model to be updated using the calculated difference.

[0104] (Supplementary Note 8) The learning model generation method according to Supplementary Note 7, wherein the first sample data includes two-dimensional joint point coordinates detected from a specific person image, and the second sample data includes two-dimensional joint point coordinates detected from the specific person image, and all or some of the values ​​of the joint point coordinates differ from those of the first sample data.

[0105] (Supplementary Note 9) The learning model generation method according to Supplementary Note 8, further comprising: a second difference calculation step of calculating, as a second difference, the difference between correct answer data including three-dimensional joint point coordinates that are correct for the first sample data and output data of the learning model to be updated when the second sample data is input; and a statistical processing step of performing statistical processing using the difference calculated by the difference calculation step and the second difference calculated by the second difference calculation step, wherein, in the parameter update step, parameters of the learning model to be updated are updated based on the results of the statistical processing.

[0106] (Supplementary Note 10) The learning model generation method according to Supplementary Note 8 or 9, wherein the specific person images are images of successive frames of video data capturing a person, and the first sample data and the second sample data each consist of a plurality of the frames, and in the first data input step, the plurality of first sample data are input in chronological order of the frames, and in the second data input step, the plurality of second sample data are input in synchronization with the first data input step.

[0107] (Supplementary Note 11) The learning model generation method according to Supplementary Note 7, wherein in the second data input step, the second sample data is generated by adding noise to the first sample data, and the generated second sample data is input to the basic learning model.

[0108] (Supplementary Note 12) A joint point detection method comprising: a joint point detection step of inputting two-dimensional joint point coordinates of a person to a learning model and detecting three-dimensional joint point coordinates of the person; and parameters of the learning model are updated using a difference between an intermediate feature value when first sample data is input to a basic learning model that has machine-learned the relationship between the two-dimensional joint point coordinates and the three-dimensional joint point coordinates, and an intermediate feature value when second sample data that differs from the first sample data is input to the learning model.

[0109] (Supplementary Note 13) A computer-readable recording medium having recorded thereon a program including instructions for causing a computer to execute the following steps: a first data input step of inputting first sample data into a basic learning model that is learning the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates by machine learning; a second data input step of inputting second sample data that has differences from the first sample data into a learning model to be updated; a difference calculation step of calculating the difference between intermediate features of the basic learning model when the first sample data is input and intermediate features of the learning model to be updated when the second sample data is input; and a parameter update step of updating parameters of the learning model to be updated using the calculated difference.

[0110] (Supplementary Note 14) The computer-readable recording medium according to Supplementary Note 13, wherein the first sample data includes two-dimensional joint point coordinates detected from a specific person image, and the second sample data includes two-dimensional joint point coordinates detected from the specific person image, and all or some of the values ​​of the joint point coordinates differ from those of the first sample data.

[0111] (Supplementary Note 15) The computer-readable recording medium according to Supplementary Note 14, wherein the program further includes instructions to cause the computer to execute: a second difference calculation step of calculating, as a second difference, the difference between correct answer data including three-dimensional joint point coordinates that are correct for the first sample data and output data of the learning model to be updated when the second sample data is input; and a statistical processing step of performing statistical processing using the difference calculated by the difference calculation step and the second difference calculated by the second difference calculation step; and wherein, in the parameter update step, parameters of the learning model to be updated are updated based on the results of the statistical processing.

[0112] (Appendix 16) The computer-readable recording medium described in Appendix 14 or 15, wherein the specific person images are images of successive frames of video data of a person, and the first sample data and the second sample data are each composed of a plurality of the frames, and in the first data input step, the plurality of first sample data are input in chronological order of the frames, and in the second data input step, the plurality of second sample data are input in synchronization with the first data input step.

[0113] (Supplementary Note 17) The computer-readable recording medium according to Supplementary Note 13, wherein in the second data input step, the second sample data is generated by adding noise to the first sample data, and the generated second sample data is input to the basic learning model.

[0114] (Supplementary Note 18) A computer-readable recording medium comprising: a program including instructions recorded on a computer to execute a joint point detection step of inputting two-dimensional joint point coordinates of a person into a learning model and detecting three-dimensional joint point coordinates of the person; and parameters of the learning model are updated using a difference between intermediate feature values ​​when first sample data is input into a basic learning model that machine-learns the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates, and intermediate feature values ​​when second sample data that differs from the first sample data is input into the learning model.

[0115] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure.

[0116] This application claims priority based on Japanese Patent Application No. 2022-211356, filed December 28, 2022, the disclosure of which is incorporated herein in its entirety.

[0117] As described above, according to the present disclosure, it is possible to improve the detection accuracy when estimating the three-dimensional coordinates of each joint point from the two-dimensional coordinates of each joint. The present disclosure is useful for systems that require estimation of a person's posture from an image, such as a video surveillance system.

[0118] DESCRIPTION OF SYMBOLS 10 Learning model generation device (first embodiment) 11 First data input unit 12 Second data input unit 13 Difference calculation unit 14 Parameter update unit 20 Basic learning model 21 Input layer 22 Intermediate layer (hidden layer) 23 Output layer 30 Updated model 31 Input layer 32 Intermediate layer (hidden layer) 33 Output layer 40 Learning model generation device (second embodiment) 41 Second difference calculation unit 42 Statistical processing unit 50 Joint point detection device 51 Joint point detection unit 60 Learning model 111 CPU 112 Main memory 113 Storage device 114 Input interface 115 Display controller 116 Data reader / writer 117 Communication interface 118 Input device 119 Display device 120 Recording medium 121 Bus

Claims

1. a first data input unit that inputs first sample data into a basic learning model that performs machine learning on the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; a second data input unit that inputs second sample data having differences from the first sample data into the learning model to be updated; a difference calculation unit that calculates a difference between an intermediate feature value of the basic learning model when the first sample data is input and an intermediate feature value of the learning model to be updated when the second sample data is input; a parameter update unit that updates parameters of the learning model to be updated using the calculated difference; A learning model generation device comprising:

2. the first sample data includes two-dimensional joint point coordinates detected from a specific person image; the second sample data includes two-dimensional joint point coordinates detected from the specific person image, and values ​​of all or part of the joint point coordinates differ from those of the first sample data; The learning model generation device according to claim 1 .

3. a second difference calculation unit that calculates, as a second difference, a difference between correct answer data including three-dimensional joint point coordinates that are correct for the first sample data and output data of the learning model to be updated when the second sample data is input; a statistical processing unit that performs statistical processing using the difference calculated by the difference calculation unit and the second difference calculated by the second difference calculation unit; Equipped with the parameter update unit updates the parameters of the learning model to be updated based on the results of the statistical processing. The learning model generating device according to claim 2 .

4. the specific person images are images of successive frames of video data capturing a person, and the first sample data and the second sample data each consist of a plurality of the frames; the first data input unit inputs a plurality of first sample data in time sequence of the frames; the second data input unit inputs the second sample data in synchronization with the first data input unit; The learning model generating device according to claim 2 or 3.

5. the second data input unit generates the second sample data by adding noise to the first sample data, and inputs the generated second sample data into the basic learning model. The learning model generation device according to claim 1 .

6. a joint point detection unit that inputs two-dimensional joint point coordinates of a person into a learning model and detects three-dimensional joint point coordinates of the person; The parameters of the learning model are: intermediate features when the first sample data is input to a basic learning model that performs machine learning on the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; intermediate features obtained when second sample data having differences from the first sample data is input to the learning model; It has been updated using the difference of A joint point detection device characterized by:

7. inputting first sample data into a basic learning model that performs machine learning on the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; inputting second sample data having differences from the first sample data into the learning model to be updated; calculating a difference between an intermediate feature value of the basic learning model when the first sample data is input and an intermediate feature value of the learning model to be updated when the second sample data is input; updating the parameters of the learning model to be updated using the calculated difference; A learning model generation method characterized by:

8. Two-dimensional joint point coordinates of a person are input to a learning model, and three-dimensional joint point coordinates of the person are detected; The parameters of the learning model are: intermediate features when the first sample data is input to a basic learning model that performs machine learning on the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; intermediate features obtained when second sample data having differences from the first sample data is input to the learning model; It has been updated using the difference of A joint point detection method characterized by:

9. On the computer, inputting the first sample data into a basic learning model that performs machine learning on the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; inputting second sample data having differences from the first sample data into a learning model to be updated; calculating a difference between an intermediate feature value of the basic learning model when the first sample data is input and an intermediate feature value of the learning model to be updated when the second sample data is input; updating the parameters of the learning model to be updated using the calculated difference; A program that executes.

10. On the computer, a joint point detection step of inputting two-dimensional joint point coordinates of a person into the learning model and detecting three-dimensional joint point coordinates of the person; The parameters of the learning model are: intermediate features when the first sample data is input to a basic learning model that performs machine learning on the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; intermediate features obtained when second sample data having differences from the first sample data is input to the learning model; It has been updated using the difference of program.