Learning model generation device, joint point detection device, learning model generation method, joint point detection method, and program

The learning model generation device improves 3D joint point coordinate detection accuracy by using a basic model to understand 2D-3D relationships, inputting noisy data, and updating parameters to correct for noise, enhancing precision in 3D coordinate estimation.

JP7861866B2Active Publication Date: 2026-05-19NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2023-12-18
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Conventional learning models struggle with low accuracy in detecting 3D joint point coordinates due to the use of noisy 2D joint point coordinates as training data, making it difficult to determine noise and maintain high detection accuracy.

Method used

A learning model generation device and method that utilizes a basic learning model to understand the relationship between 2D and 3D joint point coordinates, inputs second sample data with differences, calculates the difference in intermediate features, and updates the model's parameters to account for noise, improving detection accuracy.

Benefits of technology

The updated model enhances the detection accuracy of 3D joint point coordinates by accounting for noise in 2D input data, resulting in more precise 3D coordinate estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007861866000001
    Figure 0007861866000001
  • Figure 0007861866000002
    Figure 0007861866000002
  • Figure 0007861866000003
    Figure 0007861866000003
Patent Text Reader

Abstract

A learning model generation device 10 is provided with: a first data input unit 11 that inputs first sample data to a basic learning model that has been trained through machine learning on the relationship between two-dimensional joint point coordinates and three-dimensional joint point coordinates; a second data input unit 12 that inputs second sample data having differences from the first sample data to a learning model to be updated; a difference calculation unit 13 that calculates the difference between an intermediate feature quantity of the basic learning model when the first sample data is input thereto and an intermediate feature quantity of the learning model to be updated when the second sample data is input thereto; and a parameter update unit 14 that uses the calculated difference to update the parameters of the learning model to be updated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a learning model generation device and a learning model generation method for generating a learning model for detecting human joint points, and further to a program for realizing these. Mu Furthermore, the present disclosure relates to a joint point detection device and a joint point detection method for detecting human joint points, and further to a program for realizing these. Mu Related.

Background Art

[0002] In recent years, techniques for estimating a human posture by estimating the three-dimensional coordinates of each joint of a human from a two-dimensional image have been developed (see, for example, Patent Document 1). Such techniques are expected to be used in fields such as image monitoring systems, sports, and games. In such techniques, a learning model is used to detect the three-dimensional coordinates of each joint of a human.

[0003] The learning model is constructed, for example, by machine learning using, as training data, the two-dimensional coordinates of each joint (hereinafter referred to as "two-dimensional joint point coordinates") extracted from a human in an image and the three-dimensional coordinates of each joint (hereinafter referred to as "three-dimensional joint point coordinates"). In the training data, the three-dimensional joint point coordinates correspond to teacher data.

[0004] Also, machine learning is performed by inputting the two-dimensional joint point coordinates serving as training data into the learning model and updating the parameters of the learning model so that the difference between the output three-dimensional joint point coordinates and the three-dimensional joint point coordinates that are teacher data becomes small.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] Incidentally, in conventional machine learning models, two-dimensional joint point coordinates with minimal noise, such as errors in the position of each joint point, are used as training data. This is because if noisy two-dimensional joint point coordinates are used as training data, the learning model cannot determine whether or not something is noisy, making it impossible to build a learning model with high detection accuracy.

[0007] However, the 2D joint point coordinates extracted from actual 2D images contain noise. Therefore, conventional learning models have difficulty detecting highly accurate 3D joint point coordinates when given 2D joint point coordinates extracted from actual 2D images as input.

[0008] One example of the purpose of this disclosure is to improve the detection accuracy when estimating the 3D coordinates of each joint point from the 2D coordinates of each joint. [Means for solving the problem]

[0009] To achieve the above objective, the learning model generation device in one aspect of this disclosure is: The first data input unit inputs the first sample data into a basic learning model that uses machine learning to understand the relationship between 2D and 3D joint point coordinates. A second data input unit inputs second sample data, which has differences from the first sample data, into the learning model to be updated. A difference calculation unit calculates the difference between the intermediate features of the basic learning model when the first sample data is input and the intermediate features of the learning model to be updated when the second sample data is input. A parameter update unit updates the parameters of the learning model to be updated using the calculated difference, It is characterized by having the following features.

[0010] To achieve the above objective, the joint point detection device in one aspect of this disclosure is: The learning model is equipped with a joint point detection unit that takes the 2D joint point coordinates of a person as input and detects the 3D joint point coordinates of the person. The parameters of the aforementioned learning model are: The basic machine learning model, which is learning the relationship between 2D and 3D joint point coordinates, receives the first sample data as input, and the intermediate features are... The learning model receives intermediate features when a second sample data set having differences from the first sample data set is input, The updated version is obtained using the difference. It is characterized by the following:

[0011] To achieve the above objective, the learning model generation method in one aspect of this disclosure is: The first data input step involves inputting the first sample data into a basic learning model that uses machine learning to understand the relationship between 2D and 3D joint point coordinates. A second data input step involves inputting a second sample data set, which has differences from the first sample data set, into the learning model to be updated. A difference calculation step, which calculates the difference between the intermediate features of the basic learning model when the first sample data is input and the intermediate features of the learning model to be updated when the second sample data is input. A parameter update step in which the parameters of the learning model to be updated are updated using the calculated difference, It is characterized by having the following:

[0012] To achieve the above objective, the joint point detection method in one aspect of this disclosure is: The learning model has a joint point detection step in which the 2D joint point coordinates of a person are input and the 3D joint point coordinates of the person are detected. The parameters of the aforementioned learning model are: For a basic learning model that is learning the relationship between 2D joint coordinates and 3D joint coordinates, the intermediate feature quantity when the first sample data is input, and the intermediate feature quantity when the second sample data having a difference point from the first sample data is input to the learning model, and using the difference therebetween for updating, characterized in that.

[0013] Furthermore, to achieve the above object, a first program in one aspect of the present disclosure is to cause a computer to a first data input step of inputting first sample data to a basic learning model that is learning the relationship between 2D joint coordinates and 3D joint coordinates; a second data input step of inputting second sample data having a difference point from the first sample data to a learning model to be updated; a difference calculation step of calculating the difference between the intermediate feature quantity of the basic learning model when the first sample data is input and the intermediate feature quantity of the learning model to be updated when the second sample data is input; a parameter update step of updating the parameters of the learning model to be updated using the calculated difference; and causing ru, characterized in that.

[0014] Furthermore, to achieve the above object, a second program in one aspect of the present disclosure is to cause a computer to execute a joint point detection step of inputting 2D joint coordinates of a person to a learning model to detect 3D joint coordinates of the person height, The parameters of the learning model are the intermediate feature quantity when the first sample data is input to a basic learning model that is learning the relationship between 2D joint coordinates and 3D joint coordinates, and beforeThe learning model receives intermediate features when a second sample data set, which has differences from the first sample data set, is input. It is characterized by being updated using the difference. [Effects of the Invention]

[0015] As described above, this disclosure makes it possible to improve the detection accuracy when estimating the 3D coordinates of each joint point from the 2D coordinates of each joint. [Brief explanation of the drawing]

[0016] [Figure 1] Figure 1 is a schematic diagram showing an example of a learning model generation device. [Figure 2] Figure 2 illustrates the function of the basic learning model. [Figure 3] Figure 3 is a diagram illustrating the configuration of an example of a learning model generation device. [Figure 4] Figure 4 is an explanatory diagram illustrating the first and second sample data sets. [Figure 5] Figure 5 is a diagram showing an example of the configuration of the base learning model and the update model. [Figure 6] Figure 6 is a flowchart illustrating an example of the operation of a learning model generation device. [Figure 7] Figure 7 is a configuration diagram showing another example of a learning model generation device. [Figure 8] Figure 8 is a flowchart illustrating another example of the operation of the learning model generation device. [Figure 9] Figure 9 is a configuration diagram showing an example of a joint point detection device. [Figure 10] Figure 10 is a flowchart showing an example of the operation of the joint point detection device 50. [Figure 11] Figure 11 is a block diagram showing an example of a computer that implements a learning model generation device and an articulation point detection device. [Modes for carrying out the invention]

[0017] (Embodiment 1) The following describes an example of a learning model generation device, a learning model generation method, and a program, with reference to Figures 1 to 6.

[0018] [Device configuration] First, we will explain the schematic configuration of an example of a learning model generation device using Figures 1 and 2. Figure 1 is a configuration diagram showing the schematic configuration of an example of a learning model generation device. Figure 2 is a diagram illustrating the function of the basic learning model.

[0019] As shown in Figure 1, the learning model generation device 10 is a device for generating a learning model that takes two-dimensional human joint point coordinates as input and outputs corresponding three-dimensional human joint point coordinates. As shown in Figure 1, the learning model generation device 10 includes a first data input unit 11, a second data input unit 12, a difference calculation unit 13, and a parameter update unit 14.

[0020] The first data input unit 11 inputs the first sample data into the basic learning model. The basic learning model is a learning model that uses machine learning to determine the relationship between 2D joint point coordinates and 3D joint point coordinates. The second data input unit 12 inputs the second sample data, which has differences from the first sample data, into the learning model to be updated (hereinafter referred to as the "updated model").

[0021] The basic learning model will be explained using Figure 2. As shown in Figure 2, when the basic learning model receives the 2D joint point coordinates of a person's joint points as input, the basic learning model outputs the 3D joint point coordinates of the corresponding joint points of the person. The 2D joint point coordinates are extracted in advance from human image data, for example, using a machine learning model that has learned the relationship between image features and joint points.

[0022] The difference calculation unit 13 calculates the difference between the intermediate features of the base learning model when the first sample data is input and the intermediate features of the learning model to be updated when the second sample data is input. The parameter update unit 14 updates the parameters of the learning model to be updated using the difference calculated by the difference calculation unit 13.

[0023] Thus, in Embodiment 1, the parameters of the updated model are updated using the difference between the intermediate features of the basic learning model and the intermediate features of the updated model. As a result, machine learning that takes noise into account is performed in the updated model. Therefore, the updated model obtained in Embodiment 1 can improve the detection accuracy when estimating the 3D coordinates of each joint point from the 2D coordinates of each joint.

[0024] Next, in addition to Figures 1 and 2, Figures 3 to 5 will be used to specifically explain the configuration and function of an example of a learning model generation device. Figure 3 is a configuration diagram specifically showing the configuration of an example of a learning model generation device. Figure 4 is an explanatory diagram for explaining the first sample data and the second sample data. Figure 5 is a configuration diagram showing an example of the configuration of a base learning model and an update model.

[0025] As shown in Figure 3, the learning model generation device 10 implements a base learning model 20 and an update model 30. In Embodiment 1, the base learning model 20 and the update model 30 are, for example, neural networks. In practice, the base learning model 20 and the update model 30 are implemented by machine learning programs executed on a computer. Note that the base learning model 20 and the update model 30 may be built outside the learning model generation device 10.

[0026] Furthermore, as shown in Figure 4, the first sample data is 2D joint point data containing 2D joint point coordinates detected from a specific person image. In contrast, the second sample data is also 2D joint point data containing 2D joint point coordinates detected from a specific person image. However, as shown in Figure 4, the second sample data differs from the first sample data in all or part of the joint point coordinate values. In other words, the second sample data is the first sample data with noise added to it.

[0027] Furthermore, in the example shown in Figure 2, the second data input unit 12 acquires pre-prepared second sample data from an external source, but Embodiment 1 is not limited to this configuration. In Embodiment 1, the second data input unit 12 can also acquire first sample data and generate second sample data by adding noise to the acquired first sample data.

[0028] Specifically, the second data input unit 12 randomly selects one of the joint points of the first sample data, randomly moves the position of the selected joint point to another position, and generates the second sample data (see Figure 4).

[0029] Furthermore, as shown in Figure 5, in Embodiment 1, the base learning model 20 and the update model 30 are neural networks as described above. The base learning model 20 comprises, for example, an input layer 21, an intermediate layer (hidden layer) 22, and an output layer 23. Similarly, the update model 30 also comprises an input layer 31, an intermediate layer (hidden layer) 32, and an output layer 33.

[0030] In Embodiment 1, the difference calculation unit 13 obtains intermediate features from the intermediate layer 22 of the basic learning model 20, and further obtains intermediate features from the intermediate layer 32 of the updated model 30. The difference calculation unit 13 then calculates the difference between the intermediate features obtained from the intermediate layer 22 and the intermediate features obtained from the intermediate layer 32.

[0031] In Embodiment 1, when the difference is calculated by the difference calculation unit 13, the parameter update unit 14 updates the parameters of the update model 30, i.e., the weights of each node, so that the calculated difference becomes smaller.

[0032] [Device operation] Next, an example of the operation of the learning model generation device 10 will be explained using Figure 6. Figure 6 is a flowchart showing an example of the operation of the learning model generation device. In the following explanation, Figures 1 to 5 will be referred to as appropriate. In Embodiment 1, the learning model generation method is carried out by operating the learning model generation device. Therefore, the explanation of the learning model generation method in Embodiment 1 will be replaced by the following explanation of the operation of the learning model generation device 10.

[0033] As shown in Figure 6, first, the first data input unit 11 acquires the first sample data and inputs the acquired first sample data into the basic learning model 20 (step A1).

[0034] Next, the second data input unit 12 acquires second sample data that has differences from the first sample data, and inputs the acquired second sample data into the update model (step A2).

[0035] In the first embodiment, in step A2, the second data input unit 12 may acquire the first sample data instead of the second sample data, and generate the second sample data from the acquired first sample data.

[0036] Next, the difference calculation unit 13 obtains intermediate features from the intermediate layer 22 of the basic learning model 20 (step A3). The difference calculation unit 13 also obtains intermediate features from the intermediate layer 32 of the updated model 30 (step A4).

[0037] Next, the difference calculation unit 13 calculates the difference between the intermediate features obtained from the intermediate layer 22 in step A3 and the intermediate features obtained from the intermediate layer 32 in step A4 (step A5).

[0038] Subsequently, the parameter update unit 14 obtains the difference calculated by the difference calculation unit 13 in step A5 and updates the parameters of the updated model 30 so that this difference becomes smaller (step A6). In this embodiment 1, the updated model is completed when the parameters of the updated model are updated. In other words, the updated model is generated when the parameters of the updated model are updated.

[0039] Steps A1 to A6 are repeated for each sample data set (first and second). Ideally, the input sample data sets (first and second) should have no time difference between them. A smaller time difference makes the noise more noticeable and easier to correct.

[0040] As described above, in Embodiment 1, a learning model trained using low-noise training data is used as the base learning model 20. Furthermore, the parameters of the updated model 30 are updated using the difference between the intermediate features of the base learning model 20 and the intermediate features of the updated model 30. As a result, the parameters of the updated model 30 are updated taking into account the noise contained in the second sample data. Therefore, with the updated model 30 whose parameters have been updated in Embodiment 1, the detection accuracy when estimating the 3D coordinates of each joint point from the 2D coordinates of each joint can be improved.

[0041] [Differentiation] Here, a modified version of Embodiment 1 will be described. In this modified version, the specific person image described above is the image of each consecutive frame of video data in which the person was filmed. In this case, the first sample data and the second sample data will each be created for each frame of the video data.

[0042] The first data input unit 11 inputs N first sample data points into the base learning model 20 in accordance with the time series of the frames. The second data input unit 12 inputs N second sample data points into the update model 30 in synchronization with the first data input unit 11. N is an arbitrary integer.

[0043] In this case, the base learning model 20 and the update model 30 may output N output data corresponding to each of the N sample data, or they may output only one output data (for example, the output data corresponding to the central frame).

[0044] The difference calculation unit 13 obtains intermediate features from the intermediate layer 22 of the basic learning model 20 when N first sample data are input. The difference calculation unit 13 also obtains intermediate features from the intermediate layer 32 of the updated model 30 when N second sample data are input. Then, the difference calculation unit 13 calculates the difference between the intermediate features obtained from the intermediate layer 22 and the intermediate features obtained from the intermediate layer 32.

[0045] Subsequently, the parameter update unit 14 updates the parameters of the updated model 30 so that the difference calculated by the difference calculation unit 13 becomes smaller, similar to the case where video data is not used.

[0046] Thus, in this modified example, multiple consecutive frames are input to each learning model. Furthermore, the second sample data may contain a mix of frames with and without noise. In this case, the updated model 30 performs noise interpolation between frames, resulting in a more pronounced noise correction through parameter updates.

[0047] [program] The program in Embodiment 1 can be any program that causes a computer to execute steps A1 to A6 shown in Figure 6. By installing and running this program on a computer, the learning model generation device 10 and the learning model generation method in Embodiment 1 can be realized. In this case, the computer's processor functions as a first data input unit 11, a second data input unit 12, a difference calculation unit 13, and a parameter update unit 14, and performs processing. Examples of computers include general-purpose PCs, smartphones, and tablet devices.

[0048] The program in Embodiment 1 may be executed by a computer system constructed by multiple computers. In this case, for example, each computer may function as one of the following: the first data input unit 11, the second data input unit 12, the difference calculation unit 13, and the parameter update unit 14.

[0049] (Embodiment 2) Below, another example of a learning model generation device, learning model generation method, and program will be described with reference to Figures 7 and 8.

[0050] [Device configuration] First, we will explain the configuration of another example of a learning model generation device using Figure 7. Figure 7 is a configuration diagram showing the configuration of another example of a learning model generation device.

[0051] The learning model generation device 40 shown in Figure 7 is also a device for generating a learning model, similar to Embodiment 1, which takes two-dimensional human joint point coordinates as input and outputs corresponding three-dimensional human joint point coordinates.

[0052] As shown in Figure 7, the learning model generation device 40, like the learning model generation device 10 shown in Embodiment 1, includes a first data input unit 11, a second data input unit 12, a difference calculation unit 13, and a parameter update unit 14.

[0053] However, unlike the learning model generation device 10 shown in Embodiment 1, the learning model generation device 40 also includes a second difference calculation unit 41 and a statistical processing unit 42 in addition to the above-described configuration. The differences from Embodiment 1 will be explained below.

[0054] The second difference calculation unit 41 calculates the difference between the correct data and the output data of the updated model 30 when the second sample data is input, as the second difference. The correct data is data that includes the correct 3D joint point coordinates for the first sample data. In other words, the correct data is data composed of the actual 3D coordinates of the joint points of the person in the image from which the first and second sample data were extracted.

[0055] The statistical processing unit 42 performs statistical processing using the difference calculated by the difference calculation unit 13 and the second difference calculated by the second difference calculation unit 41. The statistical processing in this case is not particularly limited. A specific example of statistical processing is a weighted average.

[0056] In Embodiment 2, the parameter update unit 14 updates the parameters of the updated model 30 based on the results of statistical processing by the statistical processing unit 42. Specifically, the parameter update unit 14 updates the parameters of the updated model 30, i.e., the weights, so that the values ​​obtained by the statistical processing by the statistical processing unit 42 (for example, the weighted mean) become smaller.

[0057] Furthermore, in Embodiment 2, the first sample data and the second sample data may be input for each frame. In this case, the difference calculation by the difference calculation unit 13, the calculation of the second difference by the second difference calculation unit 41, the statistical processing by the statistical processing unit 42, and the update by the parameter update unit 14 are performed for each frame.

[0058] [Device operation] Next, another example of the operation of the learning model generation device 40 will be explained using Figure 8. Figure 8 is a flowchart showing another example of the operation of the learning model generation device. In the following explanation, Figure 7 will be referred to as appropriate. In Embodiment 2, the learning model generation method is carried out by operating the learning model generation device. Therefore, the explanation of the learning model generation method in Embodiment 2 will be replaced by the following explanation of the operation of the learning model generation device 4.

[0059] As shown in Figure 8, first, the first data input unit 11 acquires the first sample data and inputs the acquired first sample data into the basic learning model 20 (step B1).

[0060] Next, the second data input unit 12 acquires second sample data that has differences from the first sample data, and inputs the acquired second sample data into the update model (step B2).

[0061] Next, the difference calculation unit 13 obtains intermediate features from the intermediate layer 22 of the basic learning model 20 (step B3). The difference calculation unit 13 also obtains intermediate features from the intermediate layer 32 of the updated model 30 (step B4).

[0062] Next, the difference calculation unit 13 calculates the difference between the intermediate features obtained from the intermediate layer 22 in step A3 and the intermediate features obtained from the intermediate layer 32 in step A4 (step B5). Steps B1 to B5 described above are substantially the same as steps A1 to A5 shown in Figure 6 in Embodiment 1.

[0063] Next, in Embodiment 2, the second difference calculation unit 41 obtains correct answer data from an external source, obtains output data when the second sample data is input from the update model 30, and calculates the difference between the correct answer data and the output data of the update model 30 as the second difference (step B6).

[0064] Next, the statistical processing unit 42 performs statistical processing using the difference calculated in step B5 and the second difference calculated in step B6 (step B7).

[0065] Subsequently, the parameter update unit 14 updates the parameters of the updated model 30 based on the results of the statistical processing in step B7 (step B8).

[0066] Steps B1 to B8 are repeated for the number of first and second sample data, similar to Embodiment 1. In Embodiment 2 as well, if the first and second sample data are obtained for each frame of the video data, steps B1 to B8 are repeated for each frame.

[0067] As described above, in Embodiment 2, the parameters of the updated model 30 are updated using the difference (second difference) between the output data of the updated model and the ground truth data. As a result, in Embodiment 2, noise contained in the second sample data is taken into account even more. According to Embodiment 2, the detection accuracy when estimating the 3D coordinates of each joint point from the 2D coordinates of each joint can be further improved.

[0068] Furthermore, in Embodiment 2, as described in the modified version of Embodiment 1, the specific person image mentioned above may be an image from each consecutive frame of video data in which a person was filmed. In this case, in Embodiment 2 as well, the first data input unit 11 inputs N first sample data into the basic learning model 20 in accordance with the time series of the frames. The second data input unit 12 inputs N second sample data into the update model 30 in synchronization with the first data input unit 11.

[0069] In this case, the base learning model 20 and the update model 30 may output N output data corresponding to each of the N sample data, or they may output only one output data (for example, the output data corresponding to the central frame).

[0070] Then, in the case where N output data are output, unlike step B6 described above, the second difference calculation unit 41 calculates N differences between the ground truth data and the output data of the updated model 30 for each frame. After that, the second difference calculation unit 41 uses each of the calculated differences to perform statistical processing such as arithmetic mean and weighted mean to calculate a single (scalar) loss, which is taken as the second difference.

[0071] If only one output data is output, the second difference calculation unit 41 calculates the difference between the correct data and the output data of the updated model 30, similar to step B6 described above, and uses the calculated difference as the second difference.

[0072] [program] The program in Embodiment 2 can be any program that causes a computer to execute steps B1 to B8 shown in Figure 8. By installing and running this program on a computer, the learning model generation device 40 and the learning model generation method in Embodiment 2 can be realized. In this case, the computer's processor functions as a first data input unit 11, a second data input unit 12, a difference calculation unit 13, a parameter update unit 14, a second difference calculation unit 41, and a statistical processing unit 42, and performs the processing. Examples of computers include general-purpose PCs, smartphones, and tablet devices.

[0073] The program in Embodiment 2 may be executed by a computer system constructed by multiple computers. In this case, for example, each computer may function as one of the following: the first data input unit 11, the second data input unit 12, the difference calculation unit 13, the parameter update unit 14, the second difference calculation unit 41, and the statistical processing unit 42.

[0074] (Embodiment 3) Next, an example of a joint point detection device, joint point detection method, and program will be described with reference to Figures 9 and 10.

[0075] [Device configuration] First, we will explain the configuration of an example of a joint point detection device using Figure 9. Figure 9 is a configuration diagram showing the configuration of an example of a joint point detection device.

[0076] As shown in Figure 9, the joint point detection device 50 is a device that detects the corresponding 3D joint point coordinates of a person from the 2D joint point coordinates of a person. As shown in Figure 9, the joint point detection device 50 is equipped with a joint point detection unit 51. The joint point detection unit 51 inputs the 2D joint point coordinates of a person into a learning model 60 and detects the 3D joint point coordinates of a person.

[0077] The learning model 60 is a machine learning model that learns the relationship between 2D and 3D joint point coordinates. The parameters of the learning model 60 are updated using the difference between the first intermediate feature and the second intermediate feature. The first intermediate feature is the intermediate feature obtained when the first sample data is input to the base learning model that learns the relationship between 2D and 3D joint point coordinates. The second intermediate feature is the intermediate feature obtained when the learning model 60 is input to the second sample data, which has differences from the first sample data.

[0078] Specifically, in Embodiment 3, the learning model 60 is a neural network, and is the learning model generated by the learning model generation device in Embodiment 1 or 2, i.e., the updated model 30 (see Figures 1, 3, and 7). The learning model 60 is also implemented by a machine learning program executed on a computer.

[0079] In this way, the joint point detection device 50 detects the three-dimensional joint point coordinates of a person using a learning model generated by the learning model generation device in Embodiment 1 or 2.

[0080] [Device operation] Next, an example of the operation of the joint point detection device 50 will be explained using Figure 10. Figure 10 is a flowchart showing an example of the operation of the joint point detection device 50. In the following explanation, Figure 9 will be referred to as appropriate. In Embodiment 3, the joint point detection method is performed by operating the joint point detection device 50. Therefore, the explanation of the joint point detection method in Embodiment 3 will be replaced by the following explanation of the operation of the joint point detection device 50.

[0081] As shown in Figure 10, first, the joint point detection unit 51 obtains the 2D coordinates of the joint points of a person, which will be used as input data (step C1). Next, the joint point detection unit 51 inputs the 2D coordinates obtained in step C1 into the learning model 60 (step C2).

[0082] Next, the joint point detection unit 51 acquires the 3D coordinates of the joint points output by the learning model 60 (step C3). In Embodiment 3, step C3 indicates that the 3D coordinates of the joint points of the person have been detected.

[0083] Next, the joint point detection unit 51 outputs the 3D coordinates of the joint points acquired in step C3 to the outside (step C4).

[0084] As described above, according to Embodiment 3, the 3D coordinates of the joint points can be detected using the learning model generated by Embodiment 1 or 2. Therefore, the 3D coordinates of the joint points are detected taking into account the noise in the 2D coordinate input data, resulting in highly accurate values.

[0085] [program] The program in Embodiment 3 can be any program that causes the computer to execute steps C1 to C4 shown in Figure 10. By installing and executing this program on the computer, the joint point detection device 50 and joint point detection method in this embodiment can be realized. In this case, the computer's processor is Joint point detection unit 51 It functions and processes data. Examples of computers include general-purpose PCs, smartphones, and tablet devices.

[0086] Furthermore, the program in Embodiment 3 may be executed by a computer system constructed by multiple computers. In this case, the multiple computers function as the joint point detection unit 51.

[0087] [Physical configuration] Here, a computer that implements a learning model generation device and a joint point detection device by executing the programs in Embodiments 1 to 3 will be described with reference to Figure 11. Figure 11 is a block diagram showing an example of a computer that implements a learning model generation device and a joint point detection device.

[0088] As shown in Figure 11, the computer 110 comprises a CPU (Central Processing Unit) 111, main memory 112, storage device 113, input interface 114, display controller 115, data reader / writer 116, and communication interface 117. Each of these components is connected to the others via a bus 121, enabling data communication.

[0089] Furthermore, the computer 110 may include a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array) in addition to, or instead of, the CPU 111. In this embodiment, the GPU or FPGA can execute the program in the embodiment.

[0090] The CPU 111 loads the program in the embodiment, which consists of a set of codes stored in the storage device 113, into the main memory 112, and performs various calculations by executing each code in a predetermined order. The main memory 112 is typically a volatile storage device such as DRAM (Dynamic Random Access Memory).

[0091] Furthermore, the programs in Embodiments 1 to 3 are provided stored on a computer-readable recording medium 120. The programs in these embodiments may also be distributed over the Internet via a communication interface 117.

[0092] Specific examples of the storage device 113 include hard disk drives and semiconductor storage devices such as flash memory. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and mouse. The display controller 115 is connected to the display device 119 and controls the display on the display device 119.

[0093] The data reader / writer 116 mediates data transmission between the CPU 111 and the recording medium 120, reads programs from the recording medium 120, and writes processing results from the computer 110 to the recording medium 120. The communication interface 117 mediates data transmission between the CPU 111 and other computers.

[0094] Furthermore, specific examples of the recording medium 120 include general-purpose semiconductor memory devices such as CF (Compact Flash®) and SD (Secure Digital), magnetic recording media such as Flexible Disks, or optical recording media such as CD-ROMs (Compact Disk Read Only Memory).

[0095] Furthermore, the learning model generation device and joint point detection device in the embodiment can be implemented not by a computer with a program installed, but by using hardware corresponding to each part, such as electronic circuits. Moreover, the learning model generation device and joint point detection device in the embodiment may be partially implemented by a program and the remaining part by hardware. In the embodiment, the computer is not limited to the computer shown in Figure 11.

[0096] Some or all of the embodiments described above can be expressed by (Appendix 1) to (Appendix 18) described below, but are not limited to the following descriptions.

[0097] (Note 1) The first data input unit inputs the first sample data into a basic learning model that uses machine learning to understand the relationship between 2D and 3D joint point coordinates. A second data input unit inputs second sample data, which has differences from the first sample data, into the learning model to be updated. A difference calculation unit calculates the difference between the intermediate features of the basic learning model when the first sample data is input and the intermediate features of the learning model to be updated when the second sample data is input. A parameter update unit updates the parameters of the learning model to be updated using the calculated difference, A learning model generation device characterized by having the following features.

[0098] (Note 2) The first sample data includes two-dimensional joint point coordinates detected from a specific person image, The second sample data includes two-dimensional joint point coordinates detected from the specific person image, and all or some of the joint point coordinate values ​​differ from those of the first sample data. The learning model generation device described in Appendix 1.

[0099] (Note 3) A second difference calculation unit calculates the difference between the ground truth data, which includes the correct 3D joint point coordinates for the first sample data, and the output data of the learning model to be updated when the second sample data is input, as the second difference. A statistical processing unit performs statistical processing using the difference calculated by the difference calculation unit and the second difference calculated by the second difference calculation unit. Equipped with, The parameter update unit updates the parameters of the learning model to be updated based on the results of the statistical processing. The learning model generation device described in Appendix 2.

[0100] (Note 4) The aforementioned specific person image is an image from each of the consecutive frames of video data in which the person was filmed, and the first sample data and the second sample data are each composed of a plurality of the aforementioned frames. The first data input unit inputs a plurality of the first sample data along the time series of the frame, The second data input unit inputs a plurality of the second sample data in synchronization with the first data input unit. A learning model generation device as described in Appendix 2 or 3.

[0101] (Note 5) The second data input unit generates the second sample data by adding noise to the first sample data, and inputs the generated second sample data into the basic learning model. The learning model generation device described in Appendix 1.

[0102] (Note 6) The learning model is equipped with a joint point detection unit that takes the 2D joint point coordinates of a person as input and detects the 3D joint point coordinates of the person. The parameters of the aforementioned learning model are: The basic machine learning model, which is learning the relationship between 2D and 3D joint point coordinates, receives the first sample data as input, and the intermediate features are... The learning model receives intermediate features when a second sample data set having differences from the first sample data set is input, The updated version is obtained using the difference. A joint point detection device characterized by the following features.

[0103] (Note 7) The first data input step involves inputting the first sample data into a basic learning model that uses machine learning to understand the relationship between 2D and 3D joint point coordinates. A second data input step involves inputting a second sample data set, which has differences from the first sample data set, into the learning model to be updated. A difference calculation step, which calculates the difference between the intermediate features of the basic learning model when the first sample data is input and the intermediate features of the learning model to be updated when the second sample data is input. A parameter update step in which the parameters of the learning model to be updated are updated using the calculated difference, A method for generating a learning model, characterized by having [a certain feature].

[0104] (Note 8) The first sample data includes two-dimensional joint point coordinates detected from a specific person image, The second sample data includes two-dimensional joint point coordinates detected from the specific person image, and all or some of the joint point coordinate values ​​differ from those of the first sample data. The learning model generation method described in Appendix 7.

[0105] (Note 9) A second difference calculation step, which calculates the difference between the ground truth data, which includes the 3D joint point coordinates that are correct for the first sample data, and the output data of the learning model to be updated when the second sample data is input, as the second difference; A statistical processing step, which performs statistical processing using the difference calculated by the difference calculation step and the second difference calculated by the second difference calculation step, It further possesses, In the parameter update step, the parameters of the learning model to be updated are updated based on the results of the statistical processing. The learning model generation method described in Appendix 8.

[0106] (Note 10) The aforementioned specific person image is an image from each of the consecutive frames of video data in which the person was filmed, and the first sample data and the second sample data are each composed of a plurality of the aforementioned frames. In the first data input step, a plurality of the first sample data are input along the time series of the frame, In the second data input step, a plurality of the second sample data are input in synchronization with the first data input step. The learning model generation method described in Appendix 8 or 9.

[0107] (Note 11) In the second data input step, noise is added to the first sample data to generate the second sample data, and the generated second sample data is input to the basic learning model. The learning model generation method described in Appendix 7.

[0108] (Note 12) The learning model has a joint point detection step in which the 2D joint point coordinates of a person are input and the 3D joint point coordinates of the person are detected. The parameters of the aforementioned learning model are: The basic machine learning model, which is learning the relationship between 2D and 3D joint point coordinates, receives the first sample data as input, and the intermediate features are... The learning model receives intermediate features when a second sample data set having differences from the first sample data set is input, The updated version is obtained using the difference. A method for detecting joint points, characterized by the features described above.

[0109] (Note 13) On the computer, The first data input step involves inputting the first sample data into a basic learning model that uses machine learning to understand the relationship between 2D and 3D joint point coordinates. A second data input step involves inputting a second sample data set, which has differences from the first sample data set, into the learning model to be updated. A difference calculation step, which calculates the difference between the intermediate features of the basic learning model when the first sample data is input and the intermediate features of the learning model to be updated when the second sample data is input. A parameter update step in which the parameters of the learning model to be updated are updated using the calculated difference, Let's execute it ru, Professional Hmm.

[0110] (Note 14) The first sample data includes two-dimensional joint point coordinates detected from a specific person image, The second sample data includes two-dimensional joint point coordinates detected from the specific person image, and all or some of the joint point coordinate values ​​differ from those of the first sample data. As described in Appendix 13 program .

[0111] (Note 15) before On the computer, A second difference calculation step, which calculates the difference between the ground truth data, which includes the 3D joint point coordinates that are correct for the first sample data, and the output data of the learning model to be updated when the second sample data is input, as the second difference; A statistical processing step, which performs statistical processing using the difference calculated by the difference calculation step and the second difference calculated by the second difference calculation step, of Furthermore Execute height, In the parameter update step, the parameters of the learning model to be updated are updated based on the results of the statistical processing. As described in Appendix 14 program .

[0112] (Note 16) The aforementioned specific person image is an image from each of the consecutive frames of video data in which the person was filmed, and the first sample data and the second sample data are each composed of a plurality of the aforementioned frames. In the first data input step, a plurality of the first sample data are input along the time series of the frame, In the second data input step, a plurality of the second sample data are input in synchronization with the first data input step. As described in Appendix 14 or 15 program .

[0113] (Note 17) In the second data input step, noise is added to the first sample data to generate the second sample data, and the generated second sample data is input to the basic learning model. As described in Appendix 13 program .

[0114] (Note 18) On the computer, The learning model is given the 2D joint point coordinates of a person, and a joint point detection step is performed to detect the 3D joint point coordinates of the person. height, The parameters of the aforementioned learning model are: The basic machine learning model, which is learning the relationship between 2D and 3D joint point coordinates, receives the first sample data as input, and the intermediate features are... before The learning model receives intermediate features when a second sample data set, which has differences from the first sample data set, is input. The updated version is obtained using the difference. Characterized by program .

[0115] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the configuration and details of the present disclosure can be made that can be understood by those skilled in the art within the scope of the present disclosure.

[0116] This application claims priority based on Japanese Patent Application No. 2022-211356, filed on 28 December 2022, and incorporates all of its disclosures herein. [Industrial applicability]

[0117] As described above, this disclosure makes it possible to improve the detection accuracy when estimating the 3D coordinates of each joint point from the 2D coordinates of each joint. This disclosure is useful for systems that require the estimation of a person's posture from an image, such as video surveillance systems. [Explanation of symbols]

[0118] 10. Learning Model Generation Device (Embodiment 1) 11. First data input section 12. Second data input section 13 Difference calculation part 14 Parameter update section 20 Basic Learning Models 21 Input Layer 22. Intermediate layer (hidden layer) 23 Output Layer 30 Updated Model 31 Input Layer 32. Intermediate layer (hidden layer) 33 Output Layer 40 Learning Model Generation Device (Embodiment 2) 41 Second difference calculation unit 42 Statistical Processing Unit 50 Joint Point Detection Device 51 Joint point detection unit 60 Learning Models 111 CPU 112 Main Memory 113 Storage device 114 Input Interface 115 Display Controller 116 Data Readers / Writers 117 Communication Interface 118 Input devices 119 Display device 120 recording media 121 Bus

Claims

1. The first data input unit inputs the first sample data into a basic learning model that uses machine learning to understand the relationship between 2D and 3D joint point coordinates. A second data input unit inputs second sample data, which has differences from the first sample data, into the learning model to be updated. A difference calculation unit calculates the difference between the intermediate features of the basic learning model when the first sample data is input and the intermediate features of the learning model to be updated when the second sample data is input. A parameter update unit updates the parameters of the learning model to be updated using the calculated difference, A learning model generation device characterized by having the following features.

2. The first sample data includes two-dimensional joint point coordinates detected from a specific person image, The second sample data includes two-dimensional joint point coordinates detected from the specific person image, and all or some of the joint point coordinate values ​​differ from those of the first sample data. A learning model generation device according to claim 1.

3. A second difference calculation unit calculates the difference between the ground truth data, which includes the correct three-dimensional joint point coordinates for the first sample data, and the output data of the learning model to be updated when the second sample data is input, as the second difference. A statistical processing unit performs statistical processing using the difference calculated by the difference calculation unit and the second difference calculated by the second difference calculation unit. Equipped with, The parameter update unit updates the parameters of the learning model to be updated based on the results of the statistical processing. The learning model generation apparatus according to claim 2.

4. The aforementioned specific person image is an image from each of the consecutive frames of video data in which the person was filmed, and the first sample data and the second sample data are each composed of a plurality of the aforementioned frames. The first data input unit inputs a plurality of the first sample data along the time series of the frame, The second data input unit inputs a plurality of the second sample data in synchronization with the first data input unit. A learning model generation device according to claim 2 or 3.

5. The second data input unit generates the second sample data by adding noise to the first sample data, and inputs the generated second sample data to the learning model to be updated. A learning model generation device according to claim 1.

6. The system includes a joint point detection unit that takes the two-dimensional joint point coordinates of a person as input to a learning model and detects the three-dimensional joint point coordinates of the person. The parameters of the aforementioned learning model are: The basic learning model, which uses machine learning to determine the relationship between 2D and 3D joint point coordinates, receives the first sample data as input, and the intermediate features are... The learning model receives intermediate features when a second sample data set having differences from the first sample data set is input, The updated version is obtained using the difference. A joint point detection device characterized by the following features.

7. A method performed by a computer, The first sample data is input into a basic machine learning model that is learning the relationship between 2D and 3D joint point coordinates. The second sample data, which has differences from the first sample data, is input into the learning model to be updated. The difference between the intermediate features of the basic learning model when the first sample data is input and the intermediate features of the learning model to be updated when the second sample data is input is calculated. The calculated difference is used to update the parameters of the learning model to be updated. A method for generating a learning model characterized by the following features.

8. A method performed by a computer, The learning model is input with the 2D joint point coordinates of a person to detect the 3D joint point coordinates of the person. The parameters of the aforementioned learning model are: The basic learning model, which uses machine learning to determine the relationship between 2D and 3D joint point coordinates, receives the first sample data as input, and the intermediate features are... The learning model receives intermediate features when a second sample data set having differences from the first sample data set is input, The updated version is obtained using the difference. A method for detecting joint points, characterized by the features described above.

9. On the computer, The first step involves inputting the first sample data into a basic machine learning model that has learned the relationship between 2D and 3D joint point coordinates. The steps include: inputting a second sample data set, which has differences from the first sample data set, into the learning model to be updated; A step of calculating the difference between the intermediate features of the basic learning model when the first sample data is input and the intermediate features of the learning model to be updated when the second sample data is input. The steps include updating the parameters of the learning model to be updated using the calculated difference, A program that executes something.

10. On the computer, The learning model is given the two-dimensional joint point coordinates of a person and is instructed to perform a joint point detection step to detect the three-dimensional joint point coordinates of the person. The parameters of the aforementioned learning model are: The basic learning model, which uses machine learning to determine the relationship between 2D and 3D joint point coordinates, receives the first sample data as input, and the intermediate features are... The learning model receives intermediate features when a second sample data set having differences from the first sample data set is input, The updated version is obtained using the difference. program.