Learning model generation device, learning model generation method, and program

By calculating feature distances and updating model parameters based on camera relationships, the method improves the accuracy of three-dimensional joint point detection in images, addressing the issue of camera angle variations in training data.

JP7852753B2Active Publication Date: 2026-04-28NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2024-01-12
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing learning models for detecting three-dimensional joint points from images suffer from reduced accuracy due to variations in camera angles, leading to overfitting and inaccurate pose estimation when training data is biased by different shooting angles.

Method used

A method that calculates feature distances and similarities based on camera relationships and updates machine learning model parameters to account for variations in camera angles, using two-dimensional and three-dimensional joint point coordinate data to improve detection accuracy.

Benefits of technology

The method enhances the accuracy of detecting three-dimensional joint points by accounting for camera pose variations, resulting in a model that accurately estimates joint point coordinates regardless of shooting angles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852753000009
    Figure 0007852753000009
  • Figure 0007852753000010
    Figure 0007852753000010
  • Figure 0007852753000011
    Figure 0007852753000011
Patent Text Reader

Abstract

A learning model generating device 10 comprises: a data acquiring unit 11 for acquiring two-dimensional joint point coordinate data from among training data that includes the two-dimensional joint point coordinate data, with which it is possible to identify two-dimensional coordinates of each joint point of a person in an image, and three-dimensional joint point coordinate data with which it is possible to identify three-dimensional coordinates of each joint point, the data acquiring unit 11 inputting the two-dimensional joint point coordinate data into a machine learning model; an inter-feature quantity distance calculating unit 12 for acquiring feature quantities calculated using the machine learning model, for each item of training data, and calculating distances between the feature quantities; a similarity calculating unit 13 for using information relating to a camera that captured the image to obtain an index representing a relationship between the person and the camera, for each item of training data, and calculating a similarity between the indices; a loss calculating unit 14 for using the similarities and the distances between the feature quantities to calculate a loss of the feature quantities in the machine learning model; and a learning model generating unit 15 for using the loss to update parameters of the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a learning model generation device and a learning model generation method for generating a learning model for detecting human joint points from an image, and further relates to a program for realizing these. Mu The present disclosure also relates to a joint point detection device and a joint point detection method for detecting human joint points from an image, and further relates to a program for realizing these. Mu Related.

Background Art

[0002] In recent years, a technique for estimating a human posture by detecting three-dimensional coordinates of each joint of a human from a two-dimensional image has been developed (for example, see Patent Document 1). Such a technique is expected to be used in the fields of image monitoring systems, sports, games, and the like. In such a technique, a learning model is used to detect the three-dimensional coordinates of each joint of a human.

[0003] The learning model is constructed, for example, by machine learning using, as training data, two-dimensional coordinates of joints extracted from a person in an image (hereinafter referred to as "two-dimensional joint point coordinates") and three-dimensional coordinates of joints of this person (hereinafter referred to as "three-dimensional joint point coordinates"). In the training data, the three-dimensional joint point coordinates correspond to teacher data.

[0004] Also, machine learning is performed by inputting two-dimensional joint point coordinates serving as training data into the learning model and updating the parameters of the learning model so that the difference between the output three-dimensional joint point coordinates and the three-dimensional joint point coordinates that are teacher data becomes small. In order to improve the detection accuracy of the three-dimensional joint point coordinates by the learning model, it is necessary to prepare a large amount of training data.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] Here, assume a case of photographing a person from the front. In this case, even if the person as the subject is the same person, if the shooting angle (elevation angle, depression angle) of the camera is different, the physique of the person in the image will change, and as a result, the two-dimensional joint point coordinates extracted will also change. Specifically, when the elevation angle or depression angle of the camera is large, the difference between a tall person and a short person in the image becomes small. On the other hand, when the elevation angle or depression angle of the camera is small, the difference between a tall person and a short person in the image becomes large.

[0007] By the way, conventionally, machine learning is performed without considering at all the angle of the camera of the image from which the two-dimensional joint point coordinates are extracted. Therefore, if the shooting angles of the images that are the source of the training data are biased, the learning model will overfit with respect to this shooting angle. In this case, when two-dimensional joint points extracted from images with completely different shooting angles are input, the learning model will output three-dimensional joint point coordinates that are greatly different from the actual person. As a result, the accuracy of pose estimation will decrease.

[0008] An example of the object of the present disclosure is to improve the detection accuracy when detecting the three-dimensional coordinates of joint points from an image.

Means for Solving the Problems

[0009] To achieve the above object, a learning model generation device according to one aspect of the present disclosure acquires the two-dimensional joint point coordinate data that can identify the two-dimensional coordinates of each of a plurality of joint points of a person in an image from training data including the two-dimensional joint point coordinate data that can identify the two-dimensional coordinates of each of the plurality of joint points of the person and the three-dimensional joint point coordinate data that can identify the three-dimensional coordinates of each of the plurality of joint points of the person, and inputs the acquired two-dimensional joint point coordinate data into a machine learning model, a data acquisition unit; A feature distance calculation unit obtains the features calculated in the machine learning model for each of the training data, and further calculates the distance between the obtained features using the obtained features. A similarity calculation unit calculates an index indicating the relationship between the person and the camera that took the image, using information about the camera that took the image for each of the training data, and calculates the similarity between the calculated indices. A loss calculation unit calculates the loss for the features in the machine learning model using the calculated similarity and the calculated distance between the features. A learning model generation unit updates the parameters of the machine learning model using the calculated loss, It is characterized by having the following features.

[0010] Furthermore, in order to achieve the above objective, the learning model generation method in one aspect of this disclosure is: A data acquisition step involves acquiring the 2D joint point coordinate data from training data that includes 2D joint point coordinate data capable of identifying the 2D coordinates of each of several joint points of a person in an image, and 3D joint point coordinate data capable of identifying the 3D coordinates of each of the said multiple joint points of the person, and inputting the acquired 2D joint point coordinate data into a machine learning model. A feature distance calculation step is performed, in which, for each of the training data, the features calculated in the machine learning model are obtained, and further, the distance between the obtained features is calculated using the obtained features. A similarity calculation step is performed for each of the training data, using information about the camera that took each image, to determine an index indicating the relationship between the person and the camera that took the image, and to calculate the similarity between the determined indices. A loss calculation step, which involves calculating the loss for the features in the machine learning model using the calculated similarity and the calculated distance between the features, A learning model generation step in which the parameters of the machine learning model are updated using the calculated loss, It is characterized by having the following:

[0011] Furthermore, in order to achieve the above objectives, the first aspect of this disclosure program teeth, On the computer, A data acquisition step involves acquiring the 2D joint point coordinate data from training data that includes 2D joint point coordinate data capable of identifying the 2D coordinates of each of several joint points of a person in an image, and 3D joint point coordinate data capable of identifying the 3D coordinates of each of the said multiple joint points of the person, and inputting the acquired 2D joint point coordinate data into a machine learning model. A feature distance calculation step is performed, in which, for each of the training data, the features calculated in the machine learning model are obtained, and further, the distance between the obtained features is calculated using the obtained features. A similarity calculation step is performed for each of the training data, using information about the camera that took each image, to determine an index indicating the relationship between the person and the camera that took the image, and to calculate the similarity between the determined indices. A loss calculation step, which involves calculating the loss for the features in the machine learning model using the calculated similarity and the calculated distance between the features, A learning model generation step in which the parameters of the machine learning model are updated using the calculated loss, Let's execute it Ruko It is characterized by the following.

[0012] To achieve the above objective, the joint point detection device in one aspect of this disclosure is: A data acquisition unit obtains 2D joint point coordinate data that can identify the 2D coordinates of each of the multiple joint points of a person in an image. The joint point detection unit applies the acquired 2D joint point coordinate data to a machine learning model that has learned the relationship between the 2D and 3D coordinates of a person's joint points, and detects the 3D coordinates of each of the multiple joint points of the person. Equipped with, The parameters of the aforementioned machine learning model are: In machine learning using 2D joint point coordinate data and 3D joint point coordinate data that can identify the 3D coordinates of multiple joint points in a person, The data is updated using the distance between features obtained from the features calculated by the machine learning model, an index showing the relationship between a person and the camera that took the image, and the loss value for the features in the machine learning model, which is derived from these. It is characterized by the following:

[0013] Furthermore, in order to achieve the above objective, the joint point detection method in one aspect of this disclosure is: A data acquisition step to obtain 2D joint point coordinate data that allows the identification of the 2D coordinates of each of the multiple joint points of a person in an image, The joint point detection step involves applying the acquired 2D joint point coordinate data to a machine learning model that has learned the relationship between the 2D and 3D coordinates of a person's joint points, thereby detecting the 3D coordinates of each of the multiple joint points of the person. Equipped with, The parameters of the aforementioned machine learning model are: In machine learning using 2D joint point coordinate data and 3D joint point coordinate data that can identify the 3D coordinates of multiple joint points in a person, The data is updated using the distance between features obtained from the features calculated by the machine learning model, an index showing the relationship between a person and the camera that took the image, and the loss value for the features in the machine learning model, which is derived from these. It is characterized by the following:

[0014] Furthermore, in order to achieve the above objectives, the second aspect of this disclosure program The computer, A data acquisition step to obtain 2D joint point coordinate data that allows the identification of the 2D coordinates of each of the multiple joint points of a person in an image, The joint point detection step involves applying the acquired 2D joint point coordinate data to a machine learning model that has learned the relationship between the 2D and 3D coordinates of a person's joint points, thereby detecting the 3D coordinates of each of the multiple joint points of the person. Execute height, The parameters of the aforementioned machine learning model are: In machine learning using 2D joint point coordinate data and 3D joint point coordinate data that can identify the 3D coordinates of multiple joint points in a person, The data is updated using the distance between features obtained from the features calculated by the machine learning model, an index showing the relationship between a person and the camera that took the image, and the loss value for the features in the machine learning model, which is derived from these. It is characterized by the following: [Effects of the Invention]

[0015] As described above, this disclosure makes it possible to improve the detection accuracy when detecting the three-dimensional coordinates of joint points from an image. [Brief explanation of the drawing]

[0016] [Figure 1] Figure 1 is a schematic diagram showing an example of a learning model generation device. [Figure 2] Figure 2 is a diagram illustrating the configuration of an example of a learning model generation device. [Figure 3] Figure 3 shows an example of a camera pose vector. [Figure 4] Figure 4 shows an example of a camera pose vector. [Figure 5] Figure 5 shows an example of a camera pose vector. [Figure 6] Figure 6 illustrates an example of machine learning model generation. [Figure 7] Figure 7 is a flowchart showing the operation of an example of a learning model generation device. [Figure 8] Figure 8 is a diagram showing the configuration of another example of a learning model generation device. [Figure 9] Figure 9 is a flowchart illustrating the operation of another example of a learning model generation device. [Figure 10]Figure 10 is a configuration diagram showing yet another example of the configuration of a learning model generation device. [Figure 11] Figure 11 shows an example of a bone length vector and a bone length ratio vector. [Figure 12] Figure 12 shows an example of connected vectors. [Figure 13] Figure 13 is a flowchart illustrating the operation of yet another example of a learning model generation device. [Figure 14] Figure 14 is a diagram showing the configuration of an example of a joint point detection device. [Figure 15] Figure 15 is a flowchart showing the operation of an example of a joint point detection device. [Figure 16] Figure 16 is a block diagram showing an example of a computer that implements a learning model generation device and an articulation point detection device. [Modes for carrying out the invention]

[0017] (Embodiment 1) The learning model generation apparatus, learning model generation method, and program in Embodiment 1 will be described below with reference to Figures 1 to 7.

[0018] [Device configuration] First, the schematic configuration of the learning model generation device in Embodiment 1 will be described using Figure 1. Figure 1 is a configuration diagram showing the schematic configuration of an example of a learning model generation device.

[0019] As shown in Figure 1, the learning model generation device 10 in Embodiment 1 is a device that generates a learning model for detecting human joint points from images. As shown in Figure 1, the learning model generation device 10 includes a data acquisition unit 11, a feature distance calculation unit 12, a similarity calculation unit 13, a loss calculation unit 14, and a learning model generation unit 15.

[0020] The data acquisition unit 11 acquires 2D joint point coordinate data from the training data, which includes 2D joint point coordinate data and 3D joint point coordinate data, and inputs the acquired 2D joint point coordinate data into the machine learning model.

[0021] Two-dimensional joint point coordinate data is data that can identify the two-dimensional coordinates of each of the multiple joint points of a person in an image. Three-dimensional joint point coordinate data is data that can identify the three-dimensional coordinates of each of the multiple joint points of a person.

[0022] The feature distance calculation unit 12 obtains the features calculated by the machine learning model for each training data, and then uses the obtained features to calculate the distance between the obtained features. The similarity calculation unit 13 obtains an index that shows the relationship between a person and the camera that took the image for each training data, using information about the camera that took the image for each image, and calculates the similarity between the obtained indices.

[0023] The loss calculation unit 14 uses the calculated similarity and the calculated distance between features to calculate the loss for the features in the machine learning model. The learning model generation unit 15 uses the calculated loss to update the parameters of the machine learning model.

[0024] Thus, in Embodiment 1, a loss is calculated that reflects the variability in the relationship between the person and the camera that formed the basis of the training data, and the parameters of the machine learning model are updated based on this loss. Therefore, using the machine learning model obtained in Embodiment 1, the problem of reduced joint point detection accuracy due to the height of the person in the image appearing taller or shorter depending on the camera's shooting angle is resolved. In other words, according to this embodiment, the detection accuracy when detecting the 3D coordinates of joint points from an image is improved.

[0025] Next, the configuration and functions of the learning model generation device 10 in Embodiment 1 will be specifically described using Figures 2 to 6. Figure 2 is a configuration diagram specifically showing the configuration of an example of a learning model generation device.

[0026] As shown in Figure 2, the learning model generation device 10 includes, in addition to the data acquisition unit 11, feature distance calculation unit 12, similarity calculation unit 13, loss calculation unit 14, and learning model generation unit 15 described above, a camera parameter acquisition unit 16 and a machine learning model 20. The learning model generation device 10 is also connected to the database 30 in a data communication manner.

[0027] In Embodiment 1, the machine learning model 20 is a neural network, specifically a Deep Neural Network (DNN). The machine learning model 20 has an input layer, a hidden layer (intermediate layer), and an output layer. The machine learning model 20 is actually implemented by a machine learning program that runs on a computer. Alternatively, the machine learning model 20 may be implemented on a device (computer) separate from the learning model generation device 10.

[0028] Database 30 stores 2D and 3D joint point coordinate data, which serve as training data for the machine learning model 20. The 3D joint point coordinate data is the training data. In Embodiment 1, database 30 also stores information about the camera that captured the images from which the 2D and 3D joint point coordinate data originated, specifically camera parameters.

[0029] Here, the 2D joint point coordinate data can be an image of a person and a set of 2D coordinates for each joint point on the image. The 2D joint point coordinate data may consist of either the image of a person or the set of 2D coordinates for each joint point on the image, or both. Alternatively, instead of a set of 2D coordinates for each joint point, a map representing the probability of each joint point existing, such as a heatmap, may be used.

[0030] The 3D joint point coordinate data serves as training data. Each 3D joint point coordinate data corresponds to one 2D joint point coordinate data. The 3D joint point coordinate data can be defined as the set of 3D coordinates of each joint point of a person in the corresponding 2D joint point coordinate data.

[0031] Camera parameters include external parameters and internal parameters. As will be described later, if the similarity calculation unit 13 uses only one of the external parameters or internal parameters, it is sufficient that only the camera parameters used are stored in the database 30. Alternatively, both external and internal parameters may be stored in the database 30.

[0032] External parameters are parameters that relate the world coordinate system and the camera coordinate system. External parameters include the camera angle relative to the world coordinate system and the camera position as viewed from the origin of the world coordinate system, and are represented by a 3x4 matrix as shown in Equation 1 below. Note that in Equation 1, (X C , Y C , Z C ) indicates the coordinate point in the camera coordinate system, and (X W , Y W , Z W ) indicates a coordinate point in the world coordinate system. Since the origin of the world coordinate system is set relative to the ground, the angle of the camera relative to the ground can be determined by external parameters.

[0033]

number

[0034] Intrinsic parameters are parameters that relate the camera coordinate system to the image coordinate system. These intrinsic parameters include the camera's focal length and lens distortion. Intrinsic parameters are represented by a 3x3 matrix, as shown in Equation 2 below. In Equation 2, (u, v) represent coordinate points on the image, and s is the skew ratio. According to intrinsic parameters, a single pixel on the image can be converted to its dimensions in the real world.

[0035]

number

[0036] The data acquisition unit 11 acquires each two-dimensional joint coordinate data prepared as training data from the database 30, and sequentially inputs each of the acquired two-dimensional joint coordinate data into the machine learning model 20.

[0037] Each time the two-dimensional joint coordinate data is input into the machine learning model 20 by the data acquisition unit 11, the feature quantity distance calculation unit 12 acquires the feature quantity calculated in the machine learning model 20, specifically, the output value of the intermediate layer of the machine learning model 20 (hereinafter also referred to as "intermediate feature quantity").

[0038] Then, the feature quantity distance calculation unit 12 sets combinations between two intermediate feature quantities so that each of the acquired intermediate feature quantities is in a pairwise combination, and calculates the feature quantity distance for each combination of two intermediate feature quantities. In other words, the combination of two intermediate feature quantities is, in fact, the combination of the persons who were the source of the training data input into the machine learning model 20. Therefore, the feature quantity distance calculation unit 12 calculates the feature quantity distance for each combination of the persons who were the source of the training data input into the machine learning model 20. Hereinafter, the "combination of the persons who were the source of the training data input into the machine learning model 20" will be simply referred to as the "combination of persons".

[0039] Specifically, the feature quantity distance calculation unit 12 calculates the difference between two intermediate feature quantities as the feature quantity distance. Here, for example, if the intermediate feature quantity of person A is fea A and the intermediate feature quantity of person B is fea B , then the feature quantity distance is expressed as "L2_norm(fea A - fea B )".

[0040] The camera parameter acquisition unit 16 acquires the camera parameters corresponding to the two-dimensional joint coordinate data acquired by the data acquisition unit 11 from the database 30. Also, the camera parameter acquisition unit 16 passes each of the acquired camera parameters to the similarity calculation unit 13.

[0041] In Embodiment 1, the similarity calculation unit 13 first calculates the received camera data for each 2D joint point coordinate data (for each training data) acquired by the data acquisition unit 11. La Using a meter, the camera pose vector is determined as an indicator showing the relationship between the person and the camera that captured the image.

[0042] For example, if the camera parameters are external parameters, the camera pose vector could include a vector containing the angle between the camera's optical axis and the vertical direction, and a vector containing the angle between the camera's optical axis and a part of the person. If the camera parameters are internal parameters, the camera pose vector could include a vector containing the components of the internal parameters.

[0043] Figures 3 to 5 show examples of camera pose vectors. Of these, Figure 3 shows the camera pose vector including the angle between the camera's optical axis and the vertical direction.

[0044] Specifically, in the example in Figure 3, X is used as the world coordinate system. W Axis, Y W axis, Z W The axis is set. Z W The axis is the vertical axis, X w Axis and Y w The axis is perpendicular to the vertical direction. Also, the camera coordinate system is X C Axis, Y C axis, Z C The axis is set. Z C The axis is the axis in the direction of the camera's optical axis. C The axis is the horizontal axis of the camera's image sensor, Y. C The axis is the vertical axis of the camera's image sensor.

[0045] In Example 1 shown in Figure 3, the components of the camera pose vector are "camera coordinate Z C Axes and the world coordinate Z W These are the "angle formed with" and the "height of the camera from the ground". C Axes and the world coordinate Z WThe "angle" is determined from the camera's external parameters. The "camera height from the ground" is a preset value.

[0046] Furthermore, in Example 2 shown in Figure 3, the components of the camera pose vector are "(camera coordinate Z C Axes and the world coordinate Z W The two components are "(angle between the camera and the ground) / 360 degrees" and "(camera height from the ground) / (average camera height from the ground)". In Example 2, the ratio of angles is used as a component of the camera pose vector. Note that "average camera height from the ground" is the average height of all cameras used to capture the images that formed the basis of the training data.

[0047] Figure 4 shows the camera pose vector, which includes the angle between the camera's optical axis and a part of the human body. In Example 3 shown in Figure 4, the components of the camera pose vector are the "angle between the bone of that part and the camera's optical axis" for each body part. In Example 4, the ratio obtained by dividing the angle by the reference angle is used as a component of the camera pose vector.

[0048] Specifically, the angle between the camera's optical axis and a body part can be determined from the optical axis vector (optical axis vector) and the vector indicating the direction of the bones in each body part (bone vector). The optical axis vector is the Z coordinate system of the camera coordinate system. C This is a vector along the axis and is [0,0,1]. Bone vectors are obtained from the difference in the 3D coordinates of two joint points in the camera coordinate system. For example, the bone vector from the right shoulder to the right elbow is obtained by "3D coordinate of the right shoulder - 3D coordinate of the right elbow".

[0049] Furthermore, while Figure 4 shows an example where bone vectors for each body part are used as camera pose vectors, another example is one where only the bone vector of a representative bone of the body, such as the spine, is used.

[0050] Furthermore, in Example 4, the reference angle used for division is 90 degrees, 180 degrees, or 360 degrees. The value of the reference angle is set according to how the angle between the camera's optical axis and a person's body part is expressed. For example, if the maximum value of the angle between the camera's optical axis and a person's body part is 90 degrees or less, the reference angle is set to 90 degrees.

[0051] Figure 5 shows the camera pose vector, including the components of the camera's intrinsic parameters. As shown in Figure 5, the camera's intrinsic parameters are represented by a 3x3 matrix, as shown in equation 2 above, and can be generalized to equation 3 below. In other words, the intrinsic parameters are a 11 ~a 33 It consists of nine elements. In the example in Figure 5, the camera pose vector is composed of these internal parameter components. A camera pose vector obtained from such internal parameters is useful when a person's image is used as two-dimensional joint point coordinate data.

[0052]

number

[0053] Next, the similarity calculation unit 13 calculates an average vector for the camera pose vector obtained for each training data. Here, the average vector is "cam mean This is written as "cam". Furthermore, the similarity calculation unit 13 calculates the camera vector (hereinafter referred to as "camera vector") of the camera that photographed each person. For example, the camera vector of person A is "cam A ", the camera vector of person B is "cam B When written as , the camera vector is calculated by the following equation 4.

[0054]

number

[0055] Next, the similarity calculation unit uses the camera vector for each person to perform a brute-force calculation, calculating the similarity for each combination of people that formed the basis of the training data input to the machine learning model 20, for example, cosine similarity (cos_sim(cam A cam B )) is calculated. Note that in Embodiment 1, the similarity is not limited to cosine similarity, and for example, Euclidean distance may be used.

[0056] In Embodiment 1, the loss calculation unit 14 calculates the cosine similarity (cos_sim(cam) calculated by the similarity calculation unit 13. A cam B )) and the feature distance (L2_norm( fea A - fea B Using ), the loss for features in machine learning model 20 m Calculate.

[0057] Specifically, the loss calculation unit 14 uses, for example, the following equation 5 to calculate the loss m The following equation 5 calculates the loss. In equation 5 below, i and j are indices that represent the people from whom the training data input to the machine learning model 20 originated. (i, j) represents the combination of people from whom the training data input to the machine learning model 20 originated. Note that (i, j) and (j, i) overlap, so in equation 5 below, only one of them is calculated. The loss calculation unit 14 uses a different formula to calculate the loss loss. m It is also possible to calculate this.

[0058]

number

[0059] The learning model generation unit 15 processes each loss calculated by the loss calculation unit 14. pThe parameters of the DNN, which is the machine learning model 20, are updated so that the value becomes smaller. As a result, as shown in Figure 6, the distance between features in the feature space of the DNN becomes the distance corresponding to the camera pose of each person. Figure 6 is a diagram illustrating an example of machine learning model generation. As a result of updating the parameters in this way, a machine learning model 20 is generated that can accurately estimate the 3D coordinates of joint points without being affected by the camera pose of each person that formed the basis of the training data.

[0060] [Device operation] Next, the operation of the learning model generation device 10 in Embodiment 1 will be explained using Figure 7. Figure 7 is a flowchart showing an example of the operation of the learning model generation device. In the following explanation, Figures 1 to 6 will be referred to as appropriate. In Embodiment 1, the learning model generation method is carried out by operating the learning model generation device 10. Therefore, the explanation of the learning model generation method in Embodiment 1 will be replaced by the following explanation of the operation of the learning model generation device 10.

[0061] As shown in Figure 7, first, the data acquisition unit 11 acquires 2D joint point coordinate data for each person, which is prepared as training data (step A1). Next, the data acquisition unit 11 inputs the 2D joint point coordinate data acquired in step A1 into the machine learning model 20 (step A2). Steps A1 and A2 may be performed for all of the prepared training data, or for only a set number of training data.

[0062] Next, in step A2, when the data acquisition unit 11 inputs the 2D joint point coordinate data to the machine learning model 20, the feature distance calculation unit 12 acquires the intermediate features calculated by the machine learning model 20 (step A3). Steps A2 and A3 are repeated for the number of 2D joint point coordinate data acquired in step A1.

[0063] Next, the feature distance calculation unit 12 calculates the feature distance for each combination of two intermediate features once intermediate features have been obtained for all the 2D joint point coordinate data acquired in step A1 (step A4).

[0064] Next, the camera parameter acquisition unit 16 acquires camera parameters from the database 30 that correspond to the 2D joint point coordinate data acquired in step A1, and passes each acquired camera parameter to the similarity calculation unit 13 (step A5).

[0065] Next, the similarity calculation unit 13 calculates the received camera part for each of the two-dimensional joint point coordinate data obtained in step A1. La Using a meter, we determine the camera pose vector, which shows the relationship between the person and the camera that captured the image (Step A6).

[0066] Next, the similarity calculation unit 13 uses the camera pose vector obtained in step A6 to determine a camera vector for each person who was the source of the training data input to the machine learning model 20, and further calculates the similarity between camera vectors for each combination of people (step A7).

[0067] Next, the loss calculation unit 14 applies the similarity calculated by the similarity calculation unit 13 in step A7 and the distance between features calculated by the feature distance calculation unit 12 in step A4 to the above equation 5, and calculates the loss for features in the machine learning model. m Calculate (Step A8).

[0068] Next, the learning model generation unit 15 uses the loss calculated in step A8. m Update the parameters of the machine learning model 20 so that the value becomes smaller (Step A9).

[0069] As described above, according to Embodiment 1, a loss is calculated that reflects the difference in the feature space of the person from which the training data originated and the variation in the camera's shooting angle, and the parameters of the machine learning model are updated based on this loss. As a result, a machine learning model 20 capable of accurately estimating the 3D coordinates of joint points is generated. According to Embodiment 1, the problem that "the accuracy of detecting 3D coordinates decreases due to the variation in the shooting angle of the camera that took the images that formed the basis of the training data" is resolved, and the detection accuracy when detecting the 3D coordinates of joint points from images is improved.

[0070] [program] The program in Embodiment 1 can be any program that causes a computer to execute steps A1 to A9 shown in Figure 7. By installing and running this program on a computer, the learning model generation device 10 and the learning model generation method in Embodiment 1 can be realized. In this case, the computer's processor functions as a data acquisition unit 11, a feature distance calculation unit 12, a similarity calculation unit 13, a loss calculation unit 14, a learning model generation unit 15, and a camera parameter acquisition unit 16, and performs processing. Examples of computers include general-purpose PCs, smartphones, and tablet devices. The computer's processor also constructs the machine learning model 20.

[0071] Furthermore, the program in Embodiment 1 may be executed by a computer system constructed by multiple computers. In this case, for example, each computer may function as one of the following: a data acquisition unit 11, a feature distance calculation unit 12, a similarity calculation unit 13, a loss calculation unit 14, a learning model generation unit 15, and a camera parameter acquisition unit 16.

[0072] (Embodiment 2) Next, the learning model generation apparatus, learning model generation method, and program in Embodiment 2 will be described with reference to Figures 8 and 9.

[0073] [Device configuration] First, the configuration of the learning model generation device in Embodiment 2 will be explained using Figure 8. Figure 8 is a configuration diagram showing the configuration of another example of the learning model generation device.

[0074] The learning model generation device 40 in Embodiment 2, shown in Figure 8, is a device for generating a machine learning model 20, similar to the learning model generation device 10 shown in Figure 2 in Embodiment 1.

[0075] Furthermore, as shown in Figure 8, the learning model generation device 40, like the learning model generation device 10, includes a data acquisition unit 11, a feature distance calculation unit 12, a similarity calculation unit 13, a loss calculation unit 14, a learning model generation unit 15, a camera parameter acquisition unit 16, and a machine learning model 20.

[0076] However, as shown in Figure 8, unlike the learning model generation device 10, the learning model generation device 40 includes, in addition to the above, a loss integration unit 41 and a second loss calculation unit 42. The differences from Embodiment 1 will be explained below.

[0077] The second loss calculation unit 42 acquires the 3D joint point coordinate data output by the machine learning model 20 in response to the input of 2D joint point coordinate data (training data) from the data acquisition unit 11. Then, the second loss calculation unit 42 uses the 3D joint point coordinate data output by the machine learning model 20 and the 3D joint point coordinate data stored in the database 30 to calculate the loss for the output of the machine learning model 20 for each person from whom the training data was input to the machine learning model 20. After that, the second loss calculation unit 42 sums up the losses for each person and calculates the loss loss. p Calculate.

[0078] loss loss p The calculation process is represented by the following number 6. Below, 3D_data m This is the 3D joint point coordinate data output by the machine learning model 20, and is 3D_data tis the 3D joint point coordinate data, which is the training data. i is an index indicating the person from whom the training data was input to the machine learning model 20.

[0079]

number

[0080] The loss integration unit 41 calculates the loss for the feature calculated by the loss calculation unit 14. m and the loss for the output calculated by the second loss calculation unit 42 p The and are integrated. Specifically, the loss integration unit 41 uses the following equation 7 to calculate the loss for the feature quantities. m and loss about the output p The two values ​​are combined by calculating a weighted average of the two to obtain the final loss. In equation 7 below, λ represents the weighting coefficient in the weighted average. The value of the weighting coefficient λ is set as appropriate.

[0081]

number

[0082] In Embodiment 2, the learning model generation unit 15 uses the loss obtained through integration to update the parameters of the DNN, which is the machine learning model 20, so that the loss is reduced.

[0083] [Device operation] Next, the operation of the learning model generation device 40 in Embodiment 2 will be explained with reference to Figure 9. Figure 9 is a flowchart showing the operation of another example of the learning model generation device. In the following explanation, Figure 8 will be referred to as appropriate. In Embodiment 2, the learning model generation method is carried out by operating the learning model generation device 40. Therefore, the explanation of the learning model generation method in Embodiment 2 will be replaced by the following explanation of the operation of the learning model generation device 40.

[0084] As shown in Figure 9, first, the data acquisition unit 11 acquires 2D joint point coordinate data for each person, which is prepared as training data (step B1). Next, the data acquisition unit 11 inputs the 2D joint point coordinate data acquired in step B1 into the machine learning model 20 (step B2).

[0085] Next, in step B2, when the data acquisition unit 11 inputs the 2D joint point coordinate data to the machine learning model 20, the feature distance calculation unit 12 acquires the intermediate features calculated by the machine learning model 20 (step B3).

[0086] Next, the feature distance calculation unit 12 calculates the feature distance for each combination of two intermediate features once intermediate features have been obtained for all the 2D joint point coordinate data acquired in step B1 (step B4).

[0087] Next, the camera parameter acquisition unit 16 acquires camera parameters from the database 30 that correspond to the 2D joint point coordinate data acquired in step B1, and passes each acquired camera parameter to the similarity calculation unit 13 (step B5).

[0088] Next, the similarity calculation unit 13 calculates the received camera part for each of the two-dimensional joint point coordinate data obtained in step B1. La Using a meter, we determine the camera pose vector that shows the relationship between the person and the camera that captured the image (Step B6).

[0089] Next, the similarity calculation unit 13 uses the camera pose vector obtained in step B6 to determine a camera vector for each person who was the source of the training data input to the machine learning model 20, and further calculates the similarity between camera vectors for each combination of people (step B7).

[0090] Next, the loss calculation unit 14 applies the similarity calculated by the similarity calculation unit 13 in step A7 and the distance between features calculated by the feature distance calculation unit 12 in step A4 to the above equation 2, and calculates the loss for features in the machine learning model.m Calculate (Step B8). Note that in Embodiment 2, each of Steps B1 to B8 is the same as Steps A1 to A8 in Embodiment 1.

[0091] Next, the second loss calculation unit 42 acquires the 3D joint point coordinate data output by the output layer of the machine learning model 20 in response to the input in step B3. Furthermore, the second loss calculation unit 42 applies the acquired 3D joint point coordinate data and the training data (3D joint point coordinate data) to the above equation 6 to calculate the loss for the output of the machine learning model 20 for each person. Then, the second loss calculation unit 42 sums up the losses for each person and calculates the loss loss p Calculate (Step B9).

[0092] Next, the loss integration unit 41 uses the above equation 7 to calculate the loss for the feature calculated in step B8. p and the loss for the output calculated in step B9 p Combine these to calculate the final loss (Step B10).

[0093] Subsequently, the learning model generation unit 15 updates the parameters of the machine learning model 20 so that the loss calculated in step B9 becomes smaller (step B11).

[0094] As described above, in Embodiment 2, as in Embodiment 1, a loss is calculated that reflects the difference in the feature space of the person that formed the basis of the training data and the variation in the camera's shooting angle. In Embodiment 2, a loss due to the output of the machine learning model 20 is also calculated. In other words, in Embodiment 2, the parameters of the machine learning model are updated based on the above two losses. Therefore, according to Embodiment 2, the detection accuracy is further improved.

[0095] [program] The program in Embodiment 2 can be any program that causes a computer to execute steps B1 to B11 shown in Figure 9. By installing and running this program on a computer, the learning model generation device 10 and the learning model generation method in Embodiment 2 can be realized. In this case, the computer's processor functions as a data acquisition unit 11, a feature distance calculation unit 12, a similarity calculation unit 13, a loss calculation unit 14, a learning model generation unit 15, a camera parameter acquisition unit 16, a loss integration unit 41, and a second loss calculation unit 42, and performs processing. Examples of computers include general-purpose PCs, smartphones, and tablet devices. The computer's processor also constructs the machine learning model 20.

[0096] Furthermore, the program in Embodiment 2 may be executed by a computer system constructed by multiple computers. In this case, for example, each computer may function as one of the following: data acquisition unit 11, feature distance calculation unit 12, similarity calculation unit 13, loss calculation unit 14, learning model generation unit 15, camera parameter acquisition unit 16, loss integration unit 41, and second loss calculation unit 42.

[0097] (Embodiment 3) Next, the learning model generation apparatus, learning model generation method, and program in Embodiment 3 will be described with reference to Figure 10.

[0098] [Device configuration] First, the configuration of the learning model generation device in Embodiment 3 will be explained using Figure 10. Figure 10 is a configuration diagram showing yet another example of the configuration of the learning model generation device.

[0099] The learning model generation device 50 in Embodiment 3, shown in Figure 10, is a device for generating a machine learning model 20, similar to the learning model generation device 10 shown in Figure 2 in Embodiment 1.

[0100] By the way, in order to improve the accuracy of detecting 3D joint point coordinates using a learning model, it is necessary to prepare a large amount of training data. However, machine learning in a learning model proceeds in a way that outputs values ​​for the average body size, even if the training data is obtained from people with diverse body sizes. As a result, a problem arises in which the detection accuracy decreases when the height of the person being targeted for posture estimation is higher or lower than the height of the person from whom the training data was obtained. Embodiment 3 solves this problem with the following configuration.

[0101] As shown in Figure 10, the learning model generation device 50, like the learning model generation device 10, includes a data acquisition unit 11, a feature distance calculation unit 12, a similarity calculation unit 13, a loss calculation unit 14, a learning model generation unit 15, a camera parameter acquisition unit 16, and a machine learning model 20.

[0102] However, as shown in Figure 10, unlike the learning model generation device 10, the learning model generation device 50 is equipped with a correct answer data acquisition unit 51 in addition to the above. Furthermore, Embodiment 3 differs from Embodiment 1 in terms of the functionality of the similarity calculation unit. The differences from Embodiment 1 will be explained below.

[0103] First, in Embodiment 3, the ground truth data acquisition unit 51 acquires 3D joint point coordinate data, which is both training data and teacher data, from the database 30. The ground truth data acquisition unit 51 then passes each acquired 3D joint point coordinate data to the similarity calculation unit 13.

[0104] In Embodiment 3, the similarity calculation unit 13 first uses the camera vectors shown in Embodiments 1 and 2, as well as the 3D joint point coordinate data, to obtain a vector representing the physique of the person that formed the basis of the training data input to the machine learning model 20 (hereinafter referred to as the "physique vector").

[0105] Specifically, the similarity calculation unit 13 calculates the bone length vector for each person using 3D joint point coordinate data, and further calculates the "ratio vector of bone lengths" from the calculated bone length vector.

[0106] Figure 11 shows an example of a bone length vector and a bone length ratio vector. As shown in Figure 11, the bone length vector is composed of lengths such as "length from right shoulder to right elbow," "length from right elbow to right wrist," "length from right hip to right ankle," and "length from left hip to left ankle." Also, as shown in Figure 11, each length is calculated from the difference in coordinate values ​​(3D coordinates) between joint points. The bone length ratio vector is calculated by dividing each length that makes up the bone length vector by a reference length.

[0107] Furthermore, the similarity calculation unit 13 calculates an average vector for the bone length ratio vectors of all individuals included in the training data. Here, the average vector is "phy mean It is written as "phy". Furthermore, the similarity calculation unit 13 uses the average vector to calculate a body size vector representing the body size of each person. For example, the body size vector of person A is "phy A ", the body size vector of person B is "phy B When expressed as '', the body size vector is calculated by the following number 8.

[0108]

number

[0109] Next, the similarity calculation unit concatenates the camera vector and body size vector for each person who was the source of the training data input to the machine learning model 20, as shown in Figure 12. Hereafter, the vector obtained by concatenation will be referred to as the concatenated vector. Figure 12 shows an example of a concatenated vector. In the example in Figure 12, the camera pose vector shown in Figure 3 and the body size vector shown in Figure 11 are concatenated.

[0110] Furthermore, the similarity calculation unit uses a concatenated vector to calculate the similarity for each combination of people that formed the basis of the training data input to the machine learning model 20, for example, the cosine similarity (cos_sim(cam) i +phy i cam j +phy j )) is calculated. Note that in Embodiment 3 as well, the similarity is not limited to cosine similarity, and for example, Euclidean distance may be used.

[0111] [Device operation] Next, the operation of the learning model generation device 10 in Embodiment 3 will be explained using Figure 13. Figure 13 is a flowchart showing yet another example of the operation of the learning model generation device. In the following explanation, Figures 10 to 12 will be referred to as appropriate. In Embodiment 3, the learning model generation method is carried out by operating the learning model generation device 50. Therefore, the explanation of the learning model generation method in Embodiment 3 will be replaced by the following explanation of the operation of the learning model generation device 50.

[0112] As shown in Figure 13, first, the data acquisition unit 11 acquires 2D joint point coordinate data for each person, which is prepared as training data (step C1). Next, the data acquisition unit 11 inputs the 2D joint point coordinate data acquired in step C1 into the machine learning model 20 (step C2).

[0113] Next, in step C2, when the data acquisition unit 11 inputs the 2D joint point coordinate data to the machine learning model 20, the feature distance calculation unit 12 acquires the intermediate features calculated by the machine learning model 20 (step C3).

[0114] Next, the feature distance calculation unit 12 calculates the feature distance for each combination of two intermediate features once intermediate features have been obtained for all the 2D joint point coordinate data acquired in step A1 (step C4).

[0115] Next, the camera parameter acquisition unit 16 acquires camera parameters from the database 30 that correspond to the 2D joint point coordinate data acquired in step C1, and passes each acquired camera parameter to the similarity calculation unit 13 (step C5). Steps C1 to C5 are executed in the same way as steps A1 to A5 in Embodiment 1 shown in Figure 7.

[0116] Next, the correct answer data acquisition unit 51 acquires 3D joint point coordinate data, which is both training data and teacher data, from the database 30, and passes each acquired 3D joint point coordinate data to the similarity calculation unit 13 (step C6).

[0117] Next, the similarity calculation unit 13 uses the camera parameters obtained in step C5 and the 3D joint point coordinate data obtained in step C6 to determine a camera pose vector and a body size vector for each of the 2D joint point coordinate data obtained in step C1. Furthermore, the similarity calculation unit 13 concatenates the camera pose vector and the body size vector for each of the 2D joint point coordinate data obtained in step A1 (step C7).

[0118] Next, the similarity calculation unit 13 calculates the similarity between the linked vectors obtained in step C7 for each combination of people (step C8).

[0119] Next, the loss calculation unit 14 applies the similarity calculated by the similarity calculation unit 13 in step C8 and the distance between features calculated by the feature distance calculation unit 12 in step C4 to the above equation 5, and calculates the loss for features in the machine learning model. m Calculate (Step C9).

[0120] Next, the learning model generation unit 15 uses the loss calculated in step C9. m Update the parameters of the machine learning model 20 so that the value becomes smaller (step C10).

[0121] As described above, according to Embodiment 3, in addition to the loss that reflects the difference in the feature space of the person from which the training data was derived and the variation in the camera's shooting angle, a loss that reflects the difference in the feature space of the person from which the training data was derived and the variation in body size is also calculated. Therefore, in Embodiment 3, when the height of the person to be estimated is higher or lower than the height of the person from whom the training data was obtained, a decrease in detection accuracy is suppressed. According to Embodiment 3, the detection accuracy when detecting the 3D coordinates of joint points from an image is further improved.

[0122] Furthermore, the learning model generation device 50 in Embodiment 3 may have the same configuration and functions as the learning model generation device 40 shown in Embodiment 2. In other words, the learning model generation device 50 may include the loss integration unit 41 and the second loss calculation unit 42 shown in Embodiment 2.

[0123] [program] The program in Embodiment 3 can be any program that causes a computer to execute steps C1 to C10 shown in Figure 13. By installing and running this program on a computer, the learning model generation device 10 and learning model generation method in Embodiment 1 can be realized. In this case, the computer's processor functions as a data acquisition unit 11, a feature distance calculation unit 12, a similarity calculation unit 13, a loss calculation unit 14, a learning model generation unit 15, a camera parameter acquisition unit 16, and a ground truth data acquisition unit 51, and performs processing. Examples of computers include general-purpose PCs, smartphones, and tablet devices. The computer's processor also constructs the machine learning model 20.

[0124] Furthermore, the program in Embodiment 3 may be executed by a computer system constructed by multiple computers. In this case, for example, each computer may function as one of the following: data acquisition unit 11, feature distance calculation unit 12, similarity calculation unit 13, loss calculation unit 14, learning model generation unit 15, camera parameter acquisition unit 16, and ground truth data acquisition unit 51.

[0125] (Embodiment 4) Next, in Embodiment 4, the joint point detection device, joint point detection method, and program will be described with reference to the drawings.

[0126] [Device configuration] First, the configuration of the joint point detection device in Embodiment 4 will be explained using Figure 14. Figure 14 is a configuration diagram showing an example of the joint point detection device.

[0127] As shown in Figure 14, the joint point detection device 60 comprises a data acquisition unit 61 and a joint point detection unit 62. The joint point detection device 60 also includes a machine learning model 20.

[0128] The data acquisition unit 61 acquires two-dimensional joint point coordinate data that can identify the two-dimensional coordinates of each of the multiple joint points of a person in the image. The two-dimensional joint point coordinate data acquired by the data acquisition unit 61 is the same as the two-dimensional joint point coordinate data described in Embodiment 1, and is two-dimensional joint point coordinate data of a person for which the detection of the three-dimensional coordinates of each joint point is required. The two-dimensional joint point coordinate data is input from an external device or the like.

[0129] The joint point detection unit 62 applies the 2D joint point coordinate data acquired by the data acquisition unit 61 to the machine learning model 20 to detect the 3D coordinates of each of a person's multiple joint points.

[0130] The machine learning model 20 is a machine model that learns the relationship between the 2D and 3D coordinates of human joint points. In Embodiment 4, the machine learning model 20 is a machine learning model created according to Embodiments 1 to 3.

[0131] In other words, in Embodiment 4, the machine learning model 20 is created by machine learning using 2D joint point coordinate data and 3D joint point coordinate data, which serve as training data. The parameters of the machine learning model 20 are updated, as in Embodiment 1, using the distance between features obtained from the features calculated by the machine learning model 20, an index indicating the relationship between a person and the camera that took the image, specifically the camera parameters, and the loss on the features obtained from these.

[0132] Furthermore, the parameters of the machine learning model 20 may be updated, similar to Embodiment 2, not only using the loss on features, but also using the loss obtained from the 3D joint point coordinate data output by the machine learning model 20 and the 3D joint point coordinate data used for training.

[0133] In Embodiment 4, the machine learning model 20 is implemented by a machine learning program executed on a computer. Alternatively, the machine learning model 20 may be implemented on a separate device (computer) from the joint point detection device 60.

[0134] [Device operation] Next, the operation of the joint point detection device 60 in Embodiment 4 will be explained using Figure 15. Figure 15 is a flowchart showing the operation of an example of the joint point detection device. In the following explanation, Figure 14 will be referred to as appropriate. In Embodiment 4, the joint point detection method is performed by operating the joint point detection device 60. Therefore, the explanation of the joint point detection method in Embodiment 4 will be replaced by the following explanation of the operation of the joint point detection device 60.

[0135] As shown in Figure 15, first, the data acquisition unit 61 acquires two-dimensional joint point coordinate data for the person whose joint points are to be detected (step D1).

[0136] Next, the joint point detection unit 62 applies the 2D joint point coordinate data acquired by the data acquisition unit 81 in step D1 to the machine learning model 20 to detect the 3D coordinates of each joint point of the person to be detected (step D2).

[0137] Specifically, the joint point detection unit 62 inputs the 2D joint point coordinate data acquired by the data acquisition unit 61 in step D1 into the machine learning model 20. As a result, the machine learning model 20 outputs 3D joint point coordinate data, and the joint point detection unit 62 acquires the outputted 3D joint point coordinate data.

[0138] Thus, according to Embodiment 4, the machine learning model 20 can be used to detect the three-dimensional coordinates of each joint point of a person.

[0139] [program] The program in Embodiment 4 can be any program that causes a computer to execute steps D1 to D2 shown in Figure 15. By installing and running this program on a computer, the joint point detection device 60 and joint point detection method in Embodiment 4 can be realized. In this case, the computer's processor functions as a data acquisition unit 61 and a joint point detection unit 62, and performs processing. Examples of computers include general-purpose PCs, smartphones, and tablet devices. The computer's processor also constructs a machine learning model 20.

[0140] Furthermore, the program in Embodiment 4 may be executed by a computer system constructed by multiple computers. In this case, for example, each computer may function as either the data acquisition unit 61 or the joint point detection unit 62.

[0141] [Physical configuration] Here, a computer that implements a learning model generation device and a joint point detection device by executing the programs in Embodiments 1 to 4 will be described using Figure 16. Figure 16 is a block diagram showing an example of a computer that implements a learning model generation device and a joint point detection device.

[0142] As shown in Figure 16, the computer 110 comprises a CPU (Central Processing Unit) 111, main memory 112, storage device 113, input interface 114, display controller 115, data reader / writer 116, and communication interface 117. Each of these components is connected to the others via a bus 121, enabling data communication.

[0143] Furthermore, the computer 110 may include a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array) in addition to, or instead of, the CPU 111. In this embodiment, the GPU or FPGA can execute the program in the embodiment.

[0144] The CPU 111 loads the program in the embodiment, which consists of a set of codes stored in the storage device 113, into the main memory 112, and performs various calculations by executing each code in a predetermined order. The main memory 112 is typically a volatile storage device such as DRAM (Dynamic Random Access Memory).

[0145] Furthermore, the program in this embodiment is provided stored on a computer-readable recording medium 120. The program in this embodiment may also be distributed over the internet via a communication interface 117.

[0146] Specific examples of the storage device 113 include hard disk drives and semiconductor storage devices such as flash memory. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and mouse. The display controller 115 is connected to the display device 119 and controls the display on the display device 119.

[0147] The data reader / writer 116 mediates data transmission between the CPU 111 and the recording medium 120, reads programs from the recording medium 120, and writes processing results from the computer 110 to the recording medium 120. The communication interface 117 mediates data transmission between the CPU 111 and other computers.

[0148] Furthermore, specific examples of the recording medium 120 include general-purpose semiconductor memory devices such as CF (Compact Flash®) and SD (Secure Digital), magnetic recording media such as Flexible Disks, or optical recording media such as CD-ROMs (Compact Disk Read Only Memory).

[0149] Furthermore, the learning model generation device and joint point detection device in the embodiment can be implemented not by a computer with a program installed, but by using hardware corresponding to each part, such as electronic circuits. Moreover, the learning model generation device and joint point detection device in the embodiment may be partially implemented by a program and the remaining part by hardware. In the embodiment, the computer is not limited to the computer shown in Figure 16.

[0150] Some or all of the embodiments described above can be expressed by (Appendix 1) to (Appendix 27) described below, but are not limited to the following descriptions.

[0151] (Note 1) A data acquisition unit acquires the 2D joint point coordinate data from training data that includes 2D joint point coordinate data capable of identifying the 2D coordinates of each of several joint points of a person in an image, and 3D joint point coordinate data capable of identifying the 3D coordinates of each of the said multiple joint points of the person, and inputs the acquired 2D joint point coordinate data into a machine learning model. A feature distance calculation unit obtains the features calculated in the machine learning model for each of the training data, and further calculates the distance between the obtained features using the obtained features. A similarity calculation unit calculates an index indicating the relationship between the person and the camera that took the image, using information about the camera that took the image for each of the training data, and calculates the similarity between the calculated indices. A loss calculation unit calculates the loss for the features in the machine learning model using the calculated similarity and the calculated distance between the features. A learning model generation unit updates the parameters of the machine learning model using the calculated loss, A learning model generation device characterized by having the following features.

[0152] (Note 2) The similarity calculation unit, for each of the training data, uses the camera's external parameters as information about the camera and calculates a camera pose vector as an index, which includes the angle between the camera's optical axis direction and the vertical direction. The learning model generation device described in Appendix 1.

[0153] (Note 3) The similarity calculation unit, for each of the training data, uses the camera's external parameters as information about the camera and calculates a camera pose vector as an index, which includes the angle between the camera's optical axis and the person's body part. The learning model generation device described in Appendix 1.

[0154] (Note 4) The similarity calculation unit, for each of the training data, uses the camera's intrinsic parameters as information about the camera and calculates a camera pose vector as an index, which includes the components of the intrinsic parameters. The learning model generation device described in Appendix 1.

[0155] (Note 5) The similarity calculation unit further calculates the similarity between the body types of the people that formed the basis of the training data, using the three-dimensional joint point coordinate data which is the training data. The loss calculation unit calculates the loss for the features in the machine learning model, also using the similarity between the body types of the people that formed the basis of the training data. The learning model generation device described in Appendix 1.

[0156] (Note 6) A second loss calculation unit calculates the loss for the output of the machine learning model using the 3D joint point coordinate data output by the machine learning model and the 3D joint point coordinate data which is the training data, in response to the input of the 2D joint point coordinate data by the data acquisition unit. A loss integration unit that integrates the loss for the feature quantity and the loss for the output, Equipped with, The learning model generation unit updates the parameters of the machine learning model using the loss obtained through integration. The learning model generation device described in Appendix 1.

[0157] (Note 7) The loss integration unit integrates the loss for the feature and the loss for the output by calculating a weighted average of the two. The learning model generation device described in Appendix 6.

[0158] (Note 8) The aforementioned machine learning model is a neural network, The feature distance calculation unit obtains the feature from the intermediate layer of the neural network. The learning model generation device described in Appendix 1.

[0159] (Note 9) A data acquisition unit obtains 2D joint point coordinate data that can identify the 2D coordinates of each of the multiple joint points of a person in an image. The joint point detection unit applies the acquired 2D joint point coordinate data to a machine learning model that has learned the relationship between the 2D and 3D coordinates of a person's joint points, and detects the 3D coordinates of each of the multiple joint points of the person. Equipped with, The parameters of the aforementioned machine learning model are: In machine learning using 2D joint point coordinate data and 3D joint point coordinate data that can identify the 3D coordinates of multiple joint points in a person, The data is updated using the distance between features obtained from the features calculated by the machine learning model, an index showing the relationship between a person and the camera that took the image, and the loss value for the features in the machine learning model, which is derived from these. A joint point detection device characterized by the following features.

[0160] (Note 10) A data acquisition step involves acquiring the 2D joint point coordinate data from training data that includes 2D joint point coordinate data capable of identifying the 2D coordinates of each of several joint points of a person in an image, and 3D joint point coordinate data capable of identifying the 3D coordinates of each of the said multiple joint points of the person, and inputting the acquired 2D joint point coordinate data into a machine learning model. A feature distance calculation step is performed, in which, for each of the training data, the features calculated in the machine learning model are obtained, and further, the distance between the obtained features is calculated using the obtained features. A similarity calculation step is performed for each of the training data, using information about the camera that took each image, to determine an index indicating the relationship between the person and the camera that took the image, and to calculate the similarity between the determined indices. A loss calculation step, which involves calculating the loss for the features in the machine learning model using the calculated similarity and the calculated distance between the features, A learning model generation step in which the parameters of the machine learning model are updated using the calculated loss, A method for generating a learning model, characterized by having [a certain feature].

[0161] (Note 11) In the similarity calculation step, for each training data, the camera's external parameters are used as information about the camera, and a camera pose vector is determined, which includes the angle between the optical axis direction and the vertical direction of the camera, as the index. The learning model generation method described in Appendix 10.

[0162] (Note 12) In the similarity calculation step, for each training data, the camera's external parameters are used as information about the camera, and a camera pose vector is determined, which includes the angle between the camera's optical axis and the person's body part as the index. The learning model generation method described in Appendix 10.

[0163] (Note 13) In the similarity calculation step, for each training data, the camera's intrinsic parameters are used as information about the camera, and a camera pose vector is obtained as the index, which includes the components of the intrinsic parameters. The learning model generation method described in Appendix 10.

[0164] (Note 14) In the similarity calculation step, the similarity between the body types of the people who formed the basis of the training data is further calculated using the three-dimensional joint point coordinate data which is the training data. In the loss calculation step, the loss for the features in the machine learning model is calculated using the similarity between the physical characteristics of the people that formed the basis of the training data. The learning model generation method described in Appendix 10.

[0165] (Note 15) A second loss calculation step is performed to calculate the loss for the output of the machine learning model using the 3D joint point coordinate data output by the machine learning model in response to the input of the 2D joint point coordinate data obtained in the data acquisition step, and the 3D joint point coordinate data which is the training data. A loss integration step that integrates the loss for the aforementioned feature and the loss for the aforementioned output, It further possesses, In the learning model generation step, the parameters of the machine learning model are updated using the loss obtained by integration. The learning model generation method described in Appendix 10.

[0166] (Note 16) In the loss consolidation step, the two are consolidated by calculating a weighted average of the loss for the feature and the loss for the output. The learning model generation method described in Appendix 15.

[0167] (Note 17) The aforementioned machine learning model is a neural network, In the step of calculating the distance between features, the features are obtained from the intermediate layer of the neural network. The learning model generation method described in Appendix 10.

[0168] (Note 18) A data acquisition step to obtain 2D joint point coordinate data that allows the identification of the 2D coordinates of each of the multiple joint points of a person in an image, The joint point detection step involves applying the acquired 2D joint point coordinate data to a machine learning model that has learned the relationship between the 2D and 3D coordinates of a person's joint points, thereby detecting the 3D coordinates of each of the multiple joint points of the person. Equipped with, The parameters of the aforementioned machine learning model are: In machine learning using 2D joint point coordinate data and 3D joint point coordinate data that can identify the 3D coordinates of multiple joint points in a person, The data is updated using the distance between features obtained from the features calculated by the machine learning model, an index showing the relationship between a person and the camera that took the image, and the loss value for the features in the machine learning model, which is derived from these. A method for detecting joint points, characterized by the features described above.

[0169] (Note 19) On the computer, A data acquisition step involves acquiring the 2D joint point coordinate data from training data that includes 2D joint point coordinate data capable of identifying the 2D coordinates of each of several joint points of a person in an image, and 3D joint point coordinate data capable of identifying the 3D coordinates of each of the said multiple joint points of the person, and inputting the acquired 2D joint point coordinate data into a machine learning model. A feature distance calculation step is performed, in which, for each of the training data, the features calculated in the machine learning model are obtained, and further, the distance between the obtained features is calculated using the obtained features. A similarity calculation step is performed for each of the training data, using information about the camera that took each image, to determine an index indicating the relationship between the person and the camera that took the image, and to calculate the similarity between the determined indices. A loss calculation step, which involves calculating the loss for the features in the machine learning model using the calculated similarity and the calculated distance between the features, A learning model generation step in which the parameters of the machine learning model are updated using the calculated loss, Let's execute it ru, Professional Hmm.

[0170] (Note 20) In the similarity calculation step, for each training data, the camera's external parameters are used as information about the camera, and a camera pose vector is determined, which includes the angle between the optical axis direction and the vertical direction of the camera, as the index. As described in Appendix 19 program .

[0171] (Note 21) In the similarity calculation step, for each training data, the camera's external parameters are used as information about the camera, and a camera pose vector is determined, which includes the angle between the camera's optical axis and the person's body part as the index. As described in Appendix 19 program .

[0172] (Note 22) In the similarity calculation step, for each training data, the camera's intrinsic parameters are used as information about the camera, and a camera pose vector is obtained as the index, which includes the components of the intrinsic parameters. As described in Appendix 19 program .

[0173] (Note 23) In the similarity calculation step, the similarity between the body types of the people who formed the basis of the training data is further calculated using the three-dimensional joint point coordinate data which is the training data. In the loss calculation step, the loss for the features in the machine learning model is calculated using the similarity between the physical characteristics of the people that formed the basis of the training data. As described in Appendix 19 program .

[0174] (Note 24) before On the computer, A second loss calculation step is performed to calculate the loss for the output of the machine learning model using the 3D joint point coordinate data output by the machine learning model in response to the input of the 2D joint point coordinate data obtained in the data acquisition step, and the 3D joint point coordinate data which is the training data. A loss integration step that integrates the loss for the aforementioned feature and the loss for the aforementioned output, It further includes instructions to execute, In the learning model generation step, the parameters of the machine learning model are updated using the loss obtained by integration. As described in Appendix 19 program .

[0175] (Note 25) In the loss consolidation step, the two are consolidated by calculating a weighted average of the loss for the feature and the loss for the output. As described in Appendix 24 program .

[0176] (Note 26) The aforementioned machine learning model is a neural network, In the step of calculating the distance between features, the features are obtained from the intermediate layer of the neural network. As described in Appendix 19 program .

[0177] (Note 27) On the computer, A data acquisition step to obtain 2D joint point coordinate data that allows the identification of the 2D coordinates of each of the multiple joint points of a person in an image, The joint point detection step involves applying the acquired 2D joint point coordinate data to a machine learning model that has learned the relationship between the 2D and 3D coordinates of a person's joint points, thereby detecting the 3D coordinates of each of the multiple joint points of the person. Execute height, The parameters of the aforementioned machine learning model are: In machine learning using 2D joint point coordinate data and 3D joint point coordinate data that can identify the 3D coordinates of multiple joint points in a person, The data is updated using the distance between features obtained from the features calculated by the machine learning model, an index showing the relationship between a person and the camera that took the image, and the loss value for the features in the machine learning model, which is derived from these. program .

[0178] Although the present invention has been described above with reference to embodiments, the present invention is not limited to the above embodiments. Various modifications to the structure and details of the present invention can be made, as can be understood by those skilled in the art within the scope of the present invention.

[0179] This application claims priority based on Japanese Patent Application No. 2023-019012, filed on 10 February 2023, and incorporates all of its disclosures herein. [Industrial applicability]

[0180] As described above, this disclosure makes it possible to improve the detection accuracy when detecting the three-dimensional coordinates of joint points from an image. The present invention is useful for various systems that estimate a person's posture from an image. [Explanation of Symbols]

[0181] 10. Learning Model Generation Device (Embodiment 1) 11 Data Acquisition Unit 12 Feature Distance Calculation Unit 13 Similarity calculation unit 14 Loss calculation section 15. Learning Model Generation Unit 16 Camera parameter acquisition unit 20 Machine Learning Models 30 databases 40 Learning Model Generation Device (Embodiment 2) 41 Loss integration section 42 Second Loss Calculation Unit 50 Learning Model Generation Device (Embodiment 3) 51 Correct answer data acquisition Department 60 Joint inspection device 61 Data acquisition unit 62 Joint detection unit 110 Computer 111 CPU 112 Main memory 113 Storage device 114 Input interface 115 Display controller 116 Data reader / writer 117 Communication interface 118 Input device 119 Display device 120 Recording medium 121 Bus

Claims

1. A data acquisition unit acquires the two-dimensional joint point coordinate data from training data that includes two-dimensional joint point coordinate data capable of identifying the two-dimensional coordinates of multiple joint points of a person in an image, and three-dimensional joint point coordinate data capable of identifying the three-dimensional coordinates of the multiple joint points of the person, and inputs the acquired two-dimensional joint point coordinate data into a machine learning model. A feature distance calculation unit obtains the feature quantities calculated in the machine learning model for each of the training data, and further calculates the distance between the obtained feature quantities using the obtained feature quantities. A similarity calculation unit calculates an index indicating the relationship between the person and the camera that took the image, using information about the camera that took the image for each of the training data, and calculates the similarity between the calculated indices. A loss calculation unit calculates the loss for the features in the machine learning model using the calculated similarity and the calculated distance between the features. A learning model generation unit updates the parameters of the machine learning model using the calculated loss, A learning model generation device characterized by having the following features.

2. The similarity calculation unit, for each of the training data, uses the camera's external parameters as information about the camera and calculates a camera pose vector as an index, which includes the angle between the camera's optical axis direction and the vertical direction. A learning model generation device according to claim 1.

3. The similarity calculation unit, for each of the training data, uses the camera's external parameters as information about the camera and calculates a camera pose vector as an index, which includes the angle between the camera's optical axis and the person's body part. A learning model generation device according to claim 1.

4. The similarity calculation unit, for each of the training data, uses the camera's intrinsic parameters as information about the camera and calculates a camera pose vector as an index, which includes the components of the intrinsic parameters. A learning model generation device according to claim 1.

5. The similarity calculation unit further calculates the similarity between the body types of the people that formed the basis of the training data, using the three-dimensional joint point coordinate data which is the training data. The loss calculation unit calculates the loss for the features in the machine learning model, also using the similarity between the physical characteristics of the people that formed the basis of the training data. A learning model generation device according to claim 1.

6. A second loss calculation unit calculates the loss for the output of the machine learning model using the three-dimensional joint point coordinate data output by the machine learning model and the three-dimensional joint point coordinate data which is the training data, in response to the input of the two-dimensional joint point coordinate data by the data acquisition unit. A loss integration unit that integrates the loss for the feature quantity and the loss for the output, Equipped with, The learning model generation unit updates the parameters of the machine learning model using the loss obtained through integration. A learning model generation device according to claim 1.

7. The loss integration unit integrates the loss for the feature and the loss for the output by calculating a weighted average of the two. The learning model generation apparatus according to claim 6.

8. The aforementioned machine learning model is a neural network, The feature distance calculation unit obtains the feature from the intermediate layer of the neural network. A learning model generation device according to claim 1.

9. The training data includes two-dimensional joint point coordinate data that can identify the two-dimensional coordinates of each of the multiple joint points of a person in an image, and three-dimensional joint point coordinate data that can identify the three-dimensional coordinates of each of the multiple joint points of the person. The two-dimensional joint point coordinate data is acquired from this training data, and the acquired two-dimensional joint point coordinate data is input into a machine learning model. For each of the training data, the features calculated in the machine learning model are obtained, and further, the distance between the obtained features is calculated using the obtained features. For each of the training data, an index indicating the relationship between the person and the camera that took the image is determined using information about the camera that took the image, and the similarity between the determined indices is calculated. Using the calculated similarity and the calculated distance between the features, the loss for the features in the machine learning model is calculated. The calculated loss is used to update the parameters of the machine learning model. A method for generating a learning model characterized by the following features.

10. On the computer, The training data includes two-dimensional joint point coordinate data that can identify the two-dimensional coordinates of each of the multiple joint points of a person in an image, and three-dimensional joint point coordinate data that can identify the three-dimensional coordinates of each of the multiple joint points of the person. From this training data, the two-dimensional joint point coordinate data is acquired, and the acquired two-dimensional joint point coordinate data is input into a machine learning model. For each of the training data, the features calculated in the machine learning model are obtained, and further, the distance between the obtained features is calculated using the obtained features. For each of the training data, using information about the camera that took the image, an index indicating the relationship between the person and the camera that took the image is obtained, and the similarity between the obtained indices is calculated. Using the calculated similarity and the calculated distance between the features, the loss for the features in the machine learning model is calculated. The calculated loss is used to update the parameters of the machine learning model. program.

Citation Information

Patent Citations

  • Posture correction network study device and program thereof and posture estimation device and program thereof

    JP2021047563A

  • Method for generating data for estimating three-dimensional pose of object included in input image, computer system, and method for constructing prediction model

    JP2021111380A

  • Learning model generation device, joint point detection device, learning model generation method, joint point detection method, and program

    JP2023167320A

  • Image recognition device, image recognition method, and image recognition program

    WO2021033242A1

  • Information processing device

    WO2023277043A1