Portrait Data Analysis System
The person image data analysis system uses machine learning models to enhance posture discrimination by integrating feature extraction, behavior analysis, and keypoint extraction, achieving accurate posture recognition even in challenging orientations.
Patent Information
- Application Number
- JP2021086077
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-21
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-05-21
AI Technical Summary
Existing methods struggle to accurately discriminate a person's posture, especially when the torso is in a lateral position or twisted, making it difficult to determine the front and rear limbs and connection relationships between joints.
A person image data analysis system utilizing machine learning models for feature extraction, behavior analysis, and keypoint extraction, incorporating a loss function that considers keypoints and behavior types to enhance accuracy in posture discrimination.
The system accurately outputs keypoints of a person's posture with high precision, even in twisted or sideways positions, by considering action types, enabling precise posture analysis.
Smart Images

Figure 0007700511000001 
Figure 0007700511000002 
Figure 0007700511000003
Abstract
Description
Technical Field
[0001] The present invention relates to a system for analyzing human image data.
Background Art
[0002] In recent years, acquiring human image data and evaluating the posture of a person have been carried out. For example, Patent Document 1 describes determining the posture of a worker in order to measure the working hours of the worker at a manufacturing site. The working situation is acquired by a camera, and skeleton data including feature point data indicating the joint positions of the worker shown in the acquired image data is acquired. A posture model in which a posture label is associated with each piece of skeleton data is stored in advance. Then, based on the acquired skeleton data, the posture of the person shown in the image data is determined from the posture label determined in advance in the posture model.
[0003] In addition to determining the posture of a person in order to measure the working hours of the worker, it is also important to evaluate the posture itself of the person. For example, it is sometimes necessary to evaluate whether a person is walking in a correct posture. It is also conceivable to evaluate whether a person using a nursing device such as a walker is using the walker in a correct posture. In addition, various walking assistance devices that assist walking or operate to promote independent walking are known. It is also important to evaluate the posture of a person using a walking assistance device.
[0004] In addition, when a worker in a factory or the like wears an active power assist suit to reduce the work load, it is also important to evaluate the posture of the worker. By evaluating the posture of the worker, it is possible to evaluate whether the worker can appropriately use the assist suit and whether the assist suit is functioning properly.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] As described above, it is very important to evaluate a person's posture. In the method described in Patent Document 1, the posture is discriminated from the skeletal data of a person. In the case of an image data where a person is facing the front of the whole body or facing the back, the posture of the person can be easily discriminated from the skeletal data of the person.
[0007] However, for example, in the case where the torso is in a lateral position, it may not be possible to discriminate the posture of the person only from the skeletal data. For example, when the torso is in a lateral position, it is not easy to discriminate whether the right foot or the left foot is located in the front. Similarly, when the torso is in a lateral position, it is not easy to discriminate which of the right arm and the left arm is located in the front. Also, in the case where a person has a posture in which the upper body and the lower body are twisted, it is not easy to discriminate how each part of the person is positioned.
[0008] The present invention has been made in view of such a background, and aims to provide a person image data analysis system capable of accurately discriminating the posture of a person shown in person image data.
Means for Solving the Problems
[0009] One aspect of the present invention is a person image data analysis system configured by a computer device including an arithmetic processing device and a storage device, wherein the storage device stores a learned model related to feature amount extraction generated by performing machine learning using person image data including a person as an explanatory variable and a feature amount in the person image data as an objective variable. 1st 1st the 1st Using the feature amount extracted based on the portrait data of the person as an explanatory variable, and the plurality of time-series 1st A learned model related to behavior analysis generated by performing machine learning with the behavior type of the person in the portrait data of the person as an objective variable is memorized. Using the feature amount and the behavior type as explanatory variables, and 1st A learned model related to keypoint extraction generated by performing machine learning with the keypoint representing the posture of the person in the portrait data of the person as an objective variable is memorized. The arithmetic processing unit Using the learned model related to feature amount extraction memorized in the storage device, 2nd including a person By inputting portrait data of a person, the 2nd A feature amount extraction unit that extracts the feature amount in the portrait data of the person; Using the learned model related to behavior analysis memorized in the storage device, by inputting the feature amount extracted by the feature amount extraction unit, for the plurality of time-series 2nd A behavior type output unit that outputs the behavior type of the person in the portrait data of the person; Using the learned model related to keypoint extraction memorized in the storage device, by inputting the feature amount extracted by the feature amount extraction unit and the behavior type output by the behavior type output unit, the 2nd A keypoint output unit that outputs the keypoint of the person in the portrait data of the person; is provided 、 The learned model for feature amount extraction, the learned model for behavior analysis, and the learned model for keypoint extraction are learned by a loss function including elements of the keypoints and elements of the behavior types in a learning phase in a portrait data analysis system. Another aspect of the present invention is A person image data analysis system configured by a computer device including an arithmetic processing unit and a storage device, The storage device stores a learned model for feature amount extraction generated by performing machine learning with the 1st person image data including a person as an explanatory variable and the feature amount in the 1st person image data as an objective variable, stores a learned model for behavior analysis generated by performing machine learning with the feature amount extracted based on the 1st person image data as an explanatory variable and the behavior types of the person in a plurality of time-series 1st person image data as an objective variable, stores a learned model for keypoint extraction generated by performing machine learning with the feature amount and the behavior types as explanatory variables and the keypoints representing the posture of the person in the 1st person image data as an objective variable, The arithmetic processing unit a feature amount extraction unit that extracts the feature amount in each of a plurality of time-series 2nd person image data including a person by inputting the plurality of time-series 2nd person image data using the learned model for feature amount extraction stored in the storage device; a behavior type output unit that outputs the behavior types of the person in the plurality of time-series 2nd person image data by inputting the feature amounts for the plurality of images extracted by the feature amount extraction unit based on each of the plurality of time-series 2nd person image data using the learned model for behavior analysis stored in the storage device; a keypoint output unit that outputs the keypoints of the person in the selected 1st 2nd person image data by inputting the feature amount for 1 image extracted by the feature amount extraction unit based on the selected 1st 2nd person image data among the plurality of time-series 2nd person image data and the behavior types output by the behavior type output unit using the learned model for keypoint extraction stored in the storage device; It is in a human image data analysis system including Another aspect of the present invention is A human image data analysis system constituted by a computer device including an arithmetic processing unit and a storage device, The storage device Stores a learned model related to feature quantity extraction generated by performing machine learning with the first human image data including a person as an explanatory variable and the feature quantity in the first human image data as an objective variable, Stores a learned model related to behavior analysis generated by performing machine learning with the feature quantity for one sheet extracted based on the first human image data as an explanatory variable and the behavior types of the person in a plurality of sheets of the first human image data in time series as an objective variable, Stores a learned model related to keypoint extraction generated by performing machine learning with the feature quantity and the behavior type as explanatory variables and the keypoint representing the posture of the person in the first human image data as an objective variable, The arithmetic processing unit A feature quantity extraction unit that sequentially inputs a plurality of sheets of second human image data in time series including a person by using the learned model related to feature quantity extraction stored in the storage device, and sequentially extracts the feature quantity in each of the plurality of sheets of the second human image data, An action type output unit that sequentially inputs the feature quantity for one sheet extracted by the feature quantity extraction unit by using the learned model related to action analysis stored in the storage device, and performs a recursive operation using the result of the previous arithmetic processing, thereby outputting the action type of the person in one sheet of the second human image data that is the target of the current arithmetic processing, A keypoint output unit that inputs the feature quantity for one sheet extracted by the feature quantity extraction unit and the action type output by the action type output unit by using the learned model related to keypoint extraction stored in the storage device, thereby outputting the keypoint of the person in one sheet of the second human image data that is the target of the current arithmetic processing, It is in a human image data analysis system including
Advantages of the Invention
[0010] The keypoint output unit does not output the keypoints of the person only using the feature amounts in the person image data. The keypoint output unit inputs the action type of the person in addition to the feature amounts in the person image data and outputs the keypoints of the person.
[0011] In this way, by outputting the keypoints of the person after grasping the action type of the person, the keypoint output unit can output the keypoints of the person with high accuracy. For example, when the person is in a posture where the upper body and the lower body are twisted, there is a possibility that the connection relationship connecting adjacent joint positions, which is one of the keypoints, may be output erroneously. However, by grasping the action type of the person, even in a twisted posture, the connection relationship as one of the keypoints can be output with high accuracy. Therefore, the posture of the person can be analyzed with high accuracy.
Brief Description of the Drawings
[0012]
FIG. 1
FIG. 2
FIG. 3
FIG. 4
FIG. 5
FIG. 6
FIG. 7
FIG. 8
Modes for Carrying Out the Invention
[0013] (1. Overview of the Person Image Data Analysis System) The person image data analysis system acquires person image data and analyzes the posture of the person included in the acquired person image data. The posture of the person is classified, for example, into standing position, sitting position, lying position, kneeling position, etc., and is further classified in more detail in each case. Furthermore, the posture of the person also differs depending on whether it is a stationary state or a motion state. That is, the person image data analysis system analyzes what posture the person shown in the person image data is in.
[0014] The posture information of a person analyzed by a person portrait data analysis system is used, for example, as follows. When the person is in a stationary state, the posture of the person is evaluated. For example, when the person is in a standing posture, an evaluation is made as to whether it is an appropriate standing posture, and the person can be improved to have an appropriate standing posture. Also, when the person is in a sitting posture or a lying posture, an evaluation is made as to whether it is an appropriate sitting posture or lying posture, and it can be used for the selection of an appropriate seat or bedding, or for the development of a seat or bedding.
[0015] It can also be used to evaluate the posture of a person during movement. The postures in movements such as from a standing posture to a sitting posture, the reverse movement, from a sitting posture to a lying posture, and the reverse movement can be evaluated. Also, the postures during walking, running, jumping, etc. can be evaluated, and furthermore, various postures of a person while doing sports can be evaluated.
[0016] Furthermore, it is also possible to evaluate whether a person using a care device such as a walker is using the walker in a correct posture. Also, in a walking support device that drives to assist walking or promote independent walking, the posture of a person using the walking support device can be evaluated. Using the evaluation result of the person's posture, it is possible to evaluate whether the walking support device is functioning properly. Furthermore, the posture of the person using the walking support device can be analyzed, and using the analysis result, the control of the walking support device can also be performed.
[0017] Also, when a care recipient or a worker in a factory or the like wears an active power assist suit to reduce the operation load, the posture of the wearer can be evaluated. Using the evaluation result of the wearer's posture, it is possible to evaluate whether the assist suit is functioning properly. Furthermore, the posture of the wearer can be analyzed, and using the analysis result, the control of the assist suit can also be performed. Also, by analyzing the posture of a worker in a factory or the like, it is possible to evaluate the working hours of the worker. Furthermore, it is also possible to evaluate the working hours for each type of work performed by the worker.
[0018] (2. First Embodiment) (2-1. Configuration in the Inference Phase of the Person Image Data Analysis System 1) The configuration of the person image data analysis system 1 will be described with reference to FIGS. 1 to 6. In particular, hereinafter, the configuration in the inference phase of the person image data analysis system 1 will be described. As shown in FIG. 1, the person image data analysis system 1 is composed of an imaging device 2 and a computer device used for analysis. The computer device includes a storage device 3 and an arithmetic processing device 4.
[0019] The imaging device 2 is, for example, a moving image imaging device capable of imaging a moving image continuously in time series, a still image imaging device capable of imaging a still image in time series, or the like. The imaging device 2 is used to image so as to include a person who is the object of posture analysis. The storage device 3 stores the learned models A, B, C, and D generated by machine learning. The arithmetic processing device 4 includes a person image data generation unit 11, a feature amount extraction unit 12, an action type output unit 13, and a keypoint output unit 14.
[0020] As shown in FIG. 2, the learned model A is a machine learning model related to the extraction of person image data generated by performing machine learning. When the learned model A inputs the image data D1 (hereinafter referred to as "original image data") imaged by the imaging device 2, it extracts the person region D1a from the original image data D1. The original image data D1 includes the person region D1a and a peripheral region D1b located around the person region D1a. The person region D1a may include, in addition to the person, an object held by the person.
[0021] Then, when the original image data D1 is input, the learned model A outputs the person image data D2, which is the image data of the extracted person region D1a. The learned model A applies, for example, R-CNN (Regions with Convolutional Neural Networks) or the like. The learned model A extracts the person region D1a by, for example, a rectangular region (bounding box).
[0022] As shown in FIG. 3, the learned model B is a machine learning model for feature extraction generated by performing machine learning. The learned model B is preferably a machine learning algorithm (including deep learning) including a neural network, for example, but other machine learning algorithms may also be applied. The learned model B is a machine learning model generated by performing machine learning, using the person image data D2 including the person output by the learned model A as an explanatory variable and the feature amount in the person image data D2 as an objective variable. That is, when the person image data D2 is input, the learned model B outputs the feature amount in the person image data D2.
[0023] Note that the types of feature amounts extracted by the learned model B may be set in advance, or may be automatically extracted by machine learning. Of course, the types of feature amounts may be used in combination with automatic extraction by machine learning and setting by a setter. For example, after the types of feature amounts are automatically extracted by machine learning, a setter may perform a correction setting.
[0024] The learned model C is a machine learning model for behavior analysis generated by performing machine learning. The learned model C is preferably a machine learning algorithm (including deep learning) including a neural network, for example, but other machine learning algorithms may also be applied. The learned model C is a machine learning model generated by performing machine learning, using the feature amounts for a plurality of sheets extracted by the learned model B based on each of a plurality of sheets of person image data D2 in time series as explanatory variables and the types of behaviors of the person in the plurality of sheets of person image data D2 in time series as an objective variable. Here, the number of sheets of the feature amounts for a plurality of sheets as explanatory variables, the time in time series, etc. can be arbitrarily set.
[0025] The types of human actions can be broadly classified, for example, into standing postures, sitting postures, lying postures, kneeling postures in a static state, walking postures, running postures, jumping postures, and postures when performing various sports in a moving state. The types of human actions are further classified in more detail within this broad classification. For example, sitting postures can be classified into seiza, casual sitting, formal sitting, long sitting, edge sitting, semi-sitting, etc. Also, lying postures can be classified into supine position, lateral recumbent position, prone position, etc. Other postures are also classified in detail.
[0026] When the learned model C receives as input the feature quantities for a plurality of images corresponding to the plurality of time-series human image data D2, which were extracted by the learned model B based on each of the human image data D2, it generates a score for the type of action of the human as shown in FIG. 4. Then, the learned model C determines the type of action with the highest score value as the type of action of the human and outputs the type of action.
[0027] The learned model D is a machine learning model for keypoint extraction generated by performing machine learning. The learned model D is preferably a machine learning algorithm (including deep learning) that includes, for example, a neural network, but other machine learning algorithms may also be applied. The learned model D is a machine learning model generated by performing machine learning, with feature quantities and action types as explanatory variables and keypoints representing the posture of the human in the human image data D2 as the objective variable. The feature quantities are the information output by the learned model B. The action types are the information output by the learned model C.
[0028] Regarding the key points, they will be described with reference to FIG. 5. The key points include the joint positions of the person indicated by a part of the black circles in FIG. 5. In this embodiment, the key points include the positions of the eyes of the person as shown by a part of the black circles in FIG. 5. Further, the key points include the connection relationships indicated by the lines connecting the black circles in FIG. 5. For example, the key points include the connection relationships connecting adjacent joint positions, the connection relationships connecting adjacent eye positions, and the connection relationships connecting the eye positions and the joint positions close to the eye positions. That is, the key points are feature data including the parts for expressing the posture of the person and the connection relationships of each part. And when the learned model D is input with the feature amount and the action type, it outputs the key points shown in FIG. 5.
[0029] As shown in FIGS. 1 and 2, the person image data generation unit 11 acquires the original image data D1 from the imaging device 2. The person image data generation unit 11 uses the learned model A stored in the storage device 3 and inputs the original image data D1, thereby generating person image data D2 in which the person area D1a is extracted from the original image data D1.
[0030] When the original image data D1 is moving image data, the person image data generation unit 11 generates a plurality of still image data in time series from the acquired moving image data. Then, the person image data generation unit 11 inputs each of the plurality of generated still image data in time series (for example, times T1 to T10) to the learned model A, and generates person image data D2 in each of the plurality of still image data. That is, the person image data generation unit 11 generates, for example, a plurality of person image data D2 at times T1 to T10.
[0031] When the original image data D1 is still image data, the person image data generation unit 11 inputs each of the plurality of still image data obtained from the acquired time series (for example, T1 to T10) into the learned model A, and generates person image data D2 in each of the plurality of still image data. That is, also in this case, the person image data generation unit 11 generates, for example, a plurality of person image data D2 at times T1 to T10.
[0032] The person image data D2 includes various image data such as image data of a posture in which at least the torso of a person faces the imaging device 2, image data of a posture in which at least the torso of the person faces away, and image data of a posture in which at least the torso of the person is sideways. Here, the sideways direction is not limited to the case where it is at a 90° angle to the imaging device 2, but means excluding the cases of completely facing the imaging device 2 and completely facing away, and includes the case of facing in an oblique direction.
[0033] In addition, the person image data D2 also includes image data of a posture in which the upper body and the lower body of the person are not twisted, and image data of a twisted posture. When the person is walking, the left hand and the right foot may be positioned forward, and the right hand and the left foot may be positioned rearward. In such a case, the upper body and the lower body of the person are in a twisted posture.
[0034] As shown in FIGS. 1 and 3, the feature extraction unit 12 acquires a plurality of person image data D2 having a time series (T1 to T10) generated by the person image data generation unit 11. The feature extraction unit 12 inputs a plurality of person image data D2 having a time series (T1 to T10) using the learned model B stored in the storage device 3. Then, the feature extraction unit 12 extracts the feature amounts in each of the plurality of person image data D2, that is, the feature amounts for a plurality of sheets, as the output of the learned model B.
[0035] As shown in FIGS. 1 and 3, the action type output unit 13 acquires feature amounts for a plurality of images extracted by the feature amount extraction unit 12. The action type output unit 13 performs a process of inputting the feature amounts for a plurality of images extracted based on a plurality of pieces of person image data D2 having a time series (T1 to T10) into the learned model B using the learned model C stored in the storage device 3.
[0036] Then, the action type output unit 13 outputs the action type of the person in the plurality of pieces of person image data D2 in the time series using the feature amounts for a plurality of images in the time series (T1 to T10). Specifically, as shown in FIG. 4, the action type output unit 13 generates a score for each action type and outputs the action type with the highest score value as the action type of the person.
[0037] In this embodiment, the action type output unit 13 inputs not the feature amount in one piece of person image data D2 but the feature amounts in a plurality of pieces of person image data D2, that is, the feature amounts for a plurality of images. That is, the action type is specified by determining the change in the position of the person in the plurality of pieces of person image data D2 in the time series.
[0038] As shown in FIGS. 1 and 3, the keypoint output unit 14 acquires the feature amounts extracted by the feature amount extraction unit 12. As described above, the feature amount extraction unit 12 extracts the feature amounts for each of a plurality of pieces of person image data D2 having a time series (T1 to T10), that is, the feature amounts for a plurality of images.
[0039] However, the keypoint output unit 14 does not need to use the feature amounts for a plurality of images in the time series (T1 to T10). In this embodiment, the keypoint output unit 14 acquires the feature amount for one image extracted based on one piece of person image data D2 selected from the plurality of pieces of person image data D2 having a time series (T1 to T10). For example, the keypoint output unit 14 acquires the feature amount extracted based on the person image data D2 at the intermediate time T5 among the times T1 to T10. Note that the time selected by the keypoint output unit 14 can be arbitrarily determined.
[0040] Furthermore, the keypoint output unit 14 acquires the action type output by the action type output unit 13. The keypoint output unit 14 performs a process of inputting the acquired feature amount and action type into the learned model C stored in the storage device 3 using the learned model C, thereby outputting the keypoints of the person in the person image data D2 at time T5.
[0041] As shown in FIG. 5, the keypoint output unit 14 outputs the joint position, the eye position, and the connection relationship connecting each position as the keypoints of the person in the person image data D2 at time T5.
[0042] (2-2. Configuration in the learning phase of the person image data analysis system 1) The configuration in the learning phase of the person image data analysis system 1 will be described with reference to FIG. 6. In particular, the learning phases regarding models B, C, and D will be described.
[0043] First, a training data set for learning is prepared. As the training data set, a large number of units each consisting of a plurality of time-series person image data D2 are prepared. For example, since a plurality of moving image data includes a large number of units each consisting of a plurality of time-series person image data D2, it is suitable as a training data set. Furthermore, the training data set includes label information regarding the keypoints of the person in the person image data D2 and the action types of the person.
[0044] The loss function F(x, y) used for learning includes the element x of the keypoint and the element y of the action type. Models B, C, and D are trained by inputting the training data set so as to minimize the loss function F(x, y). Since the loss function F(x, y) has the element of the keypoint and the element of the action type, models B, C, and D are trained to output the correct answers of the keypoints and the action types. The learned models B, C, and D thus learned are stored in the storage device 3.
[0045] The learning using the loss function F(x, y) as described above does not train models B, C, and D independently, but treats models B, C, and D as an integrated model for training. Therefore, in models B, C, and D, the parts affected by the loss function (x, y) are effectively learned respectively.
[0046] (2-3. Effect) In the human image data analysis system 1, the keypoint output unit 14 does not output the keypoints of the person only using the feature amounts in the human image data D2. The keypoint output unit 14 inputs the action type of the person in addition to the feature amounts in the human image data D2, and outputs the keypoints of the person.
[0047] In this way, the keypoint output unit 14 can output the keypoints of the person with high accuracy by outputting the keypoints of the person after grasping the action type of the person. This will be described by comparing FIG. 5 which is the output result of the keypoints in this embodiment and FIG. 7 which is the output result of the keypoints as a comparative example.
[0048] FIG. 5 shows the keypoints output by the keypoint output unit 14 in this embodiment. On the other hand, FIG. 7 shows the keypoints output based only on the feature amounts in the human image data D2 without considering the action type. The human image data D2 used for the keypoints shown in FIGS. 5 and 7 is an image data of a person in a posture where the upper body and the lower body are twisted. Further, the human image data D2 is image data of a person in a posture where at least the torso is horizontal.
[0049] In the lower body of the person shown in FIG. 5, the right hip joint and the right knee joint are connected, and the left hip joint and the left knee joint are connected. In this way, in FIG. 5, the joints are correctly connected. On the other hand, in the lower body of the person shown in FIG. 7, the right hip joint and the left knee joint are connected, and the left hip joint and the right knee joint are connected. That is, in FIG. 7, the joints are incorrectly connected.
[0050] In the lower body of the person shown in FIGS. 5 and 7, the right hip joint is located closer to the left knee joint than the right knee joint, and the left hip joint is located closer to the right knee joint than the left knee joint. And since the person image data is in a posture where the person's torso is sideways, the left and right hip joints and the left and right knee joints have opposite left and right front-back positions. In FIG. 7, it seems that the joints located close to each other are connected.
[0051] As shown in FIG. 7, when the person is in a posture where the upper body and the lower body are twisted, there is a possibility that the connection relationship connecting adjacent joint positions, which is one of the key points, may be incorrectly output. If the connection of the joint positions is not correctly recognized, the person's posture cannot be correctly recognized. However, in this embodiment, as shown in FIG. 5, by grasping the type of action of the person, even in a twisted and sideways posture, the connection relationship as one of the key points can be output with high accuracy. Therefore, the person's posture can be analyzed with high accuracy.
[0052] The action type output unit 13 outputs the type of action of the person by inputting feature amounts for a plurality of consecutive frames in time series. Therefore, the action type output unit 13 can accurately identify the type of action of the person by using a plurality of consecutive person image data D2 in time series. As a result, the key points of the person can be output with high accuracy.
[0053] Also, the learned models B, C, and D are learned by the loss function F(x, y) including the elements of the key points and the elements of the action type in the learning phase. That is, the learned models B, C, and D are learned so that the key point output unit 14 can output the key points considering the action type with high accuracy. By using the learned models B, C, and D learned in this way to output the key points of the person, high-precision key points can be output.
[0054] Further, the person image data analysis system 1 does not input the original image data D1 captured by the imaging device 2 itself into the feature extraction unit 12, but inputs the person image data D2 from which the person region D1a has been extracted from the original image data D1 into the feature extraction unit 12. In this way, by generating the person image data D2 from which the person region D1a has been extracted, it leads to outputting the key points of the person in the person image data D2 with high accuracy.
[0055] (3. Second Embodiment) Regarding the configuration of the inference phase of the person image data analysis system 1 in the second embodiment, it will be described with reference to FIGS. 1 and 8.
[0056] As shown in FIG. 1, the person image data analysis system 1 includes a storage device 3 that stores learned models A, B, C, D, and an arithmetic processing device 4. The arithmetic processing device 4 includes a person image data generation unit 11, a feature extraction unit 12, an action type output unit 13, and a key point output unit 14.
[0057] The learned models A, B, D are the same as the learned models A, B, D in the first embodiment. The learned model C applies a recursive algorithm. For example, the learned model C applies RNN (Recurrent Neural Network), LSTM (Long Short Term Memory), etc.
[0058] That is, when the learned model C sequentially inputs the feature amounts for one sheet extracted based on one sheet of person image data D2 by the feature extraction unit 12, it performs a recursive operation using the result of the previous arithmetic processing, and outputs the action type of the person in one sheet of person image data D2 that is the target of the current arithmetic processing. It is a machine learning model.
[0059] In this embodiment, the processing of each part constituting the arithmetic processing unit 4 is as follows. The feature amount extraction unit 12 sequentially extracts the feature amounts in each of the plurality of piece of time-series person image data D2 by sequentially inputting the plurality of piece of time-series person image data D2 using the learned model A. That is, the feature amount extraction unit 12 sequentially extracts the feature amount of one piece of person image data D2 that is the target of arithmetic processing.
[0060] The action type output unit 13 sequentially inputs the feature amount for one piece extracted from one piece of person image data D2 by the feature amount extraction unit 12 using the learned model C, and performs a recursive operation using the result of the previous arithmetic processing, thereby outputting the action type of the person in one piece of person image data D2 that is the target of the current arithmetic processing.
[0061] The keypoint output unit 14 inputs the feature amount for one piece that is the target of the current arithmetic processing and the action type of the person in the person image data D2 including the target of the current arithmetic processing using the learned model D, thereby outputting the keypoints of the person in one piece of person image data D2 that is the target of the current arithmetic processing.
[0062] The action type output unit 13 can output the action type of the person using the feature amount for one piece that is the target of the current arithmetic processing by performing a recursive operation. Therefore, the processing in the feature amount extraction unit 12, the action type output unit 13, and the keypoint output unit 14 is executed by inputting one piece of person image data D2 that is the target of the current arithmetic processing. Therefore, each time the time-series person image data is sequentially input, the keypoints of the person in the person image data can be output. That is, the keypoints of the person can be output in real time. As a result, the posture of the person can be analyzed in real time.
Explanation of Signs
[0063] 1 Person Image Data Analysis System 3 Storage Device 4 Arithmetic Processing Unit 11 Person Image Data Generation Unit 12 Feature extraction unit 13 Action type output unit 14 Keypoint output unit D1 Original image data D2 Person image data
Claims
1. A human image data analysis system configured by a computer device including an arithmetic processing unit and a storage device, wherein the storage device stores a learned model related to feature extraction generated by performing machine learning using first human image data including a person as an explanatory variable and a feature amount in the first human image data as an objective variable, stores a learned model related to action analysis generated by performing machine learning using the feature amount extracted based on the first human image data as an explanatory variable and the action type of the person in a plurality of time-series first human image data as an objective variable, stores a learned model related to keypoint extraction generated by performing machine learning using the feature amount and the action type as explanatory variables and a keypoint representing the posture of the person in the first human image data as an objective variable, wherein the arithmetic processing unit a feature amount extraction unit that extracts the feature amount in the second human image data by inputting the second human image data including a person using the learned model related to feature extraction stored in the storage device; an action type output unit that outputs the action type of the person in a plurality of time-series second human image data by inputting the feature amount extracted by the feature amount extraction unit using the learned model related to action analysis stored in the storage device; a keypoint output unit that outputs the keypoint of the person in the second human image data by inputting the feature amount extracted by the feature amount extraction unit and the action type output by the action type output unit using the learned model related to keypoint extraction stored in the storage device; and the learned model related to feature extraction, the learned model related to action analysis, and the learned model related to keypoint extraction are learned by a loss function including elements of the keypoint and elements of the action type in a learning phase, a human image data analysis system.
2. The feature amount extraction unit extracts the feature amount in each of the plurality of time-series second human image data by inputting the plurality of time-series second human image data using the learned model related to feature extraction, The action type output unit inputs the feature quantities for a plurality of sheets extracted based on each of the plurality of pieces of the second person image data in time series, using the learned model related to the action analysis, and outputs the action type. The keypoint output unit inputs the feature quantity for one sheet extracted based on one piece of the second person image data selected from the plurality of pieces of the second person image data in time series, and the action type, and outputs the keypoints of the person in the selected one piece of the second person image data. The person image data analysis system according to claim 1.
3. The feature quantity extraction unit sequentially inputs the plurality of pieces of the second person image data in time series, using the learned model related to the feature quantity extraction, and sequentially extracts the feature quantities in each of the plurality of pieces of the second person image data. The action type output unit sequentially inputs the feature quantity for one sheet that has been extracted, using the learned model related to the action analysis, and performs a recursive operation using the result of the previous arithmetic process, and outputs the action type of the person in one piece of the second person image data that is the target of the current arithmetic process. The keypoint output unit inputs the feature quantity for one sheet and the action type, using the learned model related to the keypoint extraction, and outputs the keypoints of the person in one piece of the second person image data that is the target of the current arithmetic process. The person image data analysis system according to claim 1.
4. A person image data analysis system configured by a computer device including an arithmetic processing device and a storage device, wherein the storage device stores a learned model related to feature quantity extraction generated by performing machine learning with the first person image data including a person as an explanatory variable and the feature quantity in the first person image data as an objective variable, stores a learned model related to action analysis generated by performing machine learning with the feature quantity extracted based on the first person image data as an explanatory variable and the action type of the person in the plurality of pieces of the first person image data in time series as an objective variable. A learned model for keypoint extraction generated by performing machine learning with the feature amount and the action type as explanatory variables and the keypoints representing the posture of the person in the first person image data as the objective variable is memorized. The arithmetic processing unit A feature amount extraction unit that extracts the feature amount in each of the plurality of second person image data in a time series including a person by inputting the plurality of second person image data in a time series including a person using the learned model for feature amount extraction memorized in the storage device. An action type output unit that outputs the action type of the person in the plurality of second person image data in a time series by inputting the feature amounts for the plurality of second person image data extracted by the feature amount extraction unit based on each of the plurality of second person image data in a time series using the learned model for action analysis memorized in the storage device. A keypoint output unit that outputs the keypoints of the person in the selected one second person image data by inputting the feature amount for one second person image data extracted by the feature amount extraction unit based on the selected one second person image data among the plurality of second person image data in a time series and the action type output by the action type output unit, using the learned model for keypoint extraction memorized in the storage device. A person image data analysis system comprising the above.
5. A person image data analysis system configured by a computer device including an arithmetic processing unit and a storage device, The storage device Memorizes a learned model for feature amount extraction generated by performing machine learning with the first person image data including a person as an explanatory variable and the feature amount in the first person image data as an objective variable. Memorizes a learned model for action analysis generated by performing machine learning with the feature amount for one first person image data extracted based on the first person image data as an explanatory variable and the action type of the person in the plurality of first person image data in a time series as an objective variable. Memorizes a learned model for keypoint extraction generated by performing machine learning with the feature amount and the action type as explanatory variables and the keypoints representing the posture of the person in the first person image data as the objective variable. The arithmetic processing unit By sequentially inputting a plurality of pieces of second person image data in a time series including a person, using the learned model related to the feature amount extraction stored in the memory device, a feature amount extraction unit that sequentially extracts the feature amounts in each of the plurality of pieces of the second person image data Using the learned model related to the action analysis stored in the memory device, sequentially inputting the feature amount for one piece extracted by the feature amount extraction unit, and performing a recursive operation using the result of the previous arithmetic processing, an action type output unit that outputs the action type of the person in one piece of the second person image data that is the target of the current arithmetic processing Using the learned model related to the keypoint extraction stored in the memory device, by inputting the feature amount for one piece extracted by the feature amount extraction unit and the action type output by the action type output unit, a keypoint output unit that outputs the keypoint of the person in one piece of the second person image data that is the target of the current arithmetic processing A person image data analysis system comprising:
6. The person image data analysis system according to any one of claims 1 to 5, wherein the first person image data and the second person image data include image data of a posture in which at least the torso of the person is horizontal.
7. The arithmetic processing device further includes a person image data generation unit that inputs original image data including a person region and a peripheral region, and generates the first person image data and the second person image data from which the person region is extracted from the original image data. The person image data analysis system according to any one of claims 1 to 6.
8. The person image data analysis system according to any one of claims 1 to 7, wherein the keypoint includes a joint position of the person and a connection relationship connecting adjacent joint positions.
Citation Information
Patent Citations
Device and program for deciding personal action
JP2011100175A
Program, apparatus, and method for describing trajectory of displacement of human skeleton position from video data
JP2019219836A
Attitude analysis program and attitude analyzer
JP2020201772A
Skeleton -based action detection using recurrent neural network
US20170344829A1