Machine learning device, skill determination device, machine learning method, and machine learning program
By selecting user-edited image frames in dynamic images to generate the region of interest, the problem of unstable skill evaluation of the subject in dynamic images is solved, and more accurate and stable skill judgment is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, when evaluating the skills of subjects performing actions in dynamic images, the learning behavior is unstable and easily affected by non-focused areas, leading to inaccurate judgments of skill quality.
The data prioritization department selects user-edited image frames from the dynamic image dataset, generates regions of interest, and updates the learning model based on these regions, thereby improving the accuracy and stability of skill level assessment.
This improves the accuracy of skill level assessment for subjects performing actions in dynamic images and the stability of learning behavior, ensuring the effectiveness and accuracy of the learning model.
Smart Images

Figure CN121753064A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to machine learning devices, skill assessment devices, machine learning methods, and machine learning programs. Background Technology
[0002] Pairwise Deep Ranking (PDR) has been proposed. This technique is one of the methods for SkillAssessment, which calculates a score for a person's skill level (i.e., proficiency) to determine the quality of the skill (e.g., see Non-Patent Literature 1).
[0003] Existing technical documents
[0004] Non-patent literature
[0005] Non-patent document 1: Takada Masayuki, three others (Chubu University), "Attention Pairwise" Ranknggによる’s skill judgment and explanation of high-precision visual inspection, Chapter 23 understanding of the portrait Understand シンポジウム Summary of the Invention
[0006] The problem that the invention aims to solve
[0007] However, in the aforementioned prior art, the information provided is merely the quality of the skill and the dynamic image pair. Sometimes, the quality of the skill is judged based on the field of focus that should be paid attention to in order to determine the quality, and thus, on a field outside the field of focus (i.e., the deviated area of focus in the dynamic image pair). For example, in the skill evaluation of the action of drawing a picture, the quality of the skill is sometimes judged not based on the dynamic image during the period of moving the pen, but on the dynamic image during the period of the person moving the pen's head. Therefore, in the prior art, there is a problem that the learning behavior sometimes becomes unstable.
[0008] The purpose of this disclosure is to stabilize the learning behavior when learning a learning model for inferring the skill level of a subject's actions in a moving image.
[0009] Methods for solving problems
[0010] The machine learning apparatus disclosed herein is an apparatus for learning a learning model for inferring the skill level of an actor's actions in a moving image. The apparatus is characterized by comprising: a data priority selection unit that selects multiple pairs of moving images from a learning dataset and selects image frames from each of the selected multiple pairs of moving images for determining the quality of the skill level; a region of interest generation unit that generates regions of interest in the image frames for determining the quality of the skill level; a quality determination unit that, for each pair of moving images, uses the learning model to determine the quality of the skill level in the regions of interest; and a model learning unit that stores the learning model and updates the learning model based on the quality determination result of the skill level. The data priority selection unit selects the image frames for determining the quality of the skill level using a selection probability determined based on one or more of the following: user editing in the image frames constituting each pair of moving images, temporal direction features in consecutive multiple image frames constituting each pair of moving images, and similarity between the image frames constituting each pair of moving images.
[0011] The machine learning method disclosed herein is a method for learning a learning model for inferring the skill level of an actor's actions in a moving image. The method is characterized by the following steps: selecting multiple pairs of moving images from a dataset of moving images for learning; selecting image frames from each of the selected multiple pairs of moving images for determining the quality of the skill level; generating regions of interest in the image frames for determining the quality of the skill level; using the learning model to determine the quality of the skill level in the regions of interest for each pair of moving images; and storing the learning model and updating the learning model based on the results of the skill level determination. In the step of selecting image frames for determining the quality of the skill level, the data priority selection unit selects the image frames for determining the quality of the skill level using a selection probability determined based on one or more of the following: user editing in the image frames constituting each pair of moving images, temporal direction features in a series of consecutive image frames constituting each pair of moving images, and similarity between the image frames constituting each pair of moving images.
[0012] Invention Effects
[0013] According to this disclosure, it is possible to stabilize the learning behavior when learning a learning model for inferring the skill level of a subject's actions in a moving image. Attached Figure Description
[0014] Figure 1 This is an explanatory diagram illustrating an existing system (Comparative Example 1) used to perform skill evaluation.
[0015] Figure 2 This is a functional block diagram that roughly represents the structure of the system (Comparative Example 1).
[0016] Figure 3 This is an explanatory diagram of the actions of an existing machine learning device (Comparative Example 2) used to determine the skill level of a person's actions reflected in a dynamic image using transfer learning.
[0017] Figure 4 It is a general representation Figure 3 Functional block diagram of the structure of the machine learning device (Comparative Example 2).
[0018] Figure 5 This is a functional block diagram that roughly represents the structure of the machine learning device in Implementation 1.
[0019] Figure 6 This is a diagram illustrating an example of the hardware structure of the machine learning device (or skill determination device) according to Implementation Method 1.
[0020] Figure 7 (A) is a diagram showing a segment of a moving image selected from a pair of moving images by the random data selection unit of Comparative Example 1, and (B) is a diagram showing a portion selected from a pair of moving images by the data priority selection unit of Embodiment 1.
[0021] Figure 8 This is an explanatory diagram showing the effect of the machine learning device according to Embodiment 1.
[0022] Figure 9 This is an explanatory diagram showing the operation of the machine learning device according to Embodiment 1.
[0023] Figure 10 This is a flowchart illustrating the learning actions of the machine learning device in Implementation 1.
[0024] Figure 11 This is a flowchart illustrating the annotations of the machine learning apparatus of Implementation Method 1.
[0025] Figure 12 This is a functional block diagram that roughly represents the structure of the machine learning device in Implementation 2.
[0026] Figure 13 This is a flowchart illustrating the learning actions of the machine learning device in Implementation Method 2.
[0027] Figure 14 This is a functional block diagram that roughly represents the structure of the machine learning device in Implementation 3.
[0028] Figure 15 This is a flowchart illustrating the learning actions of the machine learning device in Implementation 3.
[0029] Figure 16 This is a functional block diagram that roughly represents the structure of the machine learning device in Implementation 4.
[0030] Figure 17 This is a flowchart illustrating the learning actions of the machine learning device in Implementation 4.
[0031] Figure 18 This is a functional block diagram that roughly represents the structure of the machine learning device in Implementation 5.
[0032] Figure 19 This is a flowchart illustrating the learning actions of the machine learning device in Implementation 5. Detailed Implementation
[0033] Hereinafter, the machine learning apparatus, skill determination apparatus, machine learning method, and machine learning program according to embodiments will be described with reference to the accompanying drawings. The following embodiments are merely examples, and embodiments can be appropriately combined and modified.
[0034] The machine learning apparatus of this embodiment is an apparatus that learns a learning model used in a deduction device (also called a "skill determination device") for deducing the skill level (i.e., proficiency) of an action subject reflected in a moving image. The machine learning apparatus of this embodiment is, for example, a computer that is an information processing device. The action subject reflected in the moving image is a person performing a task (also called a "worker"). Furthermore, the action subject reflected in the moving image may include a mechanism (e.g., a robotic arm, endoscope, etc.) that moves in conjunction with the person's movements to perform the task.
[0035] The machine learning method described in this embodiment is a method that can be implemented by a machine learning device. The machine learning method described in this embodiment is a method for learning a learning model used to derive the skill level of a subject's actions in a moving image.
[0036] The machine learning program of the implementation method is a software program that can be executed by a computer as a machine learning device. The machine learning program of the implementation method is a program that learns a learning model for inferring the skill level of a subject's actions reflected in a moving image.
[0037] Comparative Example 1
[0038] Figure 1 This is an explanatory diagram illustrating an existing system (Comparative Example 1) used for determining skill levels. Figure 1 In this context, the structure proposed in the aforementioned non-patent document 1 will be represented as system 210. Figure 2This is a functional block diagram that roughly represents the structure of system 210 (Comparative Example 1). The input to the system of Comparative Example 1 is a pair of dynamic images (i.e., two dynamic images P). i P j ), P i It is more than P j The system consists of a processing unit that segments the dynamic image, a feature extractor that extracts features from the dynamic image, a superior network that evaluates excellent actions, and an inferior network that evaluates poor actions. The superior and inferior networks each contain an attention branch and a ranking branch, respectively. The input P... i The output value obtained in this case is based on the Score (p) of the output value from the high-level network and the output value from the secondary network. i ). (When entering P) j The output value obtained in this case is based on the Score (p) of the output value from the high-level network and the output value from the secondary network. j System 210 learns the relationship between the scores of these two dynamic images. In dynamic image P... i Compared to dynamic images P j Excellent and Score (P) i ) and Score (P j In the case of a reversal of the magnitude of the scores, the difference in scores is provided to the loss function for learning. The loss function can be, for example, a marginal loss that learns only from differences greater than a certain threshold, or a soft-plus function that evaluates losses for differences less than a certain threshold.
[0039] Comparative Example 2
[0040] Figure 3 This is an illustration of an existing machine learning device (Comparative Example 2) used to determine the skill level of a person's actions reflected in a moving image using transfer learning. Figure 3 This indicates the structure of the machine learning device 220 proposed in Non-Patent Document 2. Figure 4 It means Figure 3 Functional block diagram of the learning model.
[0041] Non-patent literature 2: Masahiro Mitsuhara, et al., Embedding Human Knowledge into Deep Neural Network via Attention Map, arXiv: 1905.03540, May 9, 2019
[0042] In Comparative Example 2, a human-modified attention mechanism (the network that generates the region of interest) performs transfer learning on the region of interest generated from the dynamic image (i.e., embedding human insights into the learning model), using the modified region of interest as the corrective data. Transfer learning is a Human-in-the-Loop (HITL) type of learning. Through transfer learning, for example, a learning model can be generated to determine the skill level of a person's actions within a dynamic image while interacting with a user.
[0043] For example, in Figure 3 In the process, the data selection unit selects a pair of animated images (i.e., animated image P as animated image #1) from the dataset storage unit. i And as dynamic image #2, dynamic image P j At this time, the dynamic image P j The skill level is higher than that of dynamic images P i Excellent skill level (also described as "P") i <P j In this case, it should be determined as a dynamic image P. j The score (P) of the skills reflected in the text. j ) Compared to dynamic images P i The score (P) of the skills reflected in the text. i )big.
[0044] exist Figure 3 The image P shows a dynamic image of the region of interest that the region of interest generation department focuses on. j The score of the skills reflected in the middle. att (P) j =0.1 compared to dynamic image P i The score of the skills reflected in the middle. att (P) i Examples of small values (i.e., scores for the region of interest that are inverted relative to their expected scores) and dynamic images of the output from the FC layer, P. j The score of the skills reflected in the middle. rank (P) j =0.3 compared to dynamic image P i The score of the skills reflected in the middle. rank (P)i ) = 0.6 small examples (i.e., examples where the score output from the FC layer is reversed relative to the score that should have been).
[0045] When learning to judge skill levels, the machine learning device first starts with dynamic image #1 (i.e., dynamic image P). i In the segment S1, S2, and S3, images are selected frame by frame. In the region of interest generation unit, the dynamic image P is calculated. i The score is the score. att (P) j ) and Score rank (P) j ).exist Figure 3 In the example shown, Score att (P) i =0.8, Score rank (P) i =0.6.
[0046] Next, the machine learning device will start from dynamic image #2 (i.e., dynamic image P). j The image is selected frame by frame from each of the segments S1, S2, and S3, and the region of interest generation unit calculates the region of interest with respect to the dynamic image P. i The score is the score. att (P) j ) and Score rank (P) j ).exist Figure 3 In the example shown, Score att (P) i =0.1, Score rank (P) i =0.3.
[0047] In this example, Score att (P) i =0.8 > Score att (P) i =0.1, Score rank (P) i =0.6>Score rank (P) i =0.3.
[0048] Non-patent document 1 (where Score is denoted by f) describes an example of a learning method that calculates the difference of a loss function based on the score and uses this difference to determine the quality of a skill reflected in a dynamic image.
[0049] 3. Implementation Method 1
[0050] exist Figure 1 and Figure 2 In the technique of Comparative Example 1 shown, in order to determine the dynamic image pair (dynamic image P) i P j The quality of a person's skill level as reflected in their actions is sometimes judged based on the area of focus, and therefore, the quality of skill level is judged based on the area outside the focus (i.e., the area of focus that is deviated from), which can lead to low learning stability.
[0051] On the other hand, Figure 3 and Figure 4 In Comparative Example 2, the region of interest for manual editing (i.e., embedding human insights) in the transfer learning is considered important in determining the skill level of human actions in dynamic image pairs. For example, in the case of transfer learning in a machine learning device that generates a learning model for evaluating the skill of drawing a picture, it can be considered that manual editing is rarely performed on parts of the hand that is not reflected while moving the pen. Therefore, it can be considered that the parts edited in transfer learning are appropriate in determining the skill level of human actions in dynamic image pairs.
[0052] In Implementation 1, in a machine learning device that generates a learning model for determining the skill level of a task reflected in a pair of dynamic images, dynamic image data of parts that have been edited based on transfer learning (i.e., parts that the user focuses on from the dynamic image) are preferentially selected from the pair of dynamic images as attention data. Learning is performed based on the selected attention data, thereby improving the accuracy of skill level determination and further improving learning stability.
[0053] The data to be considered includes, for example, image frames in which the user edited (e.g., corrected, added, deleted, etc.) the animated image data during transfer learning; image data within a predetermined time range containing the edited image frames during transfer learning (i.e., image data from the X1-th image frame to the X2-th image frame, where X1 and X2 are predetermined positive integers); or data where the similarity of intermediate features contained in the animated image is greater than or equal to a predetermined value. For example, the data prioritization unit described later evaluates the weight W of the animated image data portion (image frame) edited by the user, sets the weight of the animated image data portion (image frame) that has not been edited by the user to 1, performs roulette selection, and thereby performs processing to increase the selection probability of the animated image data portion edited by the user (e.g., processing to make the weight W greater than 1).
[0054] Figure 5This is a functional block diagram that schematically represents the structure of the machine learning device 1 according to Embodiment 1. The machine learning device 1 is an apparatus for learning a learning model for inferring the skill level of an action subject in a moving image. The machine learning device 1 includes: a data priority selection unit 101, which selects multiple pairs of moving images (i.e., multiple moving image data) from a learning moving image dataset stored in a moving image dataset storage unit 110, and selects image frames for judging the quality of skill levels from each of the selected multiple pairs of moving images; a region of interest generation unit 103, which generates a region of interest in the image frame for judging the quality of skill levels; a quality judgment unit 106, which uses the learning model to judge the quality of the skill level in the region of interest for each pair of moving images; and a model learning unit 102, which stores the learning model and updates the learning model based on the judgment result of the quality of skill levels.
[0055] The machine learning device 1 includes a region of interest editing unit 105 that allows users to edit dynamic images while observing the display screen 111, a region of interest storage unit 106 that stores the edited dynamic images, and a region of interest generation unit 103 that generates regions of interest.
[0056] The data priority selection unit 101 selects image frames for judging skill level using selection probabilities determined based on user editing of image frames constituting each dynamic image pair. The data priority selection unit 101 obtains information from the attention area storage unit 104 indicating which image frame in the dynamic image has been edited by the user. For example, the data priority selection unit 101 sets the weight of edited image frames to W (a value greater than 1) and the weight of unedited image frames to 1, calculating the probability of selecting an image frame, i.e., the selection probability, for each image frame.
[0057] Weight W is an evaluation metric that takes into account factors such as the user's editing time and the consistency between the edited data and the heatmap. For example, increasing the weight W makes it easier to be selected.
[0058] The data prioritization unit 101 uses selection probability to select the image frame of the segment obtained in Comparative Example 1. The weight of the user-edited data i can be, for example, based on the editing time t. e Time t e The maximum value of max(t) e Interest region A generated by interest region generation unit 103 i With edited area of interest E i The difference, the size of the edited region of interest relative to the image area S iThe edited user's attribute r is calculated as shown in equation (1) below. In this case, the more gaps in the edited attention area and the narrower the area, the higher the probability of it being selected.
[0059] W=(t) e / max(t e ))+((E i -A i ) / A i )+S / s i +r (1)
[0060] As shown in equation (1), the data priority selection unit 101 can make the selection probability of the image frame that exists in each dynamic image pair higher than the selection probability of the image frame that does not exist in the dynamic image pair.
[0061] Furthermore, when there is user editing in the image frames that constitute each dynamic image pair, the longer the time range of user editing, the more the data priority selection unit 101 can improve the selection probability of the image frame.
[0062] Furthermore, when there is user editing in the image frames that constitute each dynamic image pair, the longer the time required for user editing, the more the data priority selection unit 101 can improve the selection probability of the image frame.
[0063] Furthermore, when there is user editing in the image frames that constitute each dynamic image pair, the greater the difference between the region of interest before and after editing, the more the data priority selection unit 101 can improve the selection probability of the image frame.
[0064] Furthermore, when there is user editing in the image frames that constitute each dynamic image pair, the smaller the area of the region of interest, the more the data priority selection unit 101 can improve the selection probability of the image frame.
[0065] The model learning unit 102 inputs dynamic image data selected by the data priority selection unit 101 into the convolutional neural network (CNN) and performs feature extraction.
[0066] The attention region generation unit 103 generates attention regions using an architecture with a class activation mapping (CAM) structure similar to an attention branch network in the intermediate branches, and saves the generation results in the attention region storage unit 104.
[0067] The model learning unit 102 extracts features related to the region of interest by masking the features of the CNN in the region of interest generated by the region of interest generation unit 103. Figure 9In this process, the data priority selection unit 101 selects a pair of animated images (i.e., animated image P, which is animated image #1) from the dataset storage unit 110. i And as dynamic image #2, dynamic image P j At this time, the dynamic image P j The skill level is higher than that of dynamic images P i Excellent skill level (also described as "P") i <P j In this case, it should be determined as a dynamic image P. j The score (P) of the skills reflected in the text. j ) Compared to dynamic images P i The score (P) of the skills reflected in the text. i )big.
[0068] exist Figure 9 The image P shown is a dynamic image of the region of interest that the region of interest generation unit 103 focuses on. j The score of the skills reflected in the middle. att (P) j =0.1 compared to dynamic image P i The score of the skills reflected in the middle. att (P) i Examples of small values (i.e., scores for the region of interest that are inverted relative to their expected scores) and dynamic images of the output from the FC layer, P. j The score of the skills reflected in the middle. rank (P) j =0.3 compared to dynamic image P i The score of the skills reflected in the middle. rank (P) i ) = 0.6 small examples (i.e., examples where the score output from the FC layer is reversed relative to the score that should have been).
[0069] The quality assessment unit 106 transforms the feature extraction results of the region of interest in the fully connected layer (FC layer) to extract quality assessment information. The quality assessment information is an evaluation result indicating which action has a higher skill level between the actions reflected in dynamic image #1 and the actions reflected in dynamic image #2.
[0070] The performance evaluation unit 106 calculates the difference between the performance evaluation result generated by the region of interest generation unit 103 and the performance evaluation result determined by the performance evaluation unit 106, based on the positive solution data related to the region of interest or performance evaluation provided by the user in advance. Based on the calculated loss, the performance evaluation unit 106 updates the parameters of the CNN in the model learning unit 102 or the CAM in the region of interest generation unit 103, as well as its own fully connected (FC) layer, through backpropagation. The performance evaluation unit 106 pre-checks whether the learning convergence condition is met; if the condition is met, the learning ends; otherwise, it repeatedly performs learning from multiple data selections.
[0071] The attention area editing unit 105 retrieves information about the attention area from the attention area storage unit 104 and visualizes it for the user. The attention area editing unit 105 accepts user input and deletes or adds attention areas. The attention area editing unit 105 saves the new attention areas edited by the user as learning data in the attention area storage unit 104.
[0072] Figure 6 This is a diagram illustrating an example of the hardware structure of the machine learning device 1 according to Embodiment 1. The machine learning device 1 of Embodiment 1 is an apparatus that performs a learning process to generate a learning model by performing machine learning. Furthermore, Figure 6 The machine learning device 1 functions as a skill determination device, which uses the learning model to deduce the skill level of an actor's actions in an input dynamic image. The machine learning device 1 is an apparatus capable of implementing the machine learning method of embodiment 1. The machine learning device 1 is, for example, a computer, but it can also be a computer system utilizing cloud computing with computer networks. Furthermore, in Figure 6 The example shown is a machine learning device that generates a learning model and a skill determination device that determines the skill level of an action subject in a motion image of an object, both located in the same computer, but they can also be located in different computers.
[0073] The machine learning device 1 has a processor 3, such as a CPU (Central Processing Unit), and a storage device 2. The storage device 2 is composed of semiconductor memory such as RAM (Random Access Memory), a hard disk drive (HDD), or a solid-state drive (SSD). The machine learning device 1 may also have a communication device for communicating with external devices. The machine learning device 1 is connected to input devices 4, such as a mouse and keyboard, and a display device 5 with a display. Additionally, the machine learning device 1 may also have a communication device for communicating with other devices.
[0074] The functions of the machine learning device 1 are implemented by processing circuitry. This processing circuitry can be, for example, dedicated hardware. Alternatively, it can be a processor 3 that executes a program stored in the storage device 2 (e.g., the machine learning program of the implementation). The processor 3 can also be a processing unit, a computing unit, a microprocessor, a microcomputer, or a DSP (Digital Signal Processor).
[0075] When the processing circuit is dedicated hardware, such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), the processing circuit is used.
[0076] When the processing circuit is processor 3, the machine learning method is executed by software, firmware, or a combination of software and firmware. The software and firmware are described as programs and stored in storage device 2. Processor 3 can implement the machine learning method of embodiment 1 by reading and executing the programs stored in storage device 2.
[0077] Furthermore, the machine learning device 1 can also implement some functions through dedicated hardware and others through software or firmware. In this way, the processing circuitry can implement the aforementioned functions through hardware, software, firmware, or any combination thereof.
[0078] The user uses a mouse or keyboard to register a dataset (dynamic image dataset) for learning via processor 3. Processor 3 reads the machine learning program stored in storage device 2 to perform learning or inference. Based on the learning dataset input by the user, processor 3 performs machine learning and inference, and stores the results as learning results in the storage device.
[0079] Processor 3 retrieves multiple data points from the dynamic image dataset. In processor 3, the data prioritization unit 101 selects the portions of the multiple data points that will be compared. In processor 3, the model learning unit 102, the merit judgment unit 106, and the region of interest generation unit 103 compare the selected data to evaluate their merits and demerits, and perform learning to obtain the evaluation value pre-assigned by the dataset. The generation result of the region of interest and the created model are recorded in the learning result. The generation result of the region of interest is displayed on the display device 5, and the user accepts the generation result and annotates the data in the input device 4. The annotation result is registered in processor 3 in the dataset as information for new machine learning through the region of interest editing unit 105.
[0080] Figure 7 (A) represents Figure 3 and Figure 4 The data selection unit of Comparative Example 2 shown is a graph of selected segments S1, S2, and S3 of a dynamic image from a dynamic image. Figure 7 (B) is a diagram showing the portion (slanted section) selected by the data priority selection unit 101 from each segment of the dynamic image pair.
[0081] Figure 8 This is an explanatory diagram showing the effect of the machine learning device 1 according to Embodiment 1. (As shown...) Figure 8 As shown, the data priority selection unit 101, by appropriately selecting image frames as paired data, can achieve a comparison score for handling subtle movements and a matching of the regions of interest, similar to a dynamic image displaying a skill. This allows for improved accuracy in skill assessment and provides appropriate explanations related to skill level. Furthermore, by comparing user-edited image frames with each other, it is easy to evaluate the skill level between the image frames desired by the user, thus ensuring stable learning. Additionally, by comparing user-edited and unedited image frames, the learning of regions of interest is also considered for unedited areas; therefore, the same regions of interest are generated even if the user does not edit the unedited areas.
[0082] Figure 9 This is an explanatory diagram illustrating the operation of the machine learning device 1 according to Embodiment 1. Additionally, Figure 10 This is a flowchart illustrating the learning operations of the machine learning device 1 in Embodiment 1. First, the data priority selection unit 101 selects multiple pairs of animated images (P as animated image #1). i And P as dynamic image #2 j (Step S101). Next, the data priority selection unit 101 obtains the edited data list from the interest area storage unit 104, which indicates which image frame in the dynamic images constituting the dynamic image pair has been edited by the user (step S102).
[0083] Next, the data priority selection unit 101 assigns a weight of W to the edited image frames and a weight of 1 to the unedited image frames to calculate the selection probability, which represents the probability of selecting an image frame. Furthermore, the weight W can be varied based on evaluation metrics such as the user's editing time or the consistency between the edited data and the heatmap. The data priority selection unit 101 then selects image frames from segments divided as in a conventional temporal segmentation network using the calculated selection probability (step S103).
[0084] The model learning unit 102 inputs the data selected by the data priority selection unit into the CNN to extract features from the image (step S104).
[0085] The region of interest generation unit 103 generates regions of interest through an architecture with a CAM structure in the intermediate branch, and saves the generation result to the region of interest storage unit 104 (step S105).
[0086] The model learning unit 102 extracts features related to the region of interest (step S106).
[0087] The quality assessment unit 106 extracts quality assessment information (step S107). The quality assessment information is an evaluation result indicating which action has a higher skill level between the actions shown in motion image #1 and the actions shown in motion image #2.
[0088] The merits / dismerits determination unit 106 calculates the difference between the merits / dismers determination results generated by the region of interest generation unit 103 and those determined by the merits / dismers determination unit 106, based on the positive solution data related to the region of interest or merits / dismers determination provided by the user in advance. Based on the calculated loss, the merits / dismers determination unit 106 updates the parameters of the CNN of the model learning unit 102 or the CAM of the region of interest generation unit 103, and its own FC layer through backpropagation (step S108).
[0089] The quality determination unit 106 pre-checks whether the learning convergence condition is met. If the condition is met, the learning ends; otherwise, the learning is repeated from multiple data selections (step S109).
[0090] Figure 11 This is a flowchart illustrating the annotations of the machine learning device 1 in Embodiment 1. The region of interest editing unit 105 retrieves information about the region of interest from the region of interest storage unit 104 and visualizes it to the user (prompt display screen 111) (step S151). Next, the region of interest editing unit 105 accepts user input and edits the region of interest (step S152). The region of interest editing unit 105 saves the newly edited region of interest as learning data to the region of interest storage unit 104 (step S153).
[0091] As explained above, according to Embodiment 1, the learning behavior during the learning of a learning model for inferring the skill level of the action subject in a moving image can be stabilized.
[0092] Implementation Method 2 (4)
[0093] Figure 12 This is a functional block diagram that schematically represents the structure of the machine learning device 1a according to Embodiment 2. Figure 12 In the middle, to and Figure 5 The elements shown are the same or corresponding element labels and Figure 5The same labels are shown. The difference between the machine learning device 1a of Embodiment 2 and the machine learning device 1 of Embodiment 1 is that it has a data priority selection unit (also called "data successive selection unit") 101a that selects consecutive image frames and a time series feature extraction unit 121 that extracts features of consecutive image frames.
[0094] In the machine learning apparatus 1a of Embodiment 2, the data priority selection unit 101a determines the selection probability of an image frame based on the temporal characteristics of multiple consecutive image frames constituting each dynamic image pair.
[0095] In the machine learning apparatus 1a of Embodiment 2, multiple consecutive image frames are selected to calculate the probability (hit rate) of selecting an appropriate dynamic image pair. In Comparative Example 1, the dynamic image data constituting the dynamic image pair is segmented to randomly select comparison target image frames. However, in the machine learning apparatus 1a of Embodiment 2, the data priority selection unit 101a selects multiple consecutive image frames. Furthermore, the machine learning apparatus 1a also includes a time-series feature extraction unit 121 to process the features of consecutive image frames. Because the machine learning apparatus 1a selects multiple consecutive image frames, the probability that any image frame contains an image frame with an edited region of interest is higher compared to the case of randomly selecting an image frame with an edited region of interest.
[0096] As a method for selecting multiple consecutive image frames, the following methods exist. In the first method, by specifying random locations, a predetermined number of consecutive data frames are selected, thus achieving a higher hit rate for the data being edited by the user compared to randomly selecting frame by frame. Table 1 below shows an example of a simulation of the hit rate.
[0097] [Table 1]
[0098]
[0099] In the machine learning apparatus 1a of embodiment 2, the surrounding data of user-edited data is highly likely to contain skills. Therefore, selection is made based on probability distribution, with the data of user-edited data at that moment as the center. By pre-preparing and selecting the variances of multiple normal distributions, features of long time series and features of short time series are selected.
[0100] Figure 13 This is a flowchart illustrating the learning actions of the machine learning device 1a in Implementation 2. Figure 13 The processing of steps S201, S203~S205, and S207~S209 in the process and Figure 10 The processing of steps S101, S104~S106, and S107~S109 is the same.
[0101] In Embodiment 2, the data priority selection unit 101b, which functions as a data continuity selection unit, selects data from multiple consecutive image frames (step S202). In this way, by selecting data that is continuous in the time direction, the edited data of the user-annotated area of interest is easily included in the judgment of quality.
[0102] In addition, in Embodiment 2, the time series feature extraction unit 121 performs convolution in the time direction on the data of multiple consecutive image frames to extract features in the time direction from the dynamic image (step S206).
[0103] As explained above, according to Embodiment 2, continuous subtle movements, such as skill motion images, can be captured. At the same time, the hit rate of the region of interest is higher compared to the case of randomly selecting moments. Therefore, the learning behavior during the learning of the learning model can be stabilized.
[0104] Apart from the above, Embodiment 2 is the same as Embodiment 1. Furthermore, the data priority selection unit 101b and the time series feature extraction unit 121 in Embodiment 2 can also be applied to Embodiment 1.
[0105] Implementation Method 3 (5)
[0106] Figure 14 This is a functional block diagram that schematically represents the structure of the machine learning device 1b in embodiment 3. Figure 14 In the middle, to and Figure 5 The elements shown are the same or corresponding element labels and Figure 5 The same labels are shown. The difference between the machine learning device 1b of Embodiment 3 and the machine learning device 1 of Embodiment 1 is that it has a region of interest comparison unit 131 that calculates the similarity between image frames edited by the user, and a data priority selection unit 101b that selects data based on similarity.
[0107] In Embodiment 2, a unit is included to improve the hit rate of image frames edited by the user. However, since the edited content is not considered, sometimes dynamic image pairs with different content are selected. Therefore, the machine learning device 1b in Embodiment 3, by providing a region of interest comparison unit 131, performs a comparison of edited content in addition to comparing regions of interest. The data priority selection unit 101b adjusts the selection probability to prioritize dynamic image pairs with similar edited content. As a result, similar regions of interest can be compared, and therefore, an improvement in the accuracy of skill level evaluation can be expected.
[0108] In Example 1 of Implementation Method 3, firstly, the region of interest comparison unit 131 calculates the similarity between image frames that have been edited by the user. The higher the similarity of the image frames that have been edited by the user, the more the data priority selection unit 101b prioritizes dynamic image pairs (increasing the selection probability) and selects dynamic image data. However, if the dynamic image pair contains data that has not been edited by the user, the selection probability is reduced by assigning a lower value (e.g., "0.01") as the similarity value. The model learning unit 102 learns for the selected dynamic image pairs, and the region of interest generation unit 103 generates regions of interest. If the selected pair is not a pair edited by the user, the region of interest comparison unit 131 calculates the similarity with other user-edited image frames or data selected as pairs.
[0109] In Example 2 of Implementation 3, firstly, the region of interest comparison unit 131 calculates the similarity between image frames that have been edited by the user. Next, the region of interest comparison unit 131 preferentially selects dissimilar pairs with a certain probability. For example, the dissimilarity is calculated using the formula "(dissimilarity) = 1.0 - (similarity)", and the selection probability is determined accordingly. This is because by selecting dissimilar animated image pairs, there is a certain probability that animated image pairs that have not been edited by the user will be included in each other, making it easier to generate regions of interest that the user has not assumed, which is related to the discovery of new regions of interest. The subsequent processing is the same as that in Example 1 of Implementation 3.
[0110] In Example 3 of Implementation 3, firstly, the region of interest comparison unit 131 has difficulty maintaining the similarity of all pairs of image frames. Therefore, it forms clusters based on similarity and selects data based on the similarity between clusters. At this time, the region of interest comparison unit 131 takes the image frames that have been edited by the user as clusters (data count is 1) and calculates the similarity between clusters. In addition, it is assumed that all image frames that have not been edited by the user belong to the unedited cluster. Next, the region of interest comparison unit 131 prioritizes selecting clusters with high similarity. However, the similarity between the unedited cluster and other clusters is considered to be a predetermined low value. For example, the region of interest comparison unit 131 randomly selects data contained in the two selected clusters A and B as pairs. Here, the region of interest comparison unit 131 may not be random, but may select data that is dissimilar to the representative point or other data within the cluster. By using such a selection method, learning is performed with the representative point and the point away from the representative as input, and it is expected that new regions of interest can be discovered. Next, the model learning unit 102 learns from the selected pairs, and the region of interest generation unit 103 generates regions of interest. Next, if the selected pair is not the pair edited by the user, the focus area comparison unit 131 calculates the similarity between the data and the representative data of other clusters, ensuring that the data belongs to the cluster with the highest similarity. Then, the focus area comparison unit 131 updates the representative data to the data with the highest similarity within the cluster.
[0111] Figure 15 This is a flowchart illustrating the learning actions of the machine learning device 1b in Implementation 3. Figure 15 The processing of steps S301, S305~S310 in the process and Figure 10 The processing of steps S101, S104~S109 is the same.
[0112] The attention region comparison unit 131 selects data from multiple dynamic image pairs (step S301). The attention region comparison unit 131 selects attention region images (Attention Maps) from the data of multiple dynamic image pairs (step S302). The attention region comparison unit 131 calculates the similarity between the attention region images (step S303). As a method for calculating similarity, the IoU (Intersection over Union) can also be used to calculate the degree of overlap between the attention region images. The attention region comparison unit 131 calculates the similarity of dynamic image #1 from a certain time t. k At a certain time t k+1 With dynamic image #2 from a certain moment s l At a certain moment s l+1 The sum of similarity values. The region comparison unit 131 performs this processing on all intervals, normalizing them so that the sum is 1. The data priority selection unit 101b determines which interval time t to select based on a random number. k ~t k+1 Or s l ~s l+1 As a range for the data, the data priority selection unit 101b can also calculate the range by weighting the similarity, so that the image edited by the user can be easily selected. In addition, the above-mentioned similarity calculation is performed on all combinations of animated image #1 and animated image #2, but one of the animated images can also be randomly selected based on its similarity to the other animated image.
[0113] As explained above, according to Embodiment 3, continuous subtle movements, such as skill motion images, can be captured. At the same time, the hit rate of the region of interest is higher compared to the case of randomly selecting moments. Therefore, the learning behavior during the learning of the learning model can be stabilized.
[0114] Apart from the above, Embodiment 3 is the same as Embodiment 1. Furthermore, the data priority selection unit 101b in Embodiment 3 can also be applied to Embodiment 1 or 2.
[0115] Implementation Method 4 (6)
[0116] Figure 16 This is a functional block diagram that schematically represents the structure of the machine learning device 1c according to embodiment 4. Figure 16In the middle, to and Figure 5 The elements shown are the same or corresponding element labels and Figure 5 The same labels are shown. The machine learning device 1c of embodiment 4 and... Figure 5 The difference in the machine learning device 1 of Embodiment 1 shown is that it has a motion extraction unit 141 and a motion comparison unit 142.
[0117] In embodiments 1 to 3, the motion of the subject in the moving image was not fully considered. In order to learn a learning model for determining skill level, it is important to determine whether the subject is performing the same motion as the excellent motion. The machine learning device 1c in embodiment 4 includes a motion extraction unit 141 and a motion comparison unit 142, and preferentially selects data with similar motion as moving image pairs.
[0118] According to the machine learning device 1c of embodiment 4, it is possible to compare data with similar movements, so it is easy to select the same skill level as the evaluation object, and it is expected to improve the accuracy of skill level evaluation.
[0119] As a first processing example for determining the similarity of actions, a hand pose tracking process can be considered. The first processing example can be executed according to the steps of processes 11 to 15 below.
[0120] (Process 11) Extract the image frames at time t (t=0, Δt, 2Δt, 3Δt, ..., NΔt) for each dynamic image. Here, N is a positive integer.
[0121] (Process 12) Calculate the motion vector (i.e., Flow) based on the image frames from time t = mΔt to time t = (m+1)Δt. Here, m = 0, 1, ..., N-1.
[0122] (Process 13) For all frames (N) of the moving image X x (N) and all frames of the dynamic image Y (N y For each pair of ( ), calculate the motion vector (Δx). l Δy l The cosine distance of (N). x ×N y ( ) cosine distances.
[0123] (Process 14) Repeat (Process 13) on all dynamic image pairs.
[0124] (Process 15) Select dynamic image pairs in a manner that prioritizes those with higher similarity as determined in (Process 13).
[0125] As a second processing example for determining the similarity of actions, we consider using clustering or other methods to reduce computational complexity. The second processing example can be executed according to the following steps (process 21) to (process 26).
[0126] (Process 21) Perform the same processing as (Process 11) of the first processing example.
[0127] (Process 22) Perform the same processing as (Process 12) of the first processing example.
[0128] (Process 23) Only calculate the similarity between adjacent image frames within the dynamic image X, and perform hierarchical clustering. The number of clusters is determined by user-defined parameters.
[0129] (Process 24) In (Process 23), the similarity between the data contained in the cluster is calculated, and the data that is on average similar to all the data in the cluster is used as the representative data.
[0130] (Process 25) Calculate the similarity between the clusters (Cx) of the dynamic image X and the clusters (Cy) of the dynamic image Y generated in (Process 23) and (Process 24).
[0131] (Process 26) Clusters are selected using the inter-cluster similarity calculated in (Process 25), and data within a cluster is randomly selected and pairs are determined. After clusters are determined, the same processing as in Implementation 3 can be applied.
[0132] In Implementation 4, user editing is not used, but motion vectors can be weighted by user editing, or data pairs can be selected by using an index that combines the similarity of Implementation 3 obtained by user editing with the similarity based on motion vectors.
[0133] Alternatively, the Dynamic Time Warping (DTW) equidistance index can be used to determine the similarity of a specific range of dynamic images.
[0134] A method to average the motion vectors of a portion of a dynamic image X to obtain the overall motion vector of one frame; selecting the dynamic image X from t=m min Δt to t=m max The interval Δt is divided into segments, and the DTW distance between each segment is calculated for clustering. Selection probabilities are set such that data belonging to the same cluster are easily selected as pairs. For example, data within the same cluster are more likely to be selected as pairs. In this case, pairs of data from different clusters are occasionally selected.
[0135] Figure 17 This is a flowchart illustrating the learning actions of the machine learning device 1c in embodiment 4. Figure 17The processing of steps S401, S405~S410 in the process and Figure 10 The processing of steps S101, S104~S109 is the same.
[0136] The motion extraction unit 141 selects data from multiple pairs of dynamic images (step S401). The motion comparison unit 142 extracts motion from the data of multiple pairs of dynamic images using methods such as optical flow (step S402). The motion extraction unit 141 can also extract the direction of movement of regions obtained by dividing the image into blocks and maintain it as feature vectors. The motion comparison unit 142 calculates similarity using the cosine distance, etc., of the extracted motion feature vectors (step S403). The motion comparison unit 142 calculates the similarity of dynamic image #1 from a certain time t. k At a certain time t k+1 With dynamic image #2 from a certain moment s l At a certain moment s l+1 The sum of similarity values. The motion comparison unit 142 performs this processing on all intervals, normalizing them so that the sum is 1. The data priority selection unit 101c determines which interval time t to select based on a random number. k ~t k+1 Or s l ~s l+1 The interval of the data (step S404).
[0137] As explained above, according to Embodiment 4, motion is extracted from the data of multiple dynamic image pairs, and the learning model is learned using this motion, thus enabling the learning behavior to be stable.
[0138] Except as described above, Embodiment 4 is the same as Embodiment 1. Furthermore, the motion extraction unit 141 and motion comparison unit 142 in Embodiment 4 can be applied to any of Embodiments 1-3.
[0139] Implementation Method 5 (7)
[0140] Figure 18 This is a functional block diagram that schematically represents the structure of the machine learning device 1d in embodiment 5. Figure 18 In the middle, to and Figure 5 The elements shown are the same or corresponding element labels and Figure 5 The same labels are shown. The machine learning device 1d of Embodiment 5 and... Figure 5 The difference in the machine learning apparatus 1 of Embodiment 1 shown is that it has a foreground extraction unit 151 and a region of interest comparison unit 152.
[0141] In embodiments 1 to 4, examples of learning actions from the entire dynamic image were described. However, when learning a learning model to deduce the skill level of the action subject, the result of analyzing the background of the dynamic image becomes noise, sometimes reducing the accuracy of skill level determination. Therefore, the machine learning apparatus 1d in embodiment 5 includes: a foreground extraction unit 151, which extracts the foreground from dynamic image pairs selected by the dynamic image dataset storage unit 110; and a region of interest comparison unit 152, which uses the foreground as a region of interest by masking the region other than the foreground, i.e., the background, and calculates the similarity of the regions of interest.
[0142] Thus, since areas unrelated to skill level are masked, selecting dynamic image pairs that are more directly correlated with skill level can lead to improved accuracy in skill level assessment and greater descriptiveness of the assessment.
[0143] Figure 19 This is a flowchart illustrating the learning actions of the machine learning device 1d in Implementation 5. Figure 19 The processing of steps S501, S507~S512 in the process and Figure 10 The processing of steps S101, S104~S109 is the same.
[0144] First, the foreground extraction unit 151 selects multiple data from the dynamic image dataset storage unit 110 and extracts the foreground (steps S501, S502). Foreground extraction can be performed, for example, by using a region that has not changed between past image frames and the current image frame as the background.
[0145] Next, the foreground extraction unit 151 obtains the region of interest image from the region of interest storage unit 104 (step S503), and performs masking processing on the foreground and region of interest images (step S504). The region of interest comparison unit 152 calculates the similarity between the masked region of interest images (step S505). The region of interest comparison unit 152 calculates the similarity of the dynamic image #1 from a certain time t. k At a certain time t k+1 With dynamic image #2 from a certain moment s l At a certain moment s l+1 The sum of similarities. The region comparison unit 152 performs this process on all intervals, normalizing them so that the sum is 1.
[0146] The data priority selection unit 101d determines which time interval t to select based on a random number. k ~t k+1 Or s l ~s l+1 As a range for data.
[0147] As explained above, according to Embodiment 5, motion is extracted from the data of multiple dynamic image pairs, and the learning model is learned using this motion, thus enabling the learning behavior to be stable.
[0148] Except as described above, Embodiment 5 is the same as Embodiment 1. Furthermore, the foreground extraction unit 151 and the region of interest comparison unit 152 in Embodiment 5 can also be applied to any of Embodiments 1-4.
[0149] Label Explanation
[0150] 1, 1a~1d: Machine learning device; 2: Storage device; 3: Processor; 4: Input device; 5: Display device; 101, 101a~101d: Data priority selection unit; 102: Model learning unit; 103: Region of interest generation unit; 104: Region of interest storage unit; 105: Region of interest editing unit; 106: Quality judgment unit; 110: Dynamic image dataset storage unit; 111; 112: Display example; 121: Time series feature extraction unit; 131: Region of interest comparison unit; 141: Motion extraction unit; 142: Motion comparison unit; 151: Foreground extraction unit.
Claims
1. A machine learning device that performs learning of a learning model for deriving a skill level of an action of a moving subject in a dynamic image, characterized by, The machine learning device has: a data priority selection section that selects a plurality of dynamic image pairs from among learning-use dynamic image data sets, selects, from each of the dynamic image pairs that constitute the selected plurality of dynamic image pairs, an image frame used to determine the merits and demerits of the skill level; an attention region generation section that generates, in the image frame, an attention region used to determine the merits and demerits of the skill level; a merit / demerit determination section that, for each dynamic image pair, uses the learning model to determine the merits and demerits of the skill level in the attention region; and a model learning section that stores the learning model, updates the learning model based on the determination results of the merits and demerits of the skill level, the data priority selection section selects the image frame used to determine the merits and demerits of the skill level using a selection probability determined based on one or more of a user edit in the image frame that constitutes each dynamic image pair, a feature in the time direction in a plurality of continuous image frames that constitute each dynamic image pair, and a similarity between the image frames that constitute each dynamic image pair.
2. The machine learning device according to claim 1, wherein the data priority selection section makes the selection probability of the image frame in which the user edit is present in the image frames that constitute each dynamic image pair higher than the selection probability of the image frame in which the user edit is not present.
3. The machine learning device according to claim 1 or 2, wherein in a case where the user edit is present in the image frames that constitute each dynamic image pair, the longer the time range in which the user edit is made is, the higher the data priority selection section makes the selection probability of the image frame.
4. The machine learning device according to any one of claims 1 to 3, wherein in a case where the user edit is present in the image frames that constitute each dynamic image pair, the longer the time required for the user edit is, the higher the data priority selection section makes the selection probability of the image frame.
5. The machine learning device according to any one of claims 1 to 4, wherein in a case where the user edit is present in the image frames that constitute each dynamic image pair, the larger the difference before and after the edit of the attention region is, the higher the data priority selection section makes the selection probability of the image frame.
6. The machine learning device according to any one of claims 1 to 5, wherein in a case where the user edit is present in the image frames that constitute each dynamic image pair, the smaller the area of the attention region is, the higher the data priority selection section makes the selection probability of the image frame.
7. The machine learning device according to any one of claims 1 to 6, wherein the data priority selection section determines the selection probability of the image frame based on a feature in the time direction in a plurality of continuous image frames that constitute each dynamic image pair.
8. The machine learning device according to any one of claims 1 to 7, wherein The machine learning device further has a region of interest comparison section that calculates a similarity of the region of interest with respect to each dynamic image pair, The higher the similarity between the regions of interest of the image frames that constitute each dynamic image pair, the higher the selection probability that the data priority selection section makes.
9. The machine learning device according to any one of claims 1 to 8, wherein The machine learning device further has: a motion extraction section that extracts a motion in the region of interest of each dynamic image pair; and a motion comparison section that calculates a similarity of the motion in the region of interest of each dynamic image pair, The higher the similarity of the motion, the higher the selection probability that the data priority selection section makes.
10. The machine learning device according to any one of claims 1 to 9, wherein The machine learning device further has: a foreground extraction section that extracts a foreground of each dynamic image pair; and a region of interest comparison section that calculates a similarity of the region of interest in the foreground, The higher the similarity of the region of interest in the foreground, the higher the selection probability that the data priority selection section makes.
11. The machine learning device according to any one of claims 1 to 10, wherein The action subject is a person or a mechanism that moves in conjunction with a motion of a body part of a person.
12. A skill determination device characterized by a dynamic image input having the learning model generated by the machine learning device according to any one of claims 1 to 11, the skill determination device determining a skill level of an action of an action subject in an object dynamic image by the learning model.
13. A machine learning method of performing learning of a learning model for inferring a skill level of an action subject in a dynamic image, characterized by, The machine learning method has the following steps: selecting a plurality of dynamic image pairs from a dynamic image dataset for learning, and selecting an image frame for determining a merit or demerit of the skill level from each dynamic image pair that constitutes the selected plurality of dynamic image pairs; generating a region of interest for determining the merit or demerit of the skill level in the image frame; determining the merit or demerit of the skill level in the region of interest using the learning model for each dynamic image pair; and storing the learning model, and updating the learning model based on a determination result of the merit or demerit of the skill level, In the step of selecting an image frame for determining a merit or demerit of the skill level, the image frame for determining the merit or demerit of the skill level is selected using a selection probability determined based on one or more of a user edit in the image frames that constitute each dynamic image pair, a feature in a time direction in a plurality of continuous image frames that constitute each dynamic image pair, and a similarity between the image frames that constitute each dynamic image pair.
14. A machine learning program that causes a computer to perform learning of a learning model for deriving a skill level of an action of a moving body in a dynamic image, characterized by, The machine learning program has the following steps: selecting a plurality of dynamic image pairs from a dynamic image dataset for learning, and selecting an image frame for determining a merit or demerit of the skill level from each dynamic image pair that constitutes the selected plurality of dynamic image pairs; generating a region of interest for determining the merit or demerit of the skill level in the image frame; determining the merit or demerit of the skill level in the region of interest using the learning model for each dynamic image pair; and storing the learning model, and updating the learning model based on a determination result of the merit or demerit of the skill level. For each dynamic image pair, the learning model is used to determine the proficiency in the skill in the attention region; And The learning model is stored, and the learning model is updated based on the determination result of the proficiency in the skill, In the step of selecting the image frame for determining the proficiency in the skill, the selection probability determined according to one or more of the user editing in the image frame constituting each dynamic image pair, the time direction feature in the continuous multiple image frames constituting each dynamic image pair, and the similarity between the image frames constituting each dynamic image pair is used to select the image frame for determining the proficiency in the skill.