Machine learning device, competence determination device, machine learning method and machine learning program
The machine learning device stabilizes competence level assessments in video-based actions by selecting frames based on user edits and similarity, enhancing learning stability and accuracy through a data preference selection and attention area generation process.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- MITSUBISHI ELECTRIC CORP
- Filing Date
- 2023-08-31
- Publication Date
- 2026-05-21
AI Technical Summary
Conventional machine learning methods for determining competence levels in video-based actions are unstable due to biased attention areas, leading to incorrect assessments and learning instability.
A machine learning device that selects video frames based on user edits, similarity, and time direction to generate attention areas for accurate competence level determination, using a data preference selection unit, attention area generation unit, and model learning unit to stabilize learning behavior.
Stabilizes learning behavior and improves accuracy in determining competence levels by focusing on user-edited frames and attention areas, ensuring consistent comparisons and explanations of competence.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
AREA OF TECHNOLOGY
[0001] The present disclosure relates to a machine learning device, a competence determination device, a machine learning method and a machine learning program. STATE OF THE ART
[0002] Pairwise Deep Ranking (PDR) has been proposed. This technology belongs to the competency assessment techniques that calculate a rating or score for a person's competency level (i.e., performance level) of an action and determine the relative quality of that competency level (see, e.g., Non-Patent Reference 1). REFERENCES ON THE STATE OF THE TECHNOLOGY - PATENT REFERENCE
[0003] Non-Patent Reference 1: Masayuki Takada and three others (Chubu University), “Attention Pairwise Ranking: Visual Explanations in Skill Assessment,” The 23rd Meeting on Image Recognition and Understanding. SUMMARY OF THE INVENTION TASK TO BE SOLVED BY THE INVENTION
[0004] However, in the conventional technology described above, the information provided only includes the superiority or inferiority of the skill and a video pair, and there are cases where the superiority or inferiority of the skill is determined based on locations other than those that should be considered for determining superiority or inferiority (i.e., biased attention areas in the video pair). For example, when assessing the action competence of drawing a picture, there are cases where the superiority or inferiority of the skill is determined not based on a video at the time a pen is moved, but on a video that captures the head of the person moving the pen. The conventional technology therefore has the problem that learning behavior can become unstable.
[0005] The aim of the present disclosure is to stabilize learning behavior when learning a learning model for deriving the competence level of an agent's action in a video. MEANS TO SOLVE THE PROBLEM
[0006] A machine learning device in the present disclosure is a device that learns a learning model to derive a competence level of an action of an acting subject in a video.The machine learning device comprises a data preference selection unit for selecting a multitude of video pairs from a video dataset for learning, and for selecting image frames used to determine superiority or inferiority of the competence level from each video pair comprising the multitude of selected video pairs; an attention area generation unit for generating an attention area used to determine the superiority or inferiority of the competence level in the image frames; a superiority / inferiority determination unit for performing a determination of superiority or inferiority of the competence level in the attention area using the learning model with respect to each video pair; and a model learning unit for storing the learning model and updating the learning model based on the result of the determination of superiority or inferiority of the competence level.The data preference selection unit selects the image frames used to determine the superiority or inferiority of the competence level by using a selection probability determined based on one or more user edits in the image frames that make up each video pair, a feature in a time direction in a multitude of sequential image frames that make up each video pair, and a similarity level between the image frames that make up each video pair.
[0007] A machine learning method in the present disclosure is a method of learning a learning model to derive a competence level of an action of an acting subject in a video.The machine learning process comprises a step of selecting a multitude of video pairs from a video dataset for learning and selecting image frames to be used to determine superiority or inferiority of the competence level from each video pair comprising the multitude of selected video pairs; a step of generating an attention area used to determine the superiority or inferiority of the competence level in the image frames; a step of performing a determination of superiority or inferiority of the competence level in the attention area using the learning model with respect to each video pair; and a step of storing the learning model and updating the learning model based on a result of the determination of superiority or inferiority of the competence level.In the step of selecting image frames to be used to determine the superiority or inferiority of the competence level, the image frames used to determine the superiority or inferiority of the competence level are selected by using a selection probability determined on the basis of one or more user edits in the image frames that make up each video pair, a feature in a time direction in a multitude of sequential image frames that make up each video pair, and a level of similarity between the image frames that make up each video pair. IMPACT OF THE INVENTION
[0008] According to the present disclosure, the learning behavior during the learning of the learning model for deriving the competence level of the action of the acting subject in a video can be stabilized. BRIEF DESCRIPTION OF THE DRAWINGS Fig.Figure 1 is an explanatory representation showing a conventional system (first comparative example) for carrying out a superiority-or-inferiority determination of competence. Fig. Figure 2 is a functional block diagram that schematically shows the configuration of the system (first comparative example). Fig. Figure 3 is an explanatory illustration showing the operation of a conventional machine learning device (second comparative example) to determine the competence level of an action performed by a person recorded in a video using transfer learning. Fig. Figure 4 is a function block diagram that schematically illustrates the configuration of the machine learning device (second comparative example) in Fig. 3 shows. Fig. Figure 5 is a functional block diagram which schematically shows the configuration of a machine learning device according to a first embodiment. Fig.Figure 6 is a representation showing an example of the hardware configuration of the machine learning device (or competency determination device) according to the first embodiment. Fig. 7A is a representation showing segments of videos selected from a video pair by a random data selection unit in the first comparison example, and Fig. Figure 7B is a representation showing parts selected from a video pair by a data preference selection unit in the first embodiment. Fig. Figure 8 is an explanatory illustration showing the effects achieved by the machine learning device according to the first embodiment. Fig. Figure 9 is an explanatory illustration showing the operation of the machine learning device according to the first embodiment. Fig.Figure 10 is a flowchart showing a learning operation performed by the machine learning device according to the first embodiment. Fig. Figure 11 is a flowchart showing an annotation performed by the machine learning device according to the first embodiment. Fig. Figure 12 is a functional block diagram which schematically shows the configuration of a machine learning device according to a second embodiment. Fig. Figure 13 is a flowchart showing a learning operation carried out by the machine learning device according to the second embodiment. Fig. Figure 14 is a functional block diagram which schematically shows the configuration of a machine learning device according to a third embodiment. Fig. Figure 15 is a flowchart showing a learning operation carried out by the machine learning device according to the third embodiment. Fig.Figure 16 is a functional block diagram which schematically shows the configuration of a machine learning device according to a fourth embodiment. Fig. Figure 17 is a flowchart showing a learning operation carried out by the machine learning device according to the fourth embodiment. Fig. Figure 18 is a functional block diagram which schematically shows the configuration of a machine learning device according to a fifth embodiment. Fig. Figure 19 is a flowchart showing a learning operation performed by the machine learning device according to the fifth embodiment. MODE FOR EXECUTING THE INVENTION
[0009] A machine learning device, a competency determination device, a machine learning method, and a machine learning program according to each embodiment are described below with reference to the drawings. The following embodiments are only examples, and it is possible to combine embodiments appropriately and to modify each embodiment appropriately.
[0010] The machine learning device according to each embodiment is a device that learns a learning model used by a derivation device (also called a "competence determination device") to derive the competence level (i.e., performance level) of an action performed by an acting subject recorded in a video. The machine learning device according to each embodiment is, for example, a computer as the information processing device. The acting subject recorded in the video is a person performing work (also called a "worker"). Furthermore, the acting subject recorded in the video may have a mechanism (e.g., a device such as a robotic arm or an endoscope) that performs work in conjunction with the person's movements.
[0011] The machine learning method according to each embodiment is a method that can be performed by the machine learning device. The machine learning method according to each embodiment is a method of learning a model to derive the competence level of the action of the acting subject in a video.
[0012] The machine learning program according to each embodiment is a software program that can be executed by a computer as the machine learning device. The machine learning program according to each embodiment is a program for learning a model in order to derive the competence level of the action of the acting subject, which was recorded in a video. (1) First comparative example
[0013] Fig. Figure 1 is an explanatory illustration showing a conventional system (first comparative example) for conducting the superiority-or-inferiority determination of competence. Fig. Figure 1 is a configuration proposed in the aforementioned non-patent reference 1, represented as System 210. Fig. Figure 2 is a functional block diagram that schematically shows the configuration of system 210 (first comparison example). The input to the system as the first comparison example is a video pair (i.e., two videos P). i and P j ), where P i a video demonstrating superiority over P j This system consists of a processing unit (processing) that divides a video into segments, a feature extractor that extracts a feature from a video, a superiority network that evaluates superiority actions, and an inferiority network that evaluates inferiority actions. Each of the superiority and inferiority networks comprises an attention branch and a ranking branch. Upon input of the video P iThe received output value is Score(p) i ), which is based on an output value of the superiority network and an output value of the inferiority network. A value of P is used as input to the video. j The received output value is Score(p) j ), which is based on an output value from the superiority network and an output value from the inferiority network. System 210 learns a magnitude relationship between the scores of these two videos. If the video P i the video P j is superior to and the size ratio between Score(P) i ) and Score(P jSince the value is inverted, the difference between the scores is fed into a loss function for learning. For example, the loss function could be Marginal Loss, which means that only differences greater than or equal to a fixed value are learned; SoftPlus, which considers a loss due to a difference less than or equal to a fixed value as a small loss; or something similar. (2) Second comparative example
[0014] Fig. Figure 3 is an explanatory illustration showing the operation of a conventional machine learning device (second comparative example) to determine the competence level of an action performed by a person recorded in a video using transfer learning. Fig. Figure 3 shows the configuration of a machine learning device 220 as proposed in non-patent reference 2. Fig. 4 is a function block diagram that shows the functions of a learning model in Fig. 3 shows.
[0015] Non-patent reference 2: Masahiro Mitsuhara and six others, “Embedding Human Knowledge into Deep Neural Network via Attention Map”, ar-Xiv:1905.03540, May 9, 2019.
[0016] In the second comparative example, transfer learning is performed, in which an attentional area generated by an attentional mechanism (a network that generates the attentional area) in relation to a video is corrected by a human (i.e., human knowledge is embedded in the learning model), and learning is carried out by using the corrected attentional area as the correct response data. Transfer learning is a form of human-in-the-loop (HITL) learning. For example, transfer learning creates a learning model that determines the competence level of a person's action in a video when interacting with a user.
[0017] In Fig.3 selects, for example, a data selection unit a video pair (i.e., a video P) i as video #1 and a video P j (as video #2) from a data storage unit. The competence level of video P j compared to the competence level of video P i superior (this relationship is also known as “P i < P j “ shown). In this case, the score should be Score(P j ) the one in the video P j The recorded competence is determined to be higher than the score. i ) the one in the video P i acquired competence.
[0018] Fig. Figure 3 shows an example where the score is... att (p j ) = 0.1 of the P in the video j The perceived competence regarding the attentional domain to which an attentional domain generation unit has drawn attention is lower than the score. att (p i ) = 0.8 of the P shown in the video irecorded competence (i.e., an example where the relationship between the scores regarding the attentional area is inverted from the originally natural score relationship) and an example where the score score- rank (P j ) = 0.3 of the P shown in the video j The recorded competence regarding the output of an FC layer is lower than the score. rank (P i ) = 0.6 of the P shown in the video i recorded competence (i.e., an example where the relationship between the score ratings issued by the FC layer is inverted from the original natural score relationship).
[0019] To learn how to determine the superiority or inferiority of the performance level, the machine learning device first selects a frame of an image from each of the segments S1, S2 and S3 of video #1 (i.e. video P). i ) and the attentional area generation unit calculates the score att (Pj ) and the score rank (P j ), which are the score ratings regarding the video P i it. Fig. Figure 3 shows an example where the score att (P i ) = 0.8 and Score rank (P i ) = 0.6 applies.
[0020] The machine learning device then selects a frame of an image from each of the segments S1, S2 and S3 of video #2 (i.e. video P). j ) and the attentional area generation unit calculates the score att (P j ) and the score rank (P j ), which are the score ratings regarding the video P i it. Fig. Figure 3 shows an example where the score att (P i ) = 0.1 and Score rank (P i ) = 0.3 applies.
[0021] In this example, the score is att ( Pi ) = 0.8 > Score att ( Pi ) = 0.1 and Score- rank ( Pi) = 0.6 > Score rank ( Pi ) = 0.3.
[0022] An example of a procedure for calculating the difference in the loss function from the score ratings and for learning the superiority-or-inferiority determination of the competence recorded in a video using the difference is described in the aforementioned non-patent reference 1 (where the score is represented by f). (3) First embodiment
[0023] At the in Fig. 1 and Fig. 2. Technology shown for determining the superiority or inferiority of the competence levels of the actions of the participants in the video pair (Videos P) i and P j) among the individuals included, there are cases where the superiority or inferiority of competence levels is determined based on locations other than those to which attention should be directed (i.e., biased attention areas), and there are cases where learning stability is low.
[0024] On the other hand, the manually processed attentional area (i.e., where human knowledge is embedded) can be used in transfer learning in the second comparative example. Fig. 3 and Fig.4. This can be considered an important part of the video pair for determining the competence level of a person's action. For example, if transfer learning is performed by a machine learning system that generates a learning model to assess the action competence of drawing a picture, the possibility of manual editing by a human being on a part where no hand moving the pen was recorded can be considered low. Therefore, the part where the editing is performed during transfer learning can be considered a part of the video pair suitable for determining the competence level of a person's action.
[0025] In the first embodiment, in the machine learning device that generates a learning model to determine the competence level of the work recorded in a video pair, video data of a part that has undergone processing by transfer learning (i.e., a part of the video to which the user directs their attention) are preferentially selected as attention data from the video pair, and the learning is carried out on the basis of the selected attention data, thereby increasing the accuracy in determining the superiority or inferiority of competence and further increasing the learning stability.
[0026] The attentional data includes, for example, image frames where a human has performed an edit (e.g., correction, addition, deletion, or similar) to video data during transfer learning; image data of a predetermined length containing image frames that were edited during transfer learning (i.e., image data from the X1st frame to the X2nd frame, where X1 and X2 are predetermined positive integers); data where the similarity of an intermediate feature contained in the video is greater than or equal to a predetermined value; or similar. For example, a data preference selection unit, described later, performs a process of increasing the selection probability of video data segments that have been edited by the user (e.g.,a process of setting a weight W to greater than 1), by evaluating the weight W of the video data parts (image frames) that have been edited by the user, setting the weight of the video data parts (image frames) that have not undergone video editing to 1, and making a roulette selection.
[0027] Fig.Figure 5 is a functional block diagram that schematically shows the configuration of a machine learning device 1 according to the first embodiment. The machine learning device 1 is a device that learns a learning model to derive the competence level of an action performed by the acting subject in a video. A machine learning device 1 comprises a data preference selection unit 101, which processes a plurality of video pairs (i.e.,a multitude of video data elements) from a video dataset for learning, which are stored in a video dataset storage unit 110, and selects image frames that are used to determine the superiority or inferiority of the competence level from each video pair that forms the multitude of selected video pairs, an attention area generation unit 103 that generates an attention area that is used to determine the superiority or inferiority of the competence level in each image frame, a superiority / inferiority determination unit 106 that performs a determination of the superiority or inferiority of the competence level in the attention area by using the learning model with respect to each video pair, and a model learning unit 102 that stores the learning model and updates the learning model based on the result of the determination of the superiority or inferiority of the competence level.
[0028] The machine learning device 1 comprises an attention area processing unit 105, which performs processing of videos according to the operations carried out by the user while viewing a display screen 111, an attention area storage unit 106, which stores the processed videos, and the attention area generation unit 103, which generates the attention area.
[0029] The Data Preference Selection Unit 101 selects the image frames used to determine the superiority or inferiority of the proficiency level by using selection probabilities determined based on user editing of image frames that comprise each video pair. The Data Preference Selection Unit 101 obtains information from the Attention Area Storage Unit 104 indicating which image frames in the videos have been edited by the user. For example, the Data Preference Selection Unit 101 sets the weight of the already edited image frames to W (a value greater than 1), sets the weight of the unedited image frames to 1, and calculates the selection probability as the probability that an image frame will be selected for each image frame.
[0030] The weight W, for example, is a rating index that takes into account the time spent by the user processing the data, the degree of similarity between the processed data and a heatmap or similar, and an image frame is more likely to be selected with increasing weight W.
[0031] The data preference selection unit 101 selects image frames from the segments obtained in the first comparison example using the selection probability. The weight of the user-processed data i can be calculated using expression (1) shown below, for example by taking into account the time t required for processing. e , the maximum value max(t e ) of time t e , the difference between an attention area A generated by the attention area generation unit 103 i and the processed attention area E i , the size s iThe edited attention area relative to an image area S and an attribute r of the user who performed the editing are used. In this case, the probability of being selected increases the greater the distance to the edited attention area and the narrower the area. W=(te / max(te))+((Ei−Ai) / Ai)+S / si+r
[0032] As shown in expression (1), the data preference selection unit 101 is able to adjust the selection probability of an image frame that has undergone user editing in the image frames that form each video pair such that it is higher than the selection probability of an image frame that has not undergone user editing.
[0033] The data preference selection unit 101 can further increase the selection probability of such image frames with increasing length of a time period of user processing, if the user processing has taken place in image frames that form each video pair.
[0034] The data preference selection unit 101 can also increase the selection probability of such image frames with increasing length of time used for user processing, if the user processing has taken place in image frames that form each video pair.
[0035] Furthermore, when user editing of image frames forming each video pair occurs, the data preference selection unit 101 can increase the selection probability of such image frames with the increase of the difference between the attention area after editing and the attention area before editing.
[0036] The data preference selection unit 101 can additionally increase the selection probability of such image frames with a decrease in the area of attention if the user processing has taken place in image frames that form each video pair.
[0037] The model learning unit 102 performs feature extraction by feeding the video data selected by the data preference selection unit 101 into a convolutional neural network (CNN).
[0038] The attention area generation unit 103 generates the attention area using an architecture with a branched class activation mapping structure (CAM), such as an attention branching network, and stores the result of the generation in the attention area storage unit 104.
[0039] Model learning unit 102 extracts a feature for the attentional domain by hiding or masking a feature of the CNN in the attentional domain that was generated by the attentional domain generation unit 103. Fig. 9 selects the data preference selection unit 101 a video pair (i.e., a video P) i as video #1 and a video P j as video #2) from data set storage unit 110. The competence level of the video is P j compared to the competence level of video P i superior (this relationship is also known as “P i < P j “ shown). In this case, the score should be Score(P j ) the one in the video P j The recorded competence is determined to be higher than the score. i ) the one in the video P i acquired competence.
[0040] Fig. 9 shows an example where the score is... att (P j ) = 0.1 of the P in the videoj The perceived competence regarding the attentional domain to which the attentional domain generation unit 103 drew attention is lower than the score. att (P i ) = 0.8 of the P shown in the video i recorded competence (i.e., an example where the relationship between the scores regarding the attentional area is inverted from the originally natural score relationship) and an example where the score score- rank (P j ) = 0.3 of the P shown in the video j The perceived competence regarding the output of the FC layer is lower than the score. rank (P i ) = 0.6 of the P shown in the video i recorded competence (i.e., an example where the relationship between the score ratings issued by the FC layer is inverted from the original natural score relationship).
[0041] Superiority / Inferiority Determination Unit 106 extracts superiority / inferiority information by transforming the result of the attentional area feature extraction in the fully connected layer (FC layer). The superiority / inferiority information is an assessment result indicating which of the competence levels of the action recorded in video #1 and the action recorded in video #2 is higher.The superiority / inferiority determination unit 106 receives the difference between the attentional area generated by the attentional area generator 103 or the superiority / inferiority determination result from the determination by the superiority / inferiority determination unit 106, based on a previously user-provided attentional area or correct response data regarding the superiority / inferiority determination. The superiority / inferiority determination unit 106 updates the CNN in the model learning unit 102 or the CAM in the attentional area generator 103 and a parameter of its own FC layer via backpropagation based on the calculated loss.The superiority / inferiority determination unit 106 checks whether a learning convergence condition has been previously met or not, and terminates the learning if the condition is met, or repeats the learning by selecting a variety of data elements if the condition is not met.
[0042] The attention area processing unit 105 retrieves information about the attention area from the attention area storage unit 104 and visualizes the attention area for the user. The attention area processing unit 105 performs deletions or additions to the attention area by receiving user input. The attention area processing unit 105 stores any new attention area obtained through user processing in the attention area storage unit 104 as training data.
[0043] Fig.Figure 6 is a diagram showing an example of the hardware configuration of the machine learning device 1 according to the first embodiment. The machine learning device 1 according to the first embodiment is a device that performs a learning process to generate a learning model through machine learning. Furthermore, the machine learning device 1 features in Fig. 6. A function as a competence determination device that derives the competence level of the action of the acting subject in an input video using the learning model. The machine learning device 1 is a device capable of performing a machine learning procedure according to the first embodiment. Although the machine learning device 1 is a computer, it can also be a computer system formed through cloud computing using a computer network. Fig.Figure 6 shows an example where the machine learning device that generates the learning model and the competence determination device that determines the competence level of the action of the acting subject in the video as an object are provided in the same computer, but the machine learning device and the competence determination device can also each be provided in different computers.
[0044] The machine learning device 1 comprises a processor 3, such as a CPU (Central Processing Unit), and a memory device 2. The memory device 2 consists of semiconductor memory such as RAM (Random Access Memory), a hard disk drive (HDD), a solid-state drive (SSD), or similar. The machine learning device 1 may include a communication device that communicates with external devices. An input device 4, such as a mouse, keyboard, or similar, and a display device 5 with a screen or monitor are connected to the machine learning device 1. Furthermore, the machine learning device 1 may include a communication device that communicates with other devices.
[0045] The functions of the machine learning device 1 are implemented by a processing circuit. This processing circuit can be, for example, specialized hardware. The processing circuit can be the processor 3, which executes a program stored in the memory device 2 (e.g., a machine learning program according to the present embodiment). The processor 3 can also be a processing device, an arithmetic device, a microprocessor, a microcomputer, or a DSP (Digital Signal Processor).
[0046] If the processing circuit is dedicated hardware, the processing circuit is, for example, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or similar.
[0047] If the processing circuit is processor 3, the machine learning process is carried out by software, firmware, or a combination of software and firmware. The software and firmware are described as programs and stored in memory device 2. Processor 3 is capable of performing the machine learning process according to the first embodiment by reading and executing the program stored in memory device 2.
[0048] It is also possible to implement part of the machine learning device 1 using dedicated hardware and another part using software or firmware. As described above, the processing circuit can implement the functions described above using hardware, software, firmware, or a combination of some of these means.
[0049] The user enters the learning data set (video data set) into Processor 3 using the mouse or keyboard. Processor 3 reads the machine learning program stored in Memory Unit 2 and performs the learning or inference. Based on the user-entered learning data set, Processor 3 performs the machine learning and inference and stores the result in Memory Unit 2 as the learning outcome.
[0050] Processor 3 extracts a multitude of data elements from the video dataset. Within Processor 3, the data preference selection unit 101 extracts portions of the data elements as comparison targets. Within Processor 3, the model learning unit 102, the superiority / inferiority determination unit 106, and the attention area generation unit 103 estimate superiority or inferiority by comparing the selected data, performing the learning process so that a previously provided estimation value can be obtained, and recording the result of the attention area generation and the generated model in the learning output. The display unit 5 shows the attention area generation result, and in response to the display, the user makes annotations to the data via the input unit 4.In processor 3, the attention area processing unit 105 enters the result of the annotation into the data set as information for new machine learning.
[0051] Fig. 7A is a representation showing the segments S1, S2, and S3 of the videos that were extracted from the video pair by the data selection unit in the second comparison example. Fig. 3 and Fig. 4 can be selected. Fig. Figure 7B is a representation showing parts (hatched parts) that are selected from the video pair segments by the data preference selection unit 101 in the first embodiment.
[0052] Fig. Figure 8 is an explanatory illustration showing the effects achieved by the machine learning device 1 according to the first embodiment. As in Fig.As shown in Figure 8, the data preference selection unit 101 appropriately selects image frames as data to be paired and is thus able to ensure the consistency of superiority-or-inferiority comparison points and the attentional area comparison in cases involving the handling of detailed movement, such as in a video demonstrating competence. Accordingly, effects of increased competence determination accuracy and obtaining an appropriate explanation of the competence level are to be expected. Furthermore, learning is stabilized, as the assessment of the competence level between user-planned image frames becomes more likely through comparison between the image frames processed by the user.Since learning, taking into account the attention area, is also carried out for unedited areas due to the comparison between a user-edited image frame and an unedited image frame, a similar attention area is also generated for unedited areas, even though the areas were not edited by the user.
[0053] Fig. Figure 9 is an explanatory illustration showing the operation of the machine learning device 1 according to the first embodiment. Fig. Figure 10 is a flowchart showing a learning operation performed by the machine learning device 1 according to the first embodiment. First, the data preference selection unit 101 selects a plurality of video pairs (i.e., P). i as video #1 and a video P jas video #2) (step S101). Then the data preference selection unit 101 obtains from the attention area storage unit 104 a list of edited data indicating which image frames in the videos forming the video pair have been edited by the user (step S102).
[0054] The data preference selection unit 101 then weights the image frames by setting the weight of a processed image frame to W and the weight of an unprocessed image frame to 1, and calculates the selection probability, which indicates the likelihood of each image frame being selected. Furthermore, the weight W can be varied based on a rating index that considers the duration of user processing, the degree of agreement between the processed data and the heatmap, or similar factors. Using the resulting selection probability, the data preference selection unit 101 selects image frames in the segments obtained through segmentation, similar to the conventional temporal segment network (step S103).
[0055] The model learning unit 102 performs image feature extraction by inputting the data selected by the data preference selection unit into the CNN (step S104).
[0056] The attention area generation unit 103 generates the attention area using an architecture with the branched CAM and stores the result of the generation in the attention area storage unit 104 (step S105).
[0057] Model learning unit 102 extracts the feature relating to the attention area (step S106).
[0058] The superiority / inferiority determination unit 106 extracts the information for the superiority-or-inferiority determination (step S107). The superiority-or-inferiority determination information is the assessment result, which indicates which of the competence levels of the action recorded in video #1 and the competence level of the action recorded in video #2 is higher.
[0059] The superiority / inferiority determination unit 106 receives the difference between the attentional area generated by the attentional area generator 103 or the superiority / inferiority determination result from the determination by the superiority / inferiority determination unit 106, based on the attentional area previously provided by the user or the correct response data regarding the superiority / inferiority determination. The superiority / inferiority determination unit 106 updates the CNN in the model learning unit 102 or the CAM in the attentional area generator 103 and the parameter of its own FC layer via backpropagation based on the calculated loss (step S108).
[0060] The superiority / inferiority determination unit 106 checks whether the learning convergence condition has been previously met or not, and terminates the learning if the condition is met, or repeats the learning by selecting a variety of data elements if the condition is not met (step S109).
[0061] Fig.Figure 11 is a flowchart showing the annotation performed by the machine learning device 1 according to the first embodiment. The attention area processing unit 105 obtains the information about the attention area from the attention area storage unit 104 and visualizes the attention area for the user (presents the display screen 111) (step S151). Then, the attention area processing unit 105 performs the processing of the attention area by receiving the user's input operations (step S152). The attention area processing unit 105 stores the new attention area obtained through the user's processing in the attention area storage unit 104 as training data (step S153).
[0062] As described above, according to the present disclosure, the learning behavior during the learning of the learning model can be stabilized to derive the competence level of the action of the acting subject in a video. (4) Second embodiment
[0063] Fig. Figure 12 is a functional block diagram which schematically shows the configuration of a machine learning device 1a according to a second embodiment. Fig. 12 is each element that is paired with an element in Fig. 5 is identical or corresponds to it, the same reference symbol as in Fig.5 assigned. The machine learning device 1a according to the second embodiment differs from the machine learning device 1 according to the first embodiment in that it has a data preference selection unit (also referred to as "sequential data pair selection unit") 101a, which selects sequential image frames, and a time series feature extraction unit 121, which extracts a feature from the sequential image frames.
[0064] The data preference selection unit 101a in the machine learning device 1a according to the second embodiment determines the selection probability of each image frame based on a feature in a time direction in a plurality of sequential image frames that form each video pair.
[0065] In the machine learning device 1a according to the second embodiment, a plurality of sequential image frames are selected to increase the probability of selecting a suitable video pair (hit rate). In the first comparison example, the video data forming a video pair is segmented into segments, and the image frames are randomly selected as the comparison targets, while the data preference selection unit 101a in the machine learning device 1a according to the second embodiment selects a plurality of sequential image frames. Furthermore, the machine learning device 1a also includes the time-series feature extraction unit 121 to process the feature of sequential image frames.Since the machine learning device 1a selects a large number of sequential image frames, the probability that an image frame will be included in some image frames after processing the attention area increases compared to the case where image frames are selected randomly after processing the attention area.
[0066] The procedure for selecting a large number of sequential image frames is as follows. In a first step, sequential data corresponding to a predetermined number of image frames is selected by naming random segments, thereby increasing the hit rate of the user-processed data compared to the method of randomly selecting individual image frames. Examples of the hit rate simulation are listed below in Table 1. Table 1 Annotation rate for 1000-frame video data 0,2 0,1 0,01 Hit rate in 1000 samples (conventional): The number of segments = 3 47% 27% 4% Hit rate in 1000 samples (second embodiment) The number of sequential frames = 3 64% 41% 6%
[0067] In the machine learning device 1a according to the second embodiment, since the probability that the competence is contained near data that has undergone user processing is high, the selection is centered around the data that underwent user processing at that time of day, based on a probability distribution. By making the selection through prior preparation of a large number of variances of the normal distribution, the feature of a long time series and the feature of a short time series are selected.
[0068] Fig. Figure 13 is a flowchart showing a learning operation performed by the machine learning device 1a according to the second embodiment. The processing in steps S201, S203 to S205 and S207 to S209 in Fig. 13 is the same as the processing in steps S101, S104 to S106 and S107 to S109 in Fig. 10.
[0069] In the second embodiment, the data preference selection unit 101b, which operates as a sequential data pair selection unit, selects data from a multitude of sequential image frames (step S202). By selecting data sequentially in the temporal direction as above, the processed data of the attention domain, which has been annotated by the user, are likely to be included in the objectives of the superiority-or-inferiority determination.
[0070] In the second embodiment, the time-series feature extraction unit 121 further extracts the feature in the time direction from the video by performing the time-direction convolution on the data of the plurality of sequential image frames (step S206).
[0071] As described above, the second embodiment allows for the capture of sequential, detailed movements, such as those shown in a skills video. Simultaneously, the hit rate in the attentional domain increases each time compared to random selection. Consequently, the learning behavior during the acquisition of the learning model can be stabilized.
[0072] With the exception of the features described above, the second embodiment is identical to the first embodiment. Furthermore, the data preference selection unit 101b and the time-series feature extraction unit 121 in the second embodiment are also applicable to the first embodiment. (5) Third embodiment
[0073] Fig. Figure 14 is a functional block diagram which schematically shows the configuration of a machine learning device 1b according to a third embodiment. Fig. 14 is each element that is paired with an element in Fig.5 is identical or corresponds to it, the same reference symbol as in Fig. 5 assigned. The machine learning device 1b according to the third embodiment differs from the machine learning device 1 according to the first embodiment in that it has an attention area comparison unit 131 which calculates a similarity level between image frames processed by the user, and in that a data preference selection unit 101b selects data on the basis of the similarity level.
[0074] In the second embodiment, the device includes means for increasing the hit rate of the image frames edited by the user, but the content of the editing is not taken into account. This means that there are cases where a video pair is selected as a pair of videos whose content differs. Therefore, according to the third embodiment, the machine learning device 1b is equipped with the attention area comparison unit 131, which compares not only the attention area but also the content of the editing. The data preference selection unit 101b adjusts the selection probability so that a video pair is preferentially selected as a pair of videos whose content is similar to the editing. This allows for the comparison of similar attention areas and thus promises an increase in the accuracy of assessing the competence level.
[0075] In a first example of the third embodiment, the attentional area comparison unit 131 first calculates the similarity level between image frames that have undergone user processing. The data preference selection unit 101b selects video data by assigning a higher priority (assigning a higher selection probability) to a video pair with the increased similarity level between image frames that have undergone user processing. However, if data that has not undergone user processing is included in a video pair, the selection probability is reduced by assigning a lower value (e.g., "0.01") than the similarity level value. The model learning unit 102 performs the learning with respect to the selected video pair, and the attentional area generation unit 103 generates the attentional area.If the selected pair consists of data that has not undergone user processing, the attention area comparison unit 131 calculates the level of similarity with another image frame that has undergone user processing or with data that has already been selected as a pair.
[0076] In a second example of the third embodiment, the attention area comparison unit 131 first calculates the similarity level between image frames that have undergone user processing. The attention area comparison unit 131 then preferentially selects a pair that is dissimilar with a certain probability. For example, a dissimilarity level is obtained using a formula "(dissimilarity level) = 1.0 - (similarity level)", and the selection probability is determined based on the dissimilarity level. This is because selecting a dissimilar video pair increases the probability of generating an attention area that was not conceived by the user and contains a video pair that, with a certain probability, has not undergone user processing, leading to the discovery of a new attention area.The subsequent processing is the same as in the first example of the third embodiment.
[0077] In a third example of the third embodiment, the attention-area comparison unit 131 first forms clusters based on similarity level, since it is difficult to store the similarity levels of all image-frame pairs, and then selects data based on the similarity level between the clusters. In this case, the attention-area comparison unit 131 considers each image frame that has undergone user processing as a cluster (in which the number of data elements is 1) and calculates the similarity level between the clusters. Additionally, all image frames that have not undergone user processing are considered unprocessed clusters. Subsequently, the attention-area comparison unit 131 preferentially selects a pair of clusters with a higher similarity level.The level of similarity between the unprocessed cluster and another cluster is considered to be a predetermined low value. For example, attentional area comparison unit 131 randomly selects data that are each contained as a pair in the two selected clusters A and B. It is also possible that attentional area comparison unit 131 does not select the data randomly, but rather selects data at a representative point or data that are dissimilar to other data in the cluster. With such a selection procedure, learning can progress by using a representative point and a point far from the representative point as inputs, and the effect of discovering a new attentional area can also be expected.With respect to the selected pair, the model learning unit 102 performs the learning, and the attentional area generation unit 103 generates the attentional area. If the selected pair consists of data that has not undergone user processing, the attentional area comparison unit 131 then calculates the similarity levels with representative data in other clusters and assigns the data to a cluster with the highest similarity level. The attentional area comparison unit 131 then updates the representative data as having the highest similarity level in the cluster.
[0078] Fig. Figure 15 is a flowchart showing a learning operation performed by the machine learning device 1b according to the third embodiment. The processing in steps S301 and S305 to S310 in Fig. 15 is the same as the processing in steps S101 and S104 to S109 in Fig. 10.
[0079] The attention area comparison unit 131 selects data from a multitude of video pairs (step S301). The attention area comparison unit 131 selects attention area images (attention maps) from the data of the multitude of video pairs (step S302). An attention area comparison unit 131 calculates the similarity level between the attention area images (step S303). As a method for calculating the similarity level, IoU (intersection over union) can be used to calculate the degree of overlap of the attention area images. The attention area comparison unit 131 obtains the total value of the similarity levels between video #1 from a specific time t. k up to a certain time t k+1 and video #2 from a certain time s l up to a certain time s l+1The attentional area comparison unit 131 performs this processing for all sections and normalizes the results so that the total sum equals 1. The data preference selection unit 101b uses random numbers to determine which of the section times t k - t k+1 and s l - s l+1 as sections of the paired data. It is also possible for the data preference selection unit 101b to obtain the sections by, for example, assigning a weight to the similarity level, so that the images edited by the user are likely to be selected. While the similarity level calculation described above is performed for all combinations of video #1 and video #2, it is also possible to select one video at random and the other video based on the similarity level.
[0080] As described above, according to the third embodiment, sequential, detailed movements, such as those in a skills video, can be captured. Simultaneously, the hit rate in the attentional domain increases each time compared to random selection. Consequently, the learning behavior during the acquisition of the learning model can be stabilized.
[0081] With the exception of the features described above, the third embodiment is identical to the first embodiment. Furthermore, the data preference selection unit 101b in the third embodiment is also applicable to the first or second embodiment. (6) Fourth embodiment
[0082] Fig. Figure 16 is a functional block diagram which schematically shows the configuration of a machine learning device 1c according to a fourth embodiment. Fig. 16 is each element that is paired with an element in Fig.5 is identical or corresponds to it, the same reference symbol as in Fig. 5 assigned. The machine learning device 1c according to the fourth embodiment differs from the machine learning device 1 according to the first embodiment in Fig. 5 by having a motion extraction unit 141 and a motion comparison unit 142.
[0083] In the first to third embodiments, the movement of the acting subject in the video is not sufficiently taken into account. To learn the learning model for determining the competence level, it is extremely important to determine whether a movement corresponding to a superiority movement is performed or not. The machine learning device 1c according to the fourth embodiment comprises the motion extraction unit 141 and the motion comparison unit 142 and preferentially selects data and data that are close to each other in terms of movement as a video pair.
[0084] With the machine learning device 1c according to the fourth embodiment, data and data that are close together in the motion can be compared, so that it becomes likely to select similar competence levels as assessment targets, and an increase in the accuracy of the assessment of the competence level can be expected.
[0085] As a first process example for determining the similarity level with respect to movement, a process using hand pose tracking can be considered. This first process example can be carried out according to the following procedure, which comprises processes 11 to 15. (Process 11) An image frame at a time t (t = 0, Δt, 2Δt, 3Δt, ..., NΔt) in each video is extracted. N is a positive integer. (Process 12) A motion vector (i.e., flow) is calculated from image frames from time t = mΔt to time t = (m + 1)Δt, where m = 0, 1, ..., N - 1. (Process 13) A cosine distance of the motion vector (Δx l , Δy l ) is applied to pairs from all frames (N x Frames) of video X and all frames (N y Frames) of video Y is calculated. (N x x N y ) Cosine intervals are calculated. (Process 14) Process 13 is repeated for all video pairs. (Process 15) A video pair is selected more favorably the higher the level of similarity obtained in Process 13.
[0086] As a second example of a procedure for determining the similarity level with respect to motion, a process is suitable that reduces the number of calculations by using clustering or similar techniques. This second process example can be carried out according to the following sequence, which comprises processes 21 to 26. (Process 21) The same process is carried out as process 11 in the first process example. (Process 22) The same process is carried out as process 12 in the first process example. (Process 23) Only the similarity levels between adjacent image frames in video X are calculated, and hierarchical clustering is performed. The number of clusters is determined by a prior definition from the user or similar. (Process 24) The level of similarity between data and data contained in a cluster in Process 23 is preserved, and data that are similar on average to all data in the cluster are called representative data. (Process 25) The similarity level is calculated with respect to the representative data between a cluster set (Cx) of video X and a cluster set (Cy) of video Y, which are generated in Process 23 and Process 24. (Process 26) The clusters are selected using the similarity level between the clusters calculated in Process 25, and a pair is obtained by randomly selecting data from the clusters. After obtaining the clusters, the same processing as in the third embodiment can be applied.
[0087] While user editing is not used in the fourth embodiment, it is also possible to select a data pair by weighting the motion vector based on user editing or by using an index obtained by combining the similarity level in the third embodiment, which is obtained based on user editing, and the similarity level based on the motion vector, or similar.
[0088] Furthermore, it is also possible to obtain the similarity level in a specific region of the video using a distance index from DTW (Dynamic Time Warping) or similar.
[0089] In a method for averaging motion vectors of parts of the video X and obtaining an overall motion vector of a frame, a section of t = m is used. min Δt to t = m maxΔt is selected in video X and segmented into segments, and clustering is performed by preserving the DTW distance between the segments obtained through segmentation. The selection probability is set so that data elements belonging to the same clusters are highly likely to be selected as a pair. For example, elements of in-cluster data from clusters are highly likely to be selected as a pair. In rare cases, a pair of data elements from different clusters will be selected.
[0090] Fig. Figure 17 is a flowchart showing a learning operation performed by the machine learning device 1c according to the fourth embodiment. The processing in steps S401 and S405 to S410 in Fig. 17 is the same as the processing in steps S101 and S104 to S109 in Fig. 10.
[0091] The motion extraction unit 141 selects data from a multitude of video pairs (step S401). The motion comparison unit 142 extracts motion from the data of the multitude of video pairs using a technique such as optical flow (step S402). The motion extraction unit 141 can extract the motion direction of a region, obtained by dividing the image into blocks, and store the direction as a feature vector. The motion comparison unit 142 calculates the similarity level using the cosine distance or similar of the feature vector of the extracted motion (step S403). The motion comparison unit 142 obtains the overall value of the similarity levels between video #1 from a given time t. k up to a certain time t k+1 and video #2 from a certain time s l up to a certain time s l+1The motion comparison unit 142 performs this processing for all segments and normalizes the results so that the total sum equals 1. A data preference selection unit 101c uses random numbers to determine which of the segment times t k - t k+1 and s l - s l+1 as sections of the pair data to be selected (step S404).
[0092] As described above, according to the fourth embodiment, movements are extracted from the data of a large number of video pairs and the learning model is learned using the extracted movements, thereby stabilizing the learning behavior.
[0093] With the exception of the features described above, the fourth embodiment is identical to the first embodiment. Furthermore, the motion extraction unit 141 and the motion comparison unit 142 in the fourth embodiment are applicable to each of the first through third embodiments. (7) Fifth embodiment
[0094] Fig. Figure 18 is a functional block diagram which schematically shows the configuration of a machine learning device 1d according to a fifth embodiment. Fig. 18 is each element that is paired with an element in Fig. 5 is identical or corresponds to it, the same reference symbol as in Fig. 5. The machine learning device 1d according to the fifth embodiment differs from the machine learning device 1 according to the first embodiment in Fig. 5 by having a foreground extraction unit 151 and an attention area comparison unit 152.
[0095] In the first to fourth embodiments, examples are described in which the learning operation is performed with respect to the entirety of a video. However, when learning the learning model for deriving the competence level of the acting subject's action, there are cases in which the result of analyzing the video's background acts as noise and degrades the accuracy of determining the competence level. Therefore, according to the fifth embodiment, the machine learning device 1d comprises the foreground extraction unit 151, which extracts foregrounds from the video pair selected from the video data storage unit 110, and the attention area comparison unit 152, which calculates the similarity level with respect to the attention area using the foreground obtained by removing the background, which is a different area than the foreground, as the attention area.
[0096] Since the area irrelevant to the competence level is hidden or masked as described above, a video pair is selected that is more likely to be directly related to the competence level, so that an improvement in the accuracy of the assessment of the competence level and an improvement in the explainability of the assessment can be expected.
[0097] Fig. Figure 19 is a flowchart showing a learning operation performed by the machine learning device 1d according to the fifth embodiment. The processing in steps S501 and S507 to S512 in Fig. 19 is the same as the processing in steps S101 and S104 to S109 in Fig. 10.
[0098] First, the foreground extraction unit 151 selects a variety of data elements from the video data set storage unit 110 and extracts the foregrounds (steps S501 and S502). Foreground extraction can be performed, for example, by considering an area where no change has occurred between the previous frame and the current frame as the background.
[0099] The foreground extraction unit 151 then obtains the attentional area images from the attentional area storage unit 104 (step S503) and performs a fade-out process with respect to the foregrounds and the attentional area images (step S504). An attentional area comparison unit 152 calculates the similarity level between the faded attentional area images (step S505). The attentional area comparison unit 152 obtains the total value of the similarity levels between video #1 from a specific time t. k up to a certain time t k+1 and video #2 from a certain time s l up to a certain time s l+1The attentional area comparison unit 152 performs this processing for all sections and normalizes the results so that the total sum equals 1. A data preference selection unit 101d uses random numbers to determine which of the section times t k - t k+1 and s l - s l+1 as sections of the pair data to be selected.
[0100] As described above, according to the fifth embodiment, movements are extracted from the data of a large number of video pairs and the learning model is learned using the extracted movements, thereby stabilizing the learning behavior.
[0101] With the exception of the features described above, the fifth embodiment is identical to the first embodiment. Furthermore, the foreground extraction unit 151 and the attention area comparison unit 152 in the fifth embodiment are also applicable to each of the first through fourth embodiments. REFERENCE MARK LIST 1, 1a - 1d Machine learning device, 2 Storage device, 3 processors, 4 Input device, 5 Display device, 101, 101a - 101d Data preference selection unit, 102 Model learning unit, 103 Attentional Area Generation Unit, 104 attention area storage unit, 105 Attention Area Processing Unit, 106 Superiority / Inferiority - Unit of Determination, 110 video data set storage unit, 111, 112 Display example, 121 Time-series feature extraction unit, 131 Attention area comparison unit, 141 Motion extraction unit, 142 motion comparison unit, 151 Foreground extraction unit.
Claims
Machine learning device (1, 1a-1d) that learns a learning model to infer a competence level of an action of an acting subject in a video, the machine learning device comprising: a data preference selection unit (101, 101a-101d) for selecting a plurality of video pairs from a video dataset for learning, and for selecting image frames that are used to determine superiority or inferiority of the competence level from each video pair that forms the plurality of selected video pairs; an attention area generation unit (103) for generating an attention area that is used to determine the superiority or inferiority of the competence level in the image frames; a superiority / inferiority determination unit (106) for performing a determination of the superiority or inferiority of the competence level in the attention area by using the learning model with respect to each video pair;and a model learning unit (102) for storing the learning model and updating the learning model based on a result of determining the superiority or inferiority of the competence level, wherein the data preference selection unit (101, 101a-101d) selects the image frames used to determine the superiority or inferiority of the competence level by using a selection probability determined on the basis of one or more user edits in the image frames forming each video pair, a feature in a time direction in a plurality of sequential image frames forming each video pair, and a level of similarity between the image frames forming each video pair. Machine learning device (1) according to claim 1, wherein the data preference selection unit (101) adjusts the selection probability of an image frame that has undergone user processing in the image frames that form each video pair such that it is higher than the selection probability of an image frame that has not undergone user processing. Machine learning device (1) according to claim 1 or 2, wherein the data preference selection unit (101) increases the selection probability of the image frames with increasing length of the time scope of the user processing, if the user processing has taken place in image frames that form each video pair. Machine learning device (1) according to one of claims 1 to 3, wherein the data preference selection unit (101) increases the selection probability of the image frames with increasing length of a time taken for user processing, if the user processing has taken place in image frames that form each video pair. Machine learning device (1) according to one of claims 1 to 4, wherein the data preference selection unit (101) increases the selection probability of the image frames with an increase in a difference between the attention area after processing and the attention area before processing when the user processing has taken place in image frames that form each video pair. Machine learning device (1) according to one of claims 1 to 5, wherein the data preference selection unit (101) increases the selection probability of the image frames with a decrease in the area of attention when user processing has taken place in image frames that form each video pair. Machine learning device (1a) according to one of claims 1 to 6, wherein the data preference selection unit (101a) determines the selection probability of each image frame based on a feature in the time direction in a plurality of sequential image frames that form each video pair. Machine learning device (1b) according to one of claims 1 to 7, further comprising an attention area comparison unit (131) for calculating the similarity level of each video pair with respect to the attention area, wherein the data preference selection unit (101b) increases the selection probability with an increase in the similarity level between the attention areas of the image frames that form each video pair. Machine learning device (1c) according to any one of claims 1 to 8, further comprising: a motion extraction unit (141) for extracting movements in the attention areas in the video pair; and a motion comparison unit (142) for calculating the level of similarity between the movements in the attention areas in the video pair, wherein the data preference selection unit (101c) increases the selection probability with an increase in the level of similarity between the movements. Machine learning device (1d) according to any one of claims 1 to 9, further comprising: a foreground extraction unit (151) for extracting foregrounds from each video pair; and an attention area comparison unit (152) for calculating the level of similarity between the attention areas in the foregrounds, wherein the data preference selection unit (101d) increases the selection probability with an increase in the level of similarity between the attention areas in the foregrounds. Machine learning device (1, 1a-1d) according to one of claims 1 to 11, wherein the acting subject is a person or a mechanism that moves in connection with a movement sequence of a body part of a person. Competence determination device, comprising: the learning model generated by the machine learning device (1, 1a-1d) according to one of claims 1 to 11, wherein the competence determination device determines a competence level of an action of an acting subject in a video using the learning model as an object. Machine learning method for learning a learning model to derive a competence level of an action of an acting subject in a video, wherein the machine learning method comprises: a step (S101, S201, S301, S401, S501) of selecting a plurality of video pairs from a video dataset for learning and selecting image frames to be used to determine superiority or inferiority of the competence level from each video pair comprising the plurality of selected video pairs; a step (S105, S204, S306, S406, S508) of generating an attention area to be used to determine the superiority or inferiority of the competence level in the image frames; a step (S107, S207, S308, S408, S510) of performing a determination of the superiority or inferiority of the competence level in the attention area by Use of the learning model in relation to each video pair;and a step (S108, S109, S208, S209, S309, S310, S409, S410, S511, S512) of saving the learning model and updating the learning model based on a result of determining the superiority or inferiority of the competence level, wherein in the step (S101, S201, S301, S401, S501) of selecting image frames to be used to determine the superiority or inferiority of the competence level, the image frames to be used to determine the superiority or inferiority of the competence level are selected using a selection probability based on one or more user edits in the image frames that make up each video pair, a feature in a temporal direction in a plurality of sequential image frames that make up each video pair, and a level of similarity between the The image frames that make up each video pair are determined. A machine learning program that instructs a computer to learn a learning model in order to derive a competence level of an action of an acting subject in a video, wherein the machine learning program comprises: a step (S101, S201, S301, S401, S501) of selecting a plurality of video pairs from a video dataset for learning and selecting image frames to be used to determine superiority or inferiority of the competence level from each video pair comprising the plurality of selected video pairs; a step (S105, S204, S306, S406, S508) of generating an attentional area to be used to determine the superiority or inferiority of the competence level in the image frames; a step (S107, S207, S308, S408, S510) of performing a determination of superiority or inferiority of the competence level in the Attention area through use of the learning model in relation to each video pair;and a step (S108, S109, S208, S209, S309, S310, S409, S410, S511, S512) of saving the learning model and updating the learning model based on a result of determining the superiority or inferiority of the competence level, wherein in the step (S101, S201, S301, S401, S501) of selecting image frames to be used to determine the superiority or inferiority of the competence level, the image frames to be used to determine the superiority or inferiority of the competence level are selected using a selection probability based on one or more user edits in the image frames that make up each video pair, a feature in a temporal direction in a plurality of sequential image frames that make up each video pair, and a level of similarity between the The image frames that make up each video pair are determined.