Support program, support device, support system, and support method

JPWO2025229736A5Active Publication Date: 2026-04-07MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing video analysis technologies for factory automation require strict shooting conditions and cannot adapt to flexible or non-standard work scenarios, limiting their applicability in real-world situations.

Method used

A support system that uses a standard work video to extract skeletal information, compares it with target work videos from different angles, and adjusts playback speed to synchronize and analyze partial tasks, allowing for flexible analysis of work performance despite variations in shooting conditions.

Benefits of technology

Enables flexible analysis of work videos in real-world situations by synchronizing and comparing standard and target work videos, identifying deviations and predicting potential issues, thereby improving work analysis efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000013_0000
    Figure 00000013_0000
  • Figure 00000013_0001
    Figure 00000013_0001
  • Figure 00000014_0000
    Figure 00000014_0000
Patent Text Reader

Abstract

The program causes the support device (10) to function as an information acquisition unit (15) that acquires reference skeletal information (32) indicating changes in skeletal posture of a reference worker performing a reference task performed as a standard for a specific task, and acquires reference task information (33) indicating a partial task performed by the reference worker among multiple partial tasks that constitute the specific task, an estimation unit (16) that estimates changes in the skeletal posture in three-dimensional space of a target worker from multiple target videos in which a target worker performing a specific task is filmed from different angles, a task identification unit (17) that identifies the partial task performed by the target worker by comparing the changes, and a display control unit (18) that plays, on the display unit (11), the portions of the reference videos in which the partial tasks are filmed, while playing, on the display unit (11), the portions of the target videos in which the same type of partial task is filmed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to an assistance program, an assistance device, an assistance system, and an assistance method. [Background technology]

[0002] In order to improve the efficiency of work performed by workers at FA (Factory Automation) sites, a technology has been proposed for analyzing the status of such work (see, for example, Patent Document 1). In Patent Document 1, temporal changes in features are extracted from a work video in which the worker's actions are recorded, and the work video is divided into multiple action sections based on the features. Then, the playback speed of the action sections is adjusted so that the playback length of each action section in the video in which the reference action is recorded matches the playback length of each action section in the work video. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2020-155961 A [Patent Document 2] International Publication No. 2021 / 053738 [Non-patent literature]

[0004] [Non-Patent Document 1] Naogo Shimizu, Analyzing worker behavior in manufacturing sites using skeletal information, Proceedings of the 81st National Conference of Information Processing Society of Japan, February 28, 2019, 5C-05 Summary of the Invention [Problem to be solved by the invention]

[0005] In the technology of Patent Document 1, based on the time change of the feature amount extracted from the video, the video of the reference action and the work video are divided into action sections and the divided action sections are associated with each other. Therefore, it is necessary to satisfy the following constraints: the shooting conditions such as the position, attitude, and viewing angle of the camera that shoots these two videos are the same, and the worker performs the same procedure as the reference action every time.

[0006] However, while it is sufficient to film the reference action once, the work to be analyzed is performed continuously, so it is desirable that the filming conditions for the work to be analyzed can be flexibly changed according to the situation. Also, analysis of the work situation is necessary precisely when the worker omits a procedure that should be performed or makes a mistake in the order, but in such cases the above-mentioned constraints are not met, so the technology of Patent Document 1 cannot be applied. In other words, there is room for flexible analysis of the video of the work in accordance with the actual situation at the site.

[0007] The present disclosure has been made in light of the above-mentioned circumstances, and aims to flexibly analyze video footage of work in accordance with the actual situation on-site. [Means for solving the problem]

[0008] In order to achieve the above-mentioned object, the assistance program disclosed herein causes a computer that assists in the analysis of performed work to function as an information acquisition means that acquires reference skeletal information indicating a first transition in the skeletal posture of a reference worker performing a reference work, extracted from a reference video that captures a reference work performed as a reference for a specific work, and acquires reference work information indicating identification information and playback time in the reference video for each of multiple partial tasks that constitute the specific work performed by the reference worker, the reference video being extracted from a reference video that captures a reference work performed as a reference for a specific work, and the reference work information indicating identification information and playback time in the reference video for each of multiple partial tasks that constitute the specific work performed by the reference worker, the reference video being acquired, the reference video being acquired, and multiple target videos that capture a target worker performing the specific work from different angles, the estimation means that estimates a second transition in the posture of the target worker's skeleton in three-dimensional space from the multiple target videos, the identification means that identifies the identification information and playback time in the target videos for each of the partial tasks performed by the target worker by comparing the first transition with the second transition, even if the specific work performed by the target worker includes the omission of any of the partial tasks or a change in the order of the partial tasks, and the display control means that plays, on the display device, a portion of the partial task that corresponds to one of the identification information in the reference video, while playing, on the display device, a portion of the partial task that corresponds to the one of the identification information in the target video. Effect of the Invention

[0009] According to the present disclosure, video footage of work can be analyzed flexibly in accordance with actual conditions at the work site. [Brief description of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram showing a configuration of a support system according to an embodiment. [Diagram 2] FIG. 1 is a diagram showing a hardware configuration of an assistance device according to an embodiment; [Diagram 3] FIG. 2 is a diagram showing a functional configuration of the support device according to the embodiment; [Figure 4] FIG. 1 is a diagram showing an example of a reference moving image according to an embodiment; [Diagram 5] FIG. 1 is a diagram showing an example of a target moving image according to an embodiment; [Figure 6] FIG. 1 is a diagram for explaining reference skeleton information according to an embodiment; [Figure 7] FIG. 1 is a diagram for explaining reference work information according to an embodiment; [Figure 8] FIG. 13 is a diagram showing an example of a registration screen for reference work information according to an embodiment; [Figure 9] FIG. 1 is a diagram for explaining target skeletal information according to an embodiment; [Figure 10] FIG. 1 is a diagram for explaining target work information according to an embodiment; [Figure 11] FIG. 1 is a diagram showing an example of target work information according to an embodiment; [Figure 12] FIG. 1 is a diagram showing an example of an allowable range according to an embodiment; [Figure 13] FIG. 1 is a diagram showing an example of a prohibited range according to an embodiment; [Figure 14] 1 is a flowchart showing a support process according to an embodiment. [Figure 15] 1 is a flowchart showing a predictive notification process according to an embodiment. [Figure 16] FIG. 13 is a diagram showing an example of an analysis screen according to an embodiment; [Figure 17] FIG. 1 is a first diagram for explaining adjustment of the playback speed of a target moving image according to an embodiment; [Figure 18] FIG. 2 is a second diagram for explaining adjustment of the playback speed of a target moving image according to an embodiment; [Figure 19] FIG. 13 is a diagram showing an example of a change operation according to an embodiment; [Figure 20] FIG. 13 is a diagram showing another example of the analysis screen according to the embodiment; [Figure 21] FIG. 13 is a diagram showing an analysis screen according to a modified example. [Figure 22] FIG. 13 is a diagram showing a functional configuration of a support device according to a modified example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Hereinafter, a support system according to an embodiment of the present disclosure will be described in detail with reference to the drawings.

[0012] Embodiment The support system 1000 according to the present embodiment is a system that supports the analysis of work performed by workers in facilities such as factories and plants. Work performed by workers is, for example, processing, inspection, or transportation of workpieces on a production line, and other work. The support system 1000 also supports the analysis by an analyst U1 by playing back, in a comparable form, videos of a reference work, which is a standard for the work performed by the workers, and a target work performed by the worker to be analyzed.

[0013] 1, the support system 1000 includes a support device 10 having a display unit 11 that displays information to an analyst U1, an imaging device 21 that images a reference task performed by a reference operator U2, and imaging devices 22 and 23 that images a target task performed by a target operator U3 who is the subject of analysis. The support device 10 may be connected to the imaging devices 21 to 23 via a communication line such as a Universal Serial Bus (USB) cable, or may be connected to the imaging devices 21 to 23 via a network such as a Local Area Network (LAN).

[0014] The support device 10 is a tablet terminal or an IPC (Industrial Personal Computer) used by an analyst U1. The support device 10 has a hardware configuration as shown in FIG. 2. That is, the support device 10 is configured as a computer having a processor 101, a main memory unit 102, an auxiliary memory unit 103, an input unit 104, an output unit 105, and a communication unit 106. The main memory unit 102, the auxiliary memory unit 103, the input unit 104, the output unit 105, and the communication unit 106 are all connected to the processor 101 via an internal bus 107.

[0015] The processor 101 includes a CPU (Central Processing Unit) as a processing circuit. The processor 101 executes a program P1 stored in the auxiliary storage unit 103 to realize various functions and execute processes described below. The program P1 corresponds to an example of an assistance program.

[0016] The main memory unit 102 includes a RAM (Random Access Memory). A program P1 is loaded into the main memory unit 102 from the auxiliary memory unit 103. The main memory unit 102 is used as a working area for the processor 101.

[0017] The auxiliary storage unit 103 includes a non-volatile memory represented by an EEPROM (Electrically Erasable Programmable Read-Only Memory) and an HDD (Hard Disk Drive). In addition to the program P1, the auxiliary storage unit 103 stores various data used in the processing of the processor 101. The auxiliary storage unit 103 supplies the data used by the processor 101 to the processor 101 according to an instruction from the processor 101. The auxiliary storage unit 103 also stores data supplied from the processor 101.

[0018] The input unit 104 includes input devices such as a hardware switch, an input key, a keyboard, and a pointing device. The input unit 104 acquires information input by the analyst U1, and notifies the processor 101 of the acquired information.

[0019] The output unit 105 includes output devices such as a light emitting diode (LED), a liquid crystal display (LCD), and a speaker. The output unit 105 presents various information to the analyst U1 according to instructions from the processor 101.

[0020] The communication unit 106 includes a communication interface circuit for communicating with an external device. The communication unit 106 receives a signal from the outside and outputs data indicated by the signal to the processor 101. The communication unit 106 also transmits a signal indicating the data output from the processor 101 to the external device.

[0021] The above-mentioned hardware configurations work together to allow the support device 10 to perform various functions. In detail, as shown in Fig. 3, the support device 10 has, as its functions, a display unit 11 on which a video is played, a video acquisition unit 12 for acquiring videos captured by the imaging devices 21 to 23, a storage unit 13 for storing various information, a skeleton extraction unit 14 for extracting the skeleton of a reference worker from a reference video capturing a reference task, an information acquisition unit 15 for acquiring reference skeleton information on the skeleton of the reference worker and reference task information showing the work procedure of the reference worker, an estimation unit 16 for estimating the skeleton of a target worker U3 from a target video capturing a target task, a task identification unit 17 for identifying the task procedure performed by the target worker U3 based on the information in the storage unit 13, a display control unit 18 for displaying support information based on the information in the storage unit 13 on the display unit 11, a reception unit 19 for receiving an operation by an analyst U1, a prediction unit 110 for predicting the future movement of the target worker, and a notification unit 111 for notifying an abnormality in the target task. The display unit 11 is realized by the output unit 105, and the storage unit 13 is mainly realized by the auxiliary storage unit 103. The display unit 105 corresponds to an example of a display device.

[0022] The video acquisition unit 12 is realized mainly by cooperation between the communication unit 106 and the processor 101. The video acquisition unit 12 acquires reference video data 31 including a reference video by receiving it from the shooting device 21, and stores the acquired reference video data 31 in the storage unit 13. The video acquisition unit 12 also acquires target video data 41 including a target video by receiving it from the shooting devices 22, 23, and stores the acquired target video data 41 in the storage unit 13. The video acquisition unit 12 corresponds to an example of a video acquisition means for acquiring a reference video and multiple target videos.

[0023] The reference video data 31 indicates, for example, a reference video captured of one task performed by a reference worker U2. The reference video is, for example, a 15-minute video captured of one task. Since the reference video is referenced as a reference for the task, it is captured from a viewpoint that clearly shows the position and posture of the reference worker U2 and the hands that are the target of the task, as illustrated in FIG. 4. The reference video corresponds to an example of a video captured of a reference task performed as a reference for a specific task.

[0024] The target video data 41 shows a plurality of target videos in which one or more tasks performed by the target worker U3 are captured from different angles by the imaging devices 22, 23. Each of the target videos is, for example, an eight-hour video captured in one day from the start of the first task by the target worker U3 to the end of the last task. The target videos are simultaneously captured from different angles as illustrated in FIG. 5. In FIG. 5, the frames connected by dashed lines indicate frames captured simultaneously. As for the target video, as long as the position and posture of the target worker U3 are known as described later, the same clarity as that of the reference video is not necessarily required for the hands, which are the work target. The viewpoint for capturing the target video may be set arbitrarily according to the situation at the site, and may be changed during the capture of the target video. The target video corresponds to an example of a plurality of videos in which the target worker performing a specific task is captured from different angles.

[0025] The target video data 41 may be video data in a streaming format showing the state of the current work by the target worker U3 in real time. The format of the reference video and the target video may be, for example, MP4 format with a resolution of 1920×1080 pixels and a frame rate of 30 fps (frames per second), or other formats. The reference video data 31 may also include information indicating the model of the imaging device 21, the position and attitude of the imaging device 21, and other information indicating the work environment of the reference work. Similarly, the target video data 41 may include information indicating the model of the imaging devices 22, 23, the positions and attitudes of the imaging devices 22, 23, and other information indicating the work environment of the target work. For example, if the imaging devices 21 to 23 are mobile terminals such as smartphones, the time-series measurement results measured by an acceleration sensor during video shooting may be included in one or both of the reference video data 31 and the target video data 41 as information indicating the attitude of the terminal.

[0026] The skeleton extraction unit 14 is mainly realized by the processor 101. The skeleton extraction unit 14 extracts information indicating the skeleton of the reference worker U2 shown in each frame of the reference video from the reference video. Here, the skeleton is a parameter that defines an outline of the worker's posture, and does not need to match the actual muscles or bones of the reference worker U2. In detail, the skeleton is a plurality of nodes 501 corresponding to representative joint positions of the reference worker U2 shown in each frame, and wires 502 connecting the nodes 501 to each other, as exemplified in FIG. 6. In the example of FIG. 6, eight nodes 501 corresponding to both shoulders, both elbows, both wrists, the base of the neck, and the waist, and seven wires 502 corresponding to the neck to both shoulders, the spine, both upper arms, and both forearms are identified in each frame. However, when a part or all of the reference worker U2 moves outside the viewing angle of the imaging device 21, the part that does not fit within the frame may not be identified. The skeleton extraction unit 14 generates and outputs reference skeleton information 32 indicating the transition of the intra-frame coordinates of the node 501 in each frame as the transition of the skeleton of the reference worker U2. The skeleton extraction unit 14 may specify the skeleton by using the method disclosed in Patent Document 2 or Non-Patent Document 1 proposed by the applicant.

[0027] The information acquisition unit 15 is realized mainly by cooperation between the processor 101 and the communication unit 106. The information acquisition unit 15 acquires reference skeleton information 32 output from the skeleton extraction unit 14 and stores it in the storage unit 13. The information acquisition unit 15 also acquires reference task information 33 indicating the work procedure of a reference task as information input by a registrant (not shown), and stores the acquired reference task information 33 in the storage unit 13.

[0028] The reference task information 33 is information for identifying partial tasks constituting the reference task captured in the reference video, as exemplified in Fig. 7. In detail, the reference task information 33 is data in a table format that associates an ID (identifier), which is identification information, a name, a start time, and an end time for each partial task. Typically, as shown by different hatching on the left side of Fig. 7, the end time of one partial task is equal to the start time of the next partial task. In the example of Fig. 7, the reference task includes five partial tasks: inspection, assembly, screw tightening, appearance check, and cleaning.

[0029] The reference work information 33 may be input to the support device 10 by a registrant who is familiar with the work registering information on the partial work while playing the reference video using a registration screen exemplified in FIG. 8. In the screen of FIG. 8, the playback time of the reference video when the button 801 is pressed is registered as the start time of the partial work, and the playback time of the reference video when the button 802 is pressed is registered as the end time of the partial work. A sequential or unique number may be automatically assigned to the partial work ID without being input by the registrant. The registrant may be the analyst U1 or another user. When the registrant is the analyst U1, the function of acquiring the reference work information 33 may be performed by the reception unit 19 described later.

[0030] The information acquisition unit 15 corresponds to an example of an information acquisition means for acquiring reference skeletal information indicating a first transition of the skeletal posture of the reference worker extracted from the reference video, and acquiring reference task information indicating identification information and playback time in the reference video of each of the partial tasks performed by the reference worker among multiple partial tasks constituting a specific task. The reference task information 33 corresponds to an example of information indicating the playback time length in the reference video of each of the partial tasks constituting the reference task.

[0031] Returning to Fig. 3, the estimation unit 16 is mainly realized by the processor 101. The estimation unit 16 estimates information indicating the skeleton of the target worker U3 appearing in each frame of the target video from the target video. The estimation unit 16 is similar to the skeleton extraction unit 14 in that it specifies a skeleton, but differs from the skeleton extraction unit 14 in that it specifies the coordinates of each node in a three-dimensional space, as exemplified in Fig. 9, based on a plurality of target videos.

[0032] The identification of coordinates in the three-dimensional space may be achieved by applying photogrammetry to a skeleton identified from each target video by the same method as the skeleton extraction unit 14, or by other methods. The origin of the coordinates in the three-dimensional space may be the viewpoint of the camera 22 or the camera 23, the intersection of the optical axes of the camera 22 and the camera 23, the midpoint between the closest points, or another point.

[0033] Furthermore, the distance unit in the three-dimensional space may be determined so that the average size of the skeleton of the human body is equivalent to the identified skeleton. The estimation unit 16 generates target skeleton information 42 indicating the transition of the three-dimensional coordinates of the nodes of each frame as the transition of the skeleton of the target worker U3, and stores the generated target skeleton information 42 in the storage unit 13. The estimation unit 16 corresponds to an example of an estimation means for estimating a second transition of the posture of the skeleton of the target worker in the three-dimensional space from a plurality of target videos.

[0034] The task identification unit 17 is mainly realized by the processor 101. The task identification unit 17 compares the reference skeletal information 32 with the target skeletal information 42, and identifies the partial tasks performed by the target worker U3 by checking the comparison result against the reference task information 33. For example, as shown in Fig. 10, the task identification unit 17 identifies the order in which the partial tasks are performed by the target worker U3 appearing in the target video, and the start time and end time of each partial task.

[0035] The comparison between the reference skeletal information 32 and the target skeletal information 42 can be made by using a priori information that the skeletons are those of a human body, and that the skeletons are similar in transition because the same work is performed in principle, and a method such as PnP (Perspective-n-Point). That is, the transition of the skeleton in the three-dimensional space indicated by the target skeletal information 42 is projected onto a two-dimensional plane equivalent to the reference video, and a transition is calculated. Then, the partial work performed by the target worker U3 is identified by the method of Patent Document 2 based on the reference work information 33.

[0036] The task identification unit 17 not only compares the skeletal transitions but also uses prior information, namely, reference task information 33, to identify partial tasks even when partial tasks are omitted or the order is changed as shown in Fig. 11. In the example of Fig. 11, the first partial task "inspection" that should be performed in the same way as the reference task is omitted, and the order of the partial tasks "visual check" and "cleaning" that should be performed next is reversed. For example, when the skeletal transitions corresponding to the partial tasks of the reference task are combined in an arbitrary order allowing for omissions, and the combination is matched by dynamic programming with the skeletal transitions of the target worker U3 allowing for extension and shortening in the time direction, the combination that produces the smallest error may be adopted as the identification result.

[0037] The task identification unit 17 generates target task information 43 indicating the IDs, names, start times, and end times of the partial tasks identified for the target task in association with each other, and stores the target task information 43 in the storage unit 13. The task identification unit 17 corresponds to an example of an identification means that identifies the identification information and the playback time in the target video of each partial task performed by the target worker by comparing the first transition with the second transition, even if the specific task performed by the target worker includes the omission of any of the partial tasks or a change in the order of the partial tasks. The task identification unit 17 also corresponds to an example of an identification means that identifies the playback time length in the target video of each partial task performed by the target worker.

[0038] The display control unit 18 is mainly realized by the processor 101. The display control unit 18 causes the display unit 11 to play back the reference moving image and the target moving image in a format that allows comparison by the analyst U1. The display contents controlled by the display control unit 18 will be described later.

[0039] The reception unit 19 is mainly realized by the input unit 104. The reception unit 19 receives operations by the analyst U1 on the analysis screen on the display unit 11 displayed by the display control unit 18. The reception unit 19 corresponds to an example of a reception means that receives operations via a user interface. The reception unit 19 also receives range information 50 input by the analyst U1.

[0040] The range information 50 is information indicating at least one of a range in which the movement of the target worker U3 is permitted or prohibited. The analyst U1 may use the target video to register range information 50 indicating a permitted range 1201 as exemplified in FIG. 12, or may register range information 50 indicating a prohibited range 1301 as exemplified in FIG. 13. The range indicated by the range information 50 is not limited to the examples in FIGS. 12 and 13. For example, the range may be defined using both the imaging devices 22 and 23, or may be defined as a range in a three-dimensional space equivalent to the coordinates indicated by the target skeleton information 42.

[0041] The prediction unit 110 is mainly realized by the processor 101. In particular, when the target video is streaming data and the analyst U1 is monitoring the work of the target worker U3, the prediction unit 110 predicts the position and posture of the target worker U3 in the near future based on the transition of the position and posture of the target worker U3 up to the present. This prediction may be made using a human body movement model learned in advance. The near future predicted by the prediction unit 110 is, for example, one second later, ten seconds later, or one minute later. Then, when the prediction unit 110 predicts that the target worker U3 will enter the prohibited range indicated by the range information 50, and when the prediction unit 110 predicts that the target worker U3 will deviate from the allowable range, the prediction unit 110 notifies the notification unit 111 of the prediction result. This notification may be made when the predicted probability of entering the prohibited range or deviating from the allowable range exceeds a predetermined threshold. The prediction target by the prediction unit 110 may be either the position or posture of the target worker U3, or both. The prediction unit 110 corresponds to an example of a prediction means for predicting at least one of the future position and posture of a target worker appearing in a target video.

[0042] The notification unit 111 is mainly realized by the output unit 105. The notification unit 111 notifies the analyst U1 or a notification destination registered in advance of the prediction result by the prediction unit 110. The notification by the notification unit 111 may be a display on the analysis screen used by the analyst U1, a notification to the work supervisor by email or a synthetic voice via a telephone line, or a warning to the target worker at the site by a buzzer sound or a light. The notification unit 111 may notify only in one case of deviation of the target worker U3 from the allowable range or intrusion into the prohibited range, or may notify in both cases. The notification unit 111 corresponds to an example of a notification means for notifying the occurrence of an abnormality in at least one case of a case where the target worker is predicted to move outside a predetermined allowable range and a case where the target worker is predicted to intrude into a predetermined prohibited range.

[0043] Next, the support process executed by the support device 10 will be described with reference to Figs. 14 to 20. The support process shown in Fig. 14 is triggered by a specific operation by an analyst U1 on an analysis application held by the support device 10. The support process corresponds to an example of a support method. Note that the support process shown in Fig. 14 is an example, and the order of each step constituting the support process may be changed as desired.

[0044] In the support process, the video acquisition unit 12 acquires a reference video and extracts reference skeletal information 32 from the reference video, whereby the information acquisition unit 15 acquires the reference skeletal information 32 (step S1). Note that the information acquisition unit 15 may acquire reference skeletal information 32 provided from an external source, instead of the reference skeletal information 32 generated by the skeleton extraction unit 14. Specifically, the information acquisition unit 15 may acquire the reference skeletal information 32 by reading it from a server on a network or from a recording medium such as a memory card inserted into the support device 10. When the reference skeletal information 32 is provided from an external source, the skeleton extraction unit 14 may be omitted from the support device 10.

[0045] Next, the information acquiring unit 15 acquires the reference work information 33, and the accepting unit 19 accepts the range information 50 (step S2). The information acquiring unit 15 may acquire the reference work information 33 registered by a registrant as described above, or may acquire the reference work information 33 by reading it from a server on a network or from a recording medium such as a memory card inserted into the support device 10.

[0046] Next, the video acquisition unit 12 acquires a plurality of target videos, and the estimation unit 16 estimates the transition of the posture of the skeleton of the target worker U3 in three-dimensional space (step S3). As a result, target skeleton information 42 indicating the estimation result is generated. When the target video data is long, the estimation unit 16 may estimate the skeleton of the target worker U3 only for a period of the target video designated by the analyst U1.

[0047] Next, the task identification unit 17 identifies the partial tasks performed by the target worker U3 by comparing the reference skeleton information 32 with the target skeleton information 42 (step S4). As a result, target task information 43 indicating the identification result is generated.

[0048] Next, the prediction unit 110 determines whether the target worker U3 is a prediction target (step S5). If it is determined that the target worker U3 is not a prediction target (step S5; No), the process by the support device 10 proceeds to step S7.

[0049] On the other hand, if it is determined that the target worker U3 is a prediction target (step S5; Yes), a prediction notification process is executed (step S6). In the prediction notification process, as shown in Fig. 15, the prediction unit 110 predicts future changes in the skeletal position and posture of the target worker U3 (step S61).

[0050] Then, the prediction unit 110 judges whether or not it is predicted that the target worker U3 will deviate from the allowable range or enter the prohibited range indicated by the range information 50 (step S62). If the judgment in step S62 is positive (step S62; Yes), the notification unit 111 notifies the occurrence of an abnormality (step S63).

[0051] After step S63 is completed, or if the determination in step S62 is negative (step S62; No), the process by the assistance device 10 returns from the prediction notification process in FIG. 15 to the assistance process in FIG.

[0052] Following step S6 in FIG. 14, the display control unit 18, in accordance with the UI (User Interface) operations received by the receiving unit 19, causes the display unit 11 to synchronously play back the partial work of the reference work and the partial work of the target work, while also synchronously displaying the target video and the skeleton of the target worker U3 on the display unit 11 (step S7).

[0053] As a result, as shown in the analysis screen illustrated in FIG. 16, when the reference video is played in window 1601 and the target video is played in window 1602 at the same time, the ID of the partial task of the reference task being played and the ID of the partial task of the target task will match. Also, as shown in the left side of FIG. 17, the time taken for each partial task may differ between the reference task and the target task, but the playback speed of the target video is adjusted as shown in the right side of FIG. 17 so that the playback times of these partial tasks are equal between the reference video and the target video. For example, for the first partial task "inspection", the recording time is 1 minute 23 seconds in the reference video and 2 minutes 46 seconds in the target video. Therefore, by making the playback speed of the part of the target video corresponding to "inspection" 2.0 times, the playback time of "inspection" on the analysis screen is set to 1 minute 23 seconds, which is common to the reference video and the target video. As a result, the start and end of the partial tasks played in windows 1601 and 1602 in FIG. 16 are synchronized. The playback speed of the other partial tasks is similarly adjusted, so that the reference video and the target video have the same playback time. Therefore, the parts of the reference video corresponding to each partial task and the parts of the target video corresponding to each partial task are played in synchronization with each other. Although an example in which the playback speed of the target video is adjusted has been described, the playback speed of the reference video may be adjusted instead of or together with the target video. The display control unit 18 corresponds to an example of a display control means that changes the playback speed of at least one of the reference video and the target video so that the playback time lengths of the partial tasks are equal.

[0054] 16, a seek bar 1603 indicating the playback time of the reference video and a seek bar 1604 indicating the playback time of the target video are divided to correspond to the playback times of the partial tasks. The portions of the seek bars 1603 and 1604 corresponding to each partial task may be hatched differently or displayed in different colors as in FIG. 16. This allows the analyst U1 to easily recognize the partial task being played back.

[0055] Furthermore, for the seek bars 1603 and 1604, a knob 1606 indicates the current playback time. When the knob 1606 of the seek bar 1603 is dragged by the analyst U1, the knob 1606 of the seek bar 1604 moves in conjunction with the dragged knob 1606. This synchronizes the partial work played in the reference video with the partial work played in the target video. The display control unit 18 corresponds to an example of a display control means that causes the display device to play back a portion of a partial work corresponding to one identification information of the reference video, while causing the display device to play back a portion of a partial work corresponding to the one identification information of the target video.

[0056] As shown in Figure 18, even if there is a omission or change in the order of partial tasks performed by the target worker U3, the playback speed of the target video is adjusted so that the playback time length of the partial tasks is equal between the reference video and the target video. Note that the target video is not played for the part corresponding to the omitted partial task. In addition, for partial tasks with a different execution order, the order of the partial tasks in the target video is changed to be the same as that of the reference task, and then the target video is played.

[0057] 16, on the analysis screen, the target moving image is played in a window 1602, and the transition of the skeleton of the target worker U3 is played in sync in a window 1605. When the playback speed of the target moving image in the window 1602 is adjusted, the playback speed of the transition in the window 1605 is also adjusted. The display control unit 18 corresponds to an example of a display control means that synchronizes the target moving image with the second transition and displays them on the display device.

[0058] Returning to FIG. 14, following step S7, the display control unit 18 changes the display form of the skeleton of the target worker U3 according to the change operation by the analyst U1 accepted by the accepting unit 19 (step S8). The change operation is any one or a combination of two or more operations of enlargement, reduction, translation, rotation in the elevation angle direction, and rotation in the azimuth angle direction, as shown in FIG. 19, for the display of the skeleton in the three-dimensional space. The display control unit 18 changes the position, size, and angle of the skeleton in the window 1605 according to the change operation, as shown in FIG. 20. This makes it possible to arbitrarily change the posture of the skeleton of the target worker U3 displayed while maintaining synchronization with the target video in the window 1602. The accepting unit 19 corresponds to an example of a accepting means for accepting at least one change operation of rotation, enlargement, reduction, and translation for the second transition displayed on the display device.

[0059] Next, the processor 101 judges whether or not an end operation for terminating the assistance process has been performed by the analyst U1 (step S9). If it is judged that the end operation has not been performed (step S9; No), the assistance device 10 repeats the processes from step S3 onwards. As a result, a new analysis screen for comparing the target video transmitted in streaming format or the target video transmitted periodically with the reference video is displayed. If it is judged that the end operation has been performed (step S9; Yes), the assistance process is terminated.

[0060] As described above, the support device 10 according to the present embodiment identifies the posture of the skeleton of the target worker U3 in three-dimensional space by photographing the target worker U3 from different angles, and compares it with the skeleton of the reference worker U2. Therefore, it is not necessary to photograph the target worker U3 from the same viewpoint as the reference video, and the degree of freedom of the viewpoint of the photographing devices 22, 23 that photograph the target worker U3 is increased. In addition, by acquiring the reference task information 33, the support device 10 identifies partial tasks even if the target task includes missing tasks or a change in order. Therefore, it becomes possible to flexibly analyze the video of the task in accordance with the actual situation at the site.

[0061] Although the embodiments of the present disclosure have been described above, the present disclosure is not limited to the above-described embodiments.

[0062] For example, the embodiment has been described in which the reference video and the target video are displayed side by side in a comparative format, but is not limited thereto. For example, if there is a target video in which the target worker U3 is shot from the same viewpoint as the reference video, the reference video and the target video may be displayed superimposed as shown in Fig. 21. In addition, by pressing buttons 2101 and 2102 on the analysis screen in Fig. 21, the skeletons of the reference worker U2 or the target worker U3 may be displayed superimposed.

[0063] Also, although an example has been described in which the moving image acquiring unit 12 directly acquires moving image data from the photographing devices 21 to 23, the present invention is not limited to this. As shown in Fig. 22, the moving image acquiring unit 12 may acquire moving image data by reading, for example, moving image data stored in a server from the server via the network NW1, or may acquire moving image data by reading the moving image data from a recording medium such as a memory card.

[0064] Moreover, the information acquisition unit 15 may acquire reference skeleton information 32 provided from an external source as shown in FIG.

[0065] 22, the display control unit 18 may control the display of information to the analyst U1 by the external UI device 112 via the network NW2, and the reception unit 19 may receive an operation by the analyst U1 on the UI device 112 via the network NW2. The support device 10 may be configured without the display unit 11.

[0066] 22 may be a cloud server on the Internet. The support device 10 may display the reference video and the target video on the UI device 112 via the network. The UI device 112 corresponds to an example of a display device.

[0067] The functions of the support device 10 according to the above-described embodiment can be realized by dedicated hardware or by a general computer system.

[0068] For example, the program P1 can be stored and distributed on a computer-readable recording medium such as a flexible disk, a CD-ROM (Compact Disk Read-Only Memory), a DVD (Digital Versatile Disk), or an MO (Magneto-Optical disk), and the program P1 can be installed on a computer to configure an apparatus that executes the above-mentioned processing.

[0069] Also, the program P1 may be stored in a disk device of a server device on a communication network such as the Internet, and may be downloaded to a computer, for example, by being superimposed on a carrier wave.

[0070] The above-mentioned processing can also be achieved by starting and executing the program P1 while transferring it via a network such as the Internet.

[0071] Furthermore, the above-mentioned processing can also be achieved by executing all or part of the program P1 on a server device, and executing the program P1 while the computer transmits and receives information related to the processing via a communications network.

[0072] In addition, when the above-mentioned functions are shared and realized by the OS (Operating System) or by the OS working together with an application, only the parts other than the OS may be stored on a medium and distributed, or may be downloaded to a computer.

[0073] Furthermore, the means for realizing the functions of the support device 10 is not limited to software, and a part or the whole of the functions may be realized by dedicated hardware or circuits.

[0074] Various embodiments and modifications of the present disclosure are possible without departing from the broad spirit and scope of the present disclosure. The above-described embodiments are for explaining the present disclosure and do not limit the scope of the present disclosure. In other words, the scope of the present disclosure is indicated by the claims, not the embodiments. Various modifications made within the scope of the claims and within the scope of the disclosure equivalent thereto are considered to be within the scope of the present disclosure. [Industrial Applicability]

[0075] This disclosure is suitable for analysis of work performed at a FA site. [Explanation of symbols]

[0076] 10 Support device, 11 Display unit, 12 Video acquisition unit, 13 Memory unit, 14 Skeleton extraction unit, 15 Information acquisition unit, 16 Estimation unit, 17 Task identification unit, 18 Display control unit, 19 Reception unit, 21-23 Shooting device, 31 Reference video data, 32 Reference skeleton information, 33 Reference task information, 41 Target video data, 42 Target skeleton information, 43 Target task information, 50 Range information, 101 Processor, 102 Main memory unit, 103 Auxiliary memory unit, 104 Input unit, 105 Output unit, 106 Communication unit, 107 Internal bus, 110 Prediction unit, 111 Notification unit, 112 UI device, 501 Node, 502 Wire, 801, 802, 2101, 2102 Button, 1000 Support system, 1201 Tolerance range, 1301 Prohibited range, 1601, 1602, 1605 windows, 1603, 1604 seek bars, 1606 knob, NW1, NW2 network, P1 program, U1 analyst, U2 reference worker, U3 target worker.

Claims

1. A computer to assist in the analysis of the work performed, Information acquisition means that acquires reference skeletal information showing the first transition of the skeletal posture of a reference worker performing a reference task, extracted from a reference video of a reference task performed as a standard for a specific task, and acquires reference task information showing the identification information and playback time in the reference video for each of the multiple subtasks constituting the specific task that was performed by the reference worker. A video acquisition means that acquires the aforementioned reference video and multiple target videos of the target worker performing the aforementioned specific task, filmed from different angles. Estimation means for estimating the second transition of the posture of the target worker's skeleton in three-dimensional space from the aforementioned multiple target videos, By comparing the first transition and the second transition, even if the specific task performed by the target worker includes the omission of any of the subtasks or a change in the order of the subtasks, the identification means for identifying the identification information and the playback time in the target video for each of the subtasks performed by the target worker, Display control means that plays back on the display device the portion of one of the reference videos in which the partial work corresponding to the identification information is filmed, while playing back on the display device the portion of the target video in which the partial work corresponding to the identification information is filmed, A support program to enable it to function as such.

2. The aforementioned computer, To further function as a means of receiving operations via a user interface, The display control means causes the second transition to be displayed on the display device. The receiving means receives at least one modification operation among rotation, enlargement, reduction, and translation of the second transition displayed on the display device. The display control means changes the display format of the second transition to be displayed on the display device in accordance with the change operation. The support program according to claim 1.

3. The display control means synchronizes the target video and the second transition and displays them on the display device. The support program according to claim 2.

4. The aforementioned reference work information indicates the playback time of each of the sub-works constituting the reference work in the reference video. The aforementioned identification means identifies the playback time of each of the partial tasks performed by the subject worker in the subject video, The display control means changes the playback speed of at least one of the reference video and the target video so that the playback time of each of the partial operations is equal. The support program according to claim 1 or 2.

5. The aforementioned computer, Predictive means for predicting at least one of the future position and posture of the target worker shown in the target video, The notification means notifies of the occurrence of an abnormality in at least one of the following cases: when the prediction means predicts that the target worker will move outside a predetermined permissible range, and when the prediction means predicts that the target worker will enter a predetermined prohibited range. The support program according to claim 1 or 2, which functions as such.

6. The aforementioned computer is a cloud server on the network, The display control means plays back on the display device via the network a portion of the reference video in which the partial work corresponding to the identification information is filmed, while playing back on the display device via the network a portion of the target video in which the partial work corresponding to the identification information is filmed, The support program according to claim 1 or 2.

7. A support device for analyzing the work performed, Information acquisition means that acquires reference skeletal information showing the first transition of the skeletal posture of a reference worker performing a reference task, extracted from a reference video of a reference task performed as a standard for a specific task, and acquires reference task information showing the identification information and playback time in the reference video for each of the multiple subtasks constituting the specific task that was performed by the reference worker, A video acquisition means that acquires the aforementioned reference video and multiple target videos of the target worker performing the aforementioned specific task, filmed from different angles. Estimation means for estimating the second transition of the posture of the target worker's skeleton in three-dimensional space from the aforementioned multiple target videos, By comparing the first transition and the second transition, even if the specific task performed by the target worker includes the omission of any of the subtasks or a change in the order of the subtasks, the identification means for identifying the identification information and the playback time in the target video for each of the subtasks performed by the target worker, Display control means that plays back on the display device the portion of one of the reference videos in which the partial work corresponding to the identification information is filmed, while playing back on the display device the portion of the target video in which the partial work corresponding to the identification information is filmed, A support device equipped with the following features.

8. The support device according to claim 7, Multiple shooting devices that photograph the target worker and transmit the target video to the support device, A support system equipped with these features.

9. A support method for assisting in the analysis of the work performed, The information acquisition means acquires reference skeletal information showing the first transition of the skeletal posture of a reference worker performing a reference task, extracted from a reference video of a reference task performed as a standard for a specific task, and acquires reference task information showing the identification information and playback time in the reference video for each of the multiple subtasks constituting the specific task that was performed by the reference worker. The video acquisition means acquires the reference video and multiple target videos of the target worker performing the specific task, filmed from different angles. The estimation means estimates the second transition of the posture of the target worker's skeleton in three-dimensional space from the multiple target videos, The identification means, by comparing the first transition and the second transition, identifies the identification information and the playback time in the target video for each of the partial tasks performed by the target worker, even if the specific task performed by the target worker includes the omission of any of the partial tasks or a change in the order of the partial tasks. The display control means causes the display device to play back the portion of the reference video in which the partial work corresponding to the identification information is filmed, while simultaneously playing back the portion of the target video in which the partial work corresponding to the identification information is filmed, Support methods that include this.