Work analysis device, program, and work analysis method

JPWO2026048084A5Active Publication Date: 2026-08-05MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MITSUBISHI ELECTRIC CORP
Filing Date
2025-01-15
Publication Date
2026-08-05

AI Technical Summary

Technical Problem

Conventional work analysis methods require laborious predefinition of determination criteria, making them time-consuming and inefficient.

Method used

A work analysis device and method that utilizes image and text encoders to automatically identify processes by calculating image and text features, comparing similarities, and extracting focus regions to enhance accuracy and efficiency.

Benefits of technology

Enables easy and accurate identification of work processes, reducing the time and effort required for analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000011_0000
    Figure 00000011_0000
  • Figure 00000011_0001
    Figure 00000011_0001
  • Figure 00000011_0002
    Figure 00000011_0002
Patent Text Reader

Abstract

The work analysis device (120) includes an image feature calculation unit (124) that calculates multiple image features by inputting multiple target image data, which show multiple target images included in video footage of the work in a time series, into an image encoder; a text feature calculation unit (126) that calculates multiple text features by inputting multiple text data, which each show the content of multiple processes constituting the work, into a text encoder; a comparison unit (127) that calculates the similarity between each of the multiple image features and each of the multiple text features; and a process identification unit (128) that identifies one or more processes in a time series by identifying the process whose content is shown in the text corresponding to the text feature with the highest similarity for each of the multiple image features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a work analysis device, a program, and a work analysis method.

Background Art

[0002] In manufacturing and the like, improving the work process for productivity improvement is important. In improving the work process, first, as a current situation analysis, work analysis for measuring the required time for each process constituting the work is necessary. Work analysis is often carried out visually based on images captured by a camera, etc., but it takes a huge amount of time. In contrast, a work analysis device that reduces the time required for work analysis has been proposed.

[0003] For example, Patent Document 1 describes a work data management system that predefines determination criteria such as the position where the operator's body parts pass through each process, and identifies the process performed at each time by satisfying the determination criteria.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, the conventional technology requires predefining determination criteria for each process, which is laborious.

[0006] Therefore, one or more aspects of the present disclosure aim to enable easy identification of processes in work with high accuracy.

Means for Solving the Problems

[0007] A work analysis device according to one aspect of the present disclosure includes: an image feature calculation unit that calculates a plurality of image features, which are a plurality of feature quantities corresponding to each of the plurality of target images, by inputting a plurality of target image data, which are a plurality of target images, contained in a video of a work captured in time series, into an image encoder; a text feature calculation unit that calculates a plurality of text features, which are a plurality of feature quantities corresponding to each of the plurality of processes constituting the work, by inputting a plurality of text data, which are a plurality of feature quantities corresponding to each of the plurality of processes, into a text encoder; a comparison unit that calculates the similarity between each of the plurality of image features and each of the plurality of text features; and a process identification unit that identifies one or more processes in time series by identifying the process whose content is shown in the text corresponding to the text feature with the highest similarity in each of the plurality of image features. A focus region extraction unit extracts one or more focus regions, which are predetermined parts, from each of the multiple frames constituting the aforementioned video, and uses the data indicating the one or more focus regions to construct the multiple target image data. The image encoder and the text encoder are equipped with a pre-trained model such that, for each of the multiple processes, the similarity between the image feature quantity calculated by inputting image data representing the image of the process into the image encoder and the text feature quantity calculated by inputting text data representing the content of the process into the text encoder is increased. Furthermore, the area of ​​focus extraction unit shall, when the worker performing the work is using their entire body to perform the work, or when the worker's posture affects the content of the work, designate the area around the worker performing the work as one or more areas of focus, and when the worker is performing the work with their hands, or when the worker is performing the work with different tools in their hands depending on the process, designate the area around the worker's hands as one or more areas of focus. It is characterized by the following.

[0008] A program according to one aspect of this disclosure includes: an image feature calculation unit that inputs a plurality of target image data representing a plurality of target images included in a video of a work being filmed in time series into an image encoder and calculates a plurality of image feature quantities which are a plurality of feature quantities corresponding to each of the plurality of target images; a text feature calculation unit that inputs a plurality of text data representing the contents of a plurality of steps constituting the work into a text encoder and calculates a plurality of text feature quantities which are a plurality of feature quantities corresponding to each of the plurality of steps; and a comparison unit that calculates the similarity between each of the plurality of image feature quantities and each of the plurality of text feature quantities. 、 A process identification unit identifies one or more processes in a time series by identifying a process whose content is shown in the text corresponding to the text feature with the highest similarity in each of the multiple image features. Furthermore, a focus region extraction unit extracts one or more focus regions, which are predetermined parts, from each of the multiple frames constituting the video, and uses the data indicating the one or more focus regions to construct the multiple target image data.A program that functions as such, wherein the image encoder and the text encoder are pre-trained models such that, for each of the multiple processes, the similarity between the image feature quantities calculated by inputting image data representing the image of one process into the image encoder and the text feature quantities calculated by inputting text data representing the content of one process into the text encoder is high. Furthermore, the area of ​​focus extraction unit shall, when the worker performing the work is using their entire body to perform the work, or when the worker's posture affects the content of the work, designate the area around the worker performing the work as one or more areas of focus, and when the worker is performing the work with their hands, or when the worker is performing the work with different tools in their hands depending on the process, designate the area around the worker's hands as one or more areas of focus. It is characterized by the following.

[0009] The work analysis method relating to one aspect of this disclosure is: The image feature calculation unit, By inputting multiple target image data, which represent multiple target images in a time series, contained in the video footage of the work, into an image encoder, multiple image features, each corresponding to one of the multiple target images, are calculated. The text feature calculation unit, By inputting multiple text data representing the contents of each of the multiple steps constituting the aforementioned work into a text encoder, multiple text features, each corresponding to one of the multiple steps, are calculated. The comparison section is, The similarity between each of the aforementioned multiple image features and each of the aforementioned multiple text features is calculated. The process identification unit, By identifying the process in which the content is shown by the text corresponding to the text feature with the highest similarity in each of the aforementioned multiple image features, one or more processes are identified in the time series. The area of ​​interest extraction unit,A work analysis method comprising extracting one or more areas of interest, which are predetermined parts, from each of a plurality of frames constituting the aforementioned video, and constructing a plurality of target image data using data indicating the one or more areas of interest, wherein the image encoder and the text encoder are pre-trained models such that for each of the plurality of processes, the similarity between the image feature quantity calculated by inputting image data indicating the image of one process into the image encoder and the text feature quantity calculated by inputting text data indicating the content of one process into the text encoder is high, wherein when the worker performing the work uses their whole body to perform the work, or when the worker's posture affects the content of the work, the area around the worker performing the work is designated as the one or more areas of interest, and when the worker performs the work with their hands, or when the worker holds different tools in their hands depending on the process to perform the work, the area around the worker's hands is designated as the one or more areas of interest. [Effects of the Invention]

[0010] According to one or more aspects of this disclosure, the steps in the work can be easily identified with high accuracy. [Brief explanation of the drawing]

[0011] [Figure 1] This is a block diagram schematically showing the configuration of the work analysis system according to Embodiments 1 and 2. [Figure 2] This is a block diagram schematically showing the configuration of the work analysis device in Embodiment 1. [Figure 3] This is a block diagram that shows the general configuration of a PC. [Figure 4] This is a flowchart showing the operation of the work analysis device in Embodiment 1. [Figure 5] This is a block diagram schematically showing the configuration of the work analysis device in Embodiment 2. [Figure 6]It is a flowchart showing the operation of the work analysis device in Embodiment 2.

Embodiments for Carrying Out the Invention

[0012] Embodiment 1. FIG. 1 is a block diagram schematically showing the configuration of a work analysis system 100 according to Embodiment 1. The work analysis system 100 includes a camera 110 as an imaging device and a work analysis device 120.

[0013] The camera 110 and the work analysis device 120 are connected to the network 101, but Embodiment 1 is not limited to such an example. For example, the camera 110 may be connected to the work analysis device 120 via a connection interface conforming to USB (Universal Serial Bus) or the like.

[0014] The camera 110 is installed at the location where the work is performed, captures an image of the work, and generates video data indicating the captured image. Then, the camera 110 transmits the video data to the work analysis device 120 via the network 101.

[0015] FIG. 2 is a block diagram schematically showing the configuration of the work analysis device 120 in Embodiment 1. The work analysis device 120 includes a communication unit 121, a video acquisition unit 122, a region of interest extraction unit 123, an image feature calculation unit 124, a text acquisition unit 125, a text feature calculation unit 126, a comparison unit 127, a process identification unit 128, and an output unit 129.

[0016] The communication unit 121 performs communication via the network 101. In this embodiment, the communication unit 121 receives video data from the camera 110 via the network 101. The received video data is provided to the video acquisition unit 122.

[0017] The video acquisition unit 122 acquires video data. The acquired video data is provided to the area of ​​interest extraction unit 123. In this example, the video acquisition unit 122 acquires video data from the camera 110 via the communication unit 121, but the embodiment is not limited to this example. For example, if the video data is stored in a storage unit (not shown), the video acquisition unit 122 may acquire the video data by reading it from the storage unit. Alternatively, if the video data is stored in a server (not shown), the video acquisition unit 122 may acquire the video data from that server via the communication unit 121.

[0018] The focus region extraction unit 123 extracts one or more focus regions, which are predetermined parts, from each of the multiple frames that make up the video, and constructs multiple target image data using one or more focus region data, which are data indicating one or more focus regions.

[0019] Here, the focus region extraction unit 123 extracts one or more portions from each of the multiple frames that make up the video data, as one or more focus regions. For example, the focus region extraction unit 123 extracts one or more focus regions from the area surrounding the worker who is performing the work being imaged by the camera 110, the area surrounding the worker's hands, and at least one of the changed areas that have changed from the previous frame.

[0020] Specifically, when a worker is using their entire body to perform a task, or when the worker's posture affects the nature of the task, it is desirable to focus on the area surrounding the worker. Furthermore, when a worker is sitting or performing tasks with their hands without changing their posture, or when a worker is using different tools depending on the process, it is desirable to focus on the area around the hands. Furthermore, when performing work using an inspection machine and the inspection results are displayed on the machine's monitor, it is desirable to designate the changed area as the region of interest. Here, we will explain an example in which the surrounding area of ​​the worker, the surrounding area of ​​the hand, and the changing area are all extracted as multiple areas of interest, but it is sufficient if at least one of these is extracted. The extracted data representing one or more areas of interest (in this case, multiple areas of interest) is provided to the image feature calculation unit 124.

[0021] The image feature calculation unit 124 inputs multiple target image data, which represent multiple target images included in the video footage of the work in a time series, into the image encoder, and calculates multiple image features, each of which corresponds to one of the multiple target images. Here, the target image is the region of interest, but the frames that make up the video may also be the target image. In this case, the region of interest extraction unit 123 is unnecessary.

[0022] For example, the image feature calculation unit 124 calculates one or more image features, which are one or more feature quantities, by inputting one or more data points of the region of interest from the region of interest extraction unit 123 into a known image encoder. The calculated one or more image features are provided to the comparison unit 127. Here, since the region of interest extraction unit 123 provides multiple regions of interest, the image feature calculation unit 124 calculates multiple image features by inputting the data representing each of these regions into a known image encoder.

[0023] The text acquisition unit 125 acquires text data that represents text describing each of the multiple steps included in the operation being imaged by the camera 110. The acquired text data is provided to the text feature calculation unit 126. For example, if the text data is stored in a storage unit (not shown), the text acquisition unit 125 may acquire the text data by reading it from the storage unit. Alternatively, if the text data is stored in a server (not shown), the text acquisition unit 125 may acquire the text data from that server via the communication unit 121.

[0024] The text feature calculation unit 126 inputs multiple text data representing the contents of multiple processes that constitute the work into a text encoder, and calculates multiple text features, each corresponding to one of those processes. Here, the text feature calculation unit 126 inputs the text data provided by the text acquisition unit 125 into a known text encoder to calculate text features, which are feature quantities for each process.

[0025] Here, we will explain the image encoder used by the image feature calculation unit 124 and the text encoder used by the text feature calculation unit 126. The image encoder and text encoder used here are pre-trained to convert image data of a certain process included in the work and text data describing the content of that process into similar feature quantities in the same space. The content of the process can be anything as long as it relates to that process.

[0026] In other words, the image encoder and text encoder are pre-trained models designed to maximize the similarity between the image features calculated by inputting image data representing the image of a single process into the image encoder, and the text features calculated by inputting text data representing the content of that single process into the text encoder. Such image encoders and text encoders are described, for example, in the following literature. Literature: Alec Radford et.al, “Leaning Transferable Visual Models From Natural Language Supervision”, arXiv:2103.00020 [cs.CV], 26 Feb 2021

[0027] The comparison unit 127 calculates the similarity between each of the multiple image features and each of the multiple text features. For example, the comparison unit 127 calculates a similarity vector for each frame, showing the similarity between one or more image features and the text features for each process, using one or more image features for each frame and the text features for each process. The similarity calculated here may be, for example, cosine similarity, but is not limited to this. The calculated similarity vector is provided to the process identification unit 128. Here, a similarity vector represents the similarity between a single image feature in a given frame and the text features for each process.

[0028] Here, the comparison unit 127 obtains multiple image features from the image feature calculation unit 124, and a similarity vector is calculated for each of the multiple image features for each frame. Specifically, the comparison unit 127 calculates a similarity vector for each frame with the area around the worker as the area of ​​focus, a similarity vector for each frame with the area around the hand as the area of ​​focus, and a similarity vector for each frame with the area of ​​change as the area of ​​focus.

[0029] The process identification unit 128 identifies one or more processes in a time series by identifying the process whose content is described by the text corresponding to the text feature with the highest similarity calculated for each of the multiple image features. Here, the process identification unit 128 identifies the process shown in each frame from the similarity vector for each frame. For example, if only one area of ​​interest is extracted, the process identification unit 128 can identify the process described by the text data with the highest similarity, as shown by the similarity vector for each frame, as the process shown in that frame.

[0030] On the other hand, if multiple areas of interest have been extracted, the process identification unit 128 determines the process indicated in that frame from the multiple similarity vectors for each frame. For example, the process identification unit 128 may determine the process described by the text data with the highest similarity among the multiple similarity vectors shown for each frame as the process shown in that frame. Furthermore, the process identification unit 128 may determine the process described by the text data with the highest summation or multiplication value of the similarity shown by multiple similarity vectors for each frame as the process shown in that frame.

[0031] The output unit 129 outputs data indicating the process determined for each frame. For example, the output unit 129 may display the data in a predetermined display format on a display (not shown). Furthermore, the output unit 129 may send such data to a predetermined destination via the communication unit 121.

[0032] The work analysis device 120 described above can be implemented, for example, by a computer such as the PC 10 shown in Figure 3. The PC10 includes storage 11 such as an HDD (Hard Disk Drive) and an SSD (Solid State Drive), memory 12, a processor 13 such as a CPU (Central Processing Unit), a communication I / F (Interface) 14 such as a NIC (Network Interface Card), an input I / F 15 such as a keyboard and mouse, and a display 16.

[0033] For example, the communication unit 121 can be implemented using the communication interface 14. The video acquisition unit 122, the area of ​​interest extraction unit 123, the image feature calculation unit 124, the text acquisition unit 125, the text feature calculation unit 126, the comparison unit 127, the process identification unit 128, and the output unit 129 can be realized by loading a program stored in the storage unit 11 into the memory unit 12, and then having the processor 13 execute that program.

[0034] The program may be downloaded to storage 11 via a reader / writer (not shown) from a recording medium (not shown), or via a communication interface 14 from network 101, and then loaded into memory 12 and executed by processor 13. Alternatively, it may be loaded directly into memory 12 via a reader / writer from a recording medium, or via a communication interface 14 from network 101, and then executed by processor 13. In other words, the program may be provided by a computer program product such as a recording medium.

[0035] Figure 4 is a flowchart showing the operation of the work analysis device 120 in Embodiment 1. First, the video acquisition unit 122 acquires video data (S10). The acquired video data is then provided to the area of ​​interest extraction unit 123.

[0036] The area of ​​interest extraction unit 123 sequentially selects one unprocessed frame from among multiple frames that make up the video data (S11). The focus region extraction unit 123 then extracts one or more focus regions from the selected frame (S12). Here, we will explain assuming that multiple focus regions are extracted. The multiple focus region data representing the extracted multiple focus regions are provided to the image feature calculation unit 124.

[0037] The image feature calculation unit 124 inputs each of the multiple focus region data from the focus region extraction unit 123 into a known image encoder to calculate multiple image feature quantities corresponding to each of the multiple focus regions (S13). The calculated multiple image feature quantities are provided to the comparison unit 127.

[0038] The text acquisition unit 125 acquires text data that describes each of the multiple processes included in the operation being imaged by the camera 110, and the text feature calculation unit 126 inputs this text data into a known text encoder to calculate text features, which are characteristic quantities for each process (S14). The calculated text features are provided to the comparison unit 127.

[0039] The comparison unit 127 calculates the similarity between each of the multiple image features and the text features for each process, thereby calculating multiple similarity vectors corresponding to each of the multiple image features (S15). The multiple similarity vectors are provided to the process identification unit 128.

[0040] The process identification unit 128 identifies the process shown in the frame from a plurality of similarity vectors (S16). Here, the process identification unit 128 uses the similarity of each process shown in each of the plurality of similarity vectors to identify the process that has the highest similarity among the processes shown in the frame, according to predetermined criteria.

[0041] The area of ​​interest extraction unit 123 then determines whether all frames of the video data were selected in step S11 (S17). If all frames are selected (Yes in S17), the process proceeds to step S18; if there are still frames that have not been selected (No in S17), the process returns to step S11.

[0042] In step S18, the output unit 129 outputs data indicating the process determined for each frame.

[0043] As described above, according to Embodiment 1, the process in the work can be easily identified with high accuracy.

[0044] Embodiment 2. As shown in Figure 1, the work analysis system 200 according to Embodiment 2 includes a camera 110 and a work analysis device 220. The camera 110 of the work analysis system 200 according to Embodiment 2 is the same as the camera 110 of the work analysis system 100 according to Embodiment 1. However, in Embodiment 2, the camera 110 transmits video data to the work analysis device 220 via the network 101.

[0045] Figure 5 is a block diagram schematically showing the configuration of the work analysis device 220 in Embodiment 2. The work analysis device 220 includes a communication unit 121, an image acquisition unit 122, a focus area extraction unit 123, an image feature calculation unit 124, a text acquisition unit 125, a text feature calculation unit 126, a comparison unit 127, a process identification unit 128, an output unit 129, and a process correction unit 230.

[0046] The communication unit 121, video acquisition unit 122, area of ​​interest extraction unit 123, image feature calculation unit 124, text acquisition unit 125, text feature calculation unit 126, comparison unit 127, process identification unit 128, and output unit 129 of the work analysis device 220 in Embodiment 2 are the same as those of the communication unit 121, video acquisition unit 122, area of ​​interest extraction unit 123, image feature calculation unit 124, text acquisition unit 125, text feature calculation unit 126, comparison unit 127, process identification unit 128, and output unit 129 of the work analysis device 120 in Embodiment 1. However, in Embodiment 2, the process identification unit 128 notifies the process correction unit 230 of the process identified for each frame, and the output unit 129 outputs data indicating the process corrected by the process correction unit 230.

[0047] The process correction unit 230 corrects errors included in the process for each frame identified by the process identification unit 128. For example, the process correction unit 230 corrects a predetermined number of processes by using the most frequent process among a predetermined number of processes that have been identified consecutively in a time series. Specifically, the process correction unit 230 may use the process that is identified most frequently for each predetermined number of frames as the process for that predetermined number of frames.

[0048] Furthermore, the process correction unit 230 may statistically correct the errors in the identified process. Specifically, the process correction unit 230 may statistically correct the errors in the identified process using a known algorithm such as a hidden Markov model.

[0049] The work analysis device 220 described above can also be implemented using a computer such as the PC10 shown in Figure 2. In Embodiment 2, the process correction unit 230 can also be realized by loading a program stored in the storage 11 into the memory 12, and then having the processor 13 execute that program.

[0050] Figure 6 is a flowchart showing the operation of the work analysis device 220 in Embodiment 2. Furthermore, among the processes included in the flowchart shown in Figure 6, those processes that are the same as those included in the flowchart shown in Figure 4 are given the same symbols as those used in the flowchart shown in Figure 4.

[0051] The processes in steps S10 to S17 in Figure 6 are the same as those in steps S10 to S17 in Figure 4. However, in Figure 6, if it is determined in step S17 that all frames have been selected (Yes in S17), the process proceeds to step S20.

[0052] In step S20, the process correction unit 230 corrects errors included in the process for each frame identified by the process identification unit 128. The corrected process for each frame is provided to the output unit 129. Then, the process proceeds to step S18.

[0053] In step S18, the output unit 129 outputs data indicating the process for each frame.

[0054] As described above, according to Embodiment 2, since the process for each frame is corrected, it is possible to identify the process with higher accuracy.

[0055] In Embodiments 1 and 2 described above, the process is specified for each frame, but Embodiments 1 and 2 are not limited to such examples. For example, one sample frame may be extracted for each predetermined number of frames contained in the video, and the process may be specified for each sample frame using the above operation. [Explanation of Symbols]

[0056] 100,200 Work analysis system, 110 Camera, 120,220 Work analysis device, 121 Communication unit, 122 Image acquisition unit, 123 Area of ​​interest extraction unit, 124 Image feature calculation unit, 125 Text acquisition unit, 126 Text feature calculation unit, 127 Comparison unit, 128 Process identification unit, 129 Output unit, 230 Process correction unit.

Claims

1. An image feature calculation unit inputs multiple target image data, which represent multiple target images in a time series, contained in video footage of the work being performed, into an image encoder, and calculates multiple image feature quantities, each of which corresponds to one of the multiple target images. A text feature calculation unit inputs multiple text data representing the contents of multiple steps constituting the aforementioned work into a text encoder, thereby calculating multiple text features that are multiple feature quantities corresponding to each of the multiple steps, A comparison unit that calculates the similarity between each of the plurality of image features and each of the plurality of text features, A process identification unit identifies one or more processes in a time series by identifying a process whose content is shown in the text corresponding to the text feature with the highest similarity in each of the multiple image features, The system includes a focus region extraction unit that extracts one or more focus regions, which are predetermined parts, from each of the multiple frames constituting the video, and uses the data indicating the one or more focus regions to construct the multiple target image data, The image encoder and the text encoder are pre-trained models such that, for each of the multiple processes, the similarity between the image features calculated by inputting image data representing the image of one process into the image encoder and the text features calculated by inputting text data representing the content of one process into the text encoder is high. The area of ​​focus extraction unit shall, when the worker performing the work is using their entire body to perform the work, or when the worker's posture affects the content of the work, designate the area around the worker performing the work as one or more areas of focus, and when the worker is performing the work with their hands, or when the worker is performing the work with different tools in their hands depending on the process, designate the area around the worker's hands as one or more areas of focus. A work analysis device characterized by the following.

2. The aforementioned area of ​​interest extraction unit selects the portion of the video that has changed from the previous frame as one of the one or more areas of interest. The work analysis apparatus according to claim 1, characterized by the following:

3. The system further includes a process correction unit for correcting errors in the identified process. A work analysis apparatus according to claim 1 or 2, characterized by the above.

4. The process correction unit corrects the predetermined number of processes by using the most frequent process among the predetermined number of processes that are consecutively identified in the time series. The work analysis apparatus according to claim 3, characterized by the following:

5. The process correction unit statistically corrects the errors in the identified process. The work analysis apparatus according to claim 3, characterized by the following:

6. Computers, An image feature calculation unit inputs multiple target image data, which represent multiple target images in a time series, contained in video footage of the work being performed, into an image encoder, and calculates multiple image feature quantities, each of which corresponds to one of the multiple target images. A text feature calculation unit inputs multiple text data representing the contents of multiple steps constituting the aforementioned work into a text encoder, and calculates multiple text features that are multiple feature quantities corresponding to each of the multiple steps. A comparison unit that calculates the similarity between each of the plurality of image features and each of the plurality of text features. A process identification unit identifies one or more processes in the time series by identifying a process whose content is shown in the text corresponding to the text feature with the highest similarity in each of the multiple image features, and A program that extracts one or more areas of interest, which are predetermined parts, from each of the multiple frames constituting the aforementioned video, and functions as an area of ​​interest extraction unit that constitutes the multiple target image data using the data indicating the one or more areas of interest, The image encoder and the text encoder are pre-trained models such that, for each of the multiple processes, the similarity between the image features calculated by inputting image data representing the image of one process into the image encoder and the text features calculated by inputting text data representing the content of one process into the text encoder is high. The area of ​​focus extraction unit shall, when the worker performing the work is using their entire body to perform the work, or when the worker's posture affects the content of the work, designate the area around the worker performing the work as one or more areas of focus, and when the worker is performing the work with their hands, or when the worker is performing the work with different tools in their hands depending on the process, designate the area around the worker's hands as one or more areas of focus. A program characterized by the following.

7. By inputting multiple target image data, which represent multiple target images in a time series, contained in the video footage of the work, into an image encoder, multiple image features, each corresponding to one of the multiple target images, are calculated. By inputting multiple text data representing the contents of each of the multiple steps constituting the aforementioned work into a text encoder, multiple text features, each corresponding to one of the multiple steps, are calculated. The similarity between each of the aforementioned multiple image features and each of the aforementioned multiple text features is calculated. By identifying the process in which the content is shown by the text corresponding to the text feature with the highest similarity in each of the aforementioned multiple image features, one or more processes are identified in the time series. A work analysis method which involves extracting one or more areas of interest, which are predetermined parts, from each of the multiple frames constituting the aforementioned video, and constructing the multiple target image data using the data indicating the one or more areas of interest, The image encoder and the text encoder are pre-trained models such that, for each of the multiple processes, the similarity between the image features calculated by inputting image data representing the image of one process into the image encoder and the text features calculated by inputting text data representing the content of one process into the text encoder is high. If the worker performing the aforementioned task uses their entire body to perform the task, or if the worker's posture affects the content of the task, the area surrounding the worker performing the task is designated as one or more areas of focus. If the worker is performing the work by hand, or if the worker is performing the work with different tools in their hand depending on the process, the area around the worker's hand is considered one or more areas of interest. A work analysis method characterized by the following.