Program and information processing method

WO2026191937A1PCT designated stage Publication Date: 2026-09-17DAIKIN INDUSTRIES LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/009273
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-14
Filing Date
2026-03-11
Publication Date
2026-09-17

Smart Images

  • Figure JP2026009273_17092026_PF_FP_ABST
    Figure JP2026009273_17092026_PF_FP_ABST
Patent Text Reader

Abstract

A program according to a first aspect of the present disclosure causes a computer to execute processing for acquiring a moving image captured by a mobile terminal of a worker, extracting, from the moving image, a work scene in which the worker works on equipment, inputting the work scene into a language model and thereby outputting first explanatory data regarding the work scene, and inputting the work scene or the first explanatory data into the language model and thereby outputting second explanatory data which is shorter than the first explanatory data.
Need to check novelty before this filing date? Find Prior Art

Description

Program and Information Processing Method

[0001] The present technology relates to a program and an information processing method.

[0002] Conventionally, programs for managing or viewing video content or audio content have been proposed. For example, the program described in Patent Document 1 acquires video or audio data, the field of the video or audio data, and text information of the video or audio data, divides the video or audio data into one or more sections, and assigns section tags based on the text information.

[0003] Japanese Unexamined Patent Application Publication No. 2023-122236

[0004] However, the program described in Patent Document 1 does not consider at all adding description data to a worker's video that does not include text data.

[0005] The present disclosure has been made in view of such circumstances, and an object of the present disclosure is to provide a program or the like capable of adding description data to a worker's video that does not include text data.

[0006] The program according to a first aspect of the present disclosure causes a computer to execute processing of: acquiring a video captured by a worker's mobile terminal; extracting, from the video, a work scene in which the worker performs work on equipment; inputting the work scene into a language model to output first description data of the work scene; and inputting the work scene or the first description data into a language model to output second description data shorter than the first description data.

[0007] The program according to a second aspect of the present disclosure outputs the first description data in accordance with the skill level of a viewer who views the video.

[0008] The program according to a third aspect of the present disclosure extracts the work scene based on the captured worker's hand, tool, or the equipment.

[0009] In the program according to a fourth aspect of the present disclosure, the second description data includes a noun indicating the equipment or a component of the equipment, and a verb or noun indicating work content.

[0010] A program according to a fifth aspect of the present disclosure outputs the second descriptive data by inputting into a language model a list of nouns indicating the equipment or parts of the equipment, and a list of verbs or nouns indicating the work content.

[0011] The program according to the sixth aspect of this disclosure applies a mosaic effect to parts of the video other than the operator's hands, tools, or equipment that have been filmed.

[0012] A program according to a seventh aspect of the present disclosure extracts a confirmation scene in which the worker performs a confirmation task from the video, and outputs second explanatory data for the confirmation scene by inputting the confirmation scene into a language model or an object detection model.

[0013] The program according to the eighth aspect of this disclosure recognizes the worker's pointing motion in the video and extracts the confirmation scene based on the pointing motion.

[0014] The program according to the ninth aspect of this disclosure stores the video with tags indicating the content of the video.

[0015] A program according to a tenth aspect of the present disclosure accepts the selection of the work scene to be played back based on the second explanatory data, plays back the work scene, and displays the first explanatory data at the same time.

[0016] The program according to the eleventh aspect of this disclosure has different areas in which the work scene to be played is displayed and in which the first explanatory data is displayed.

[0017] A program according to a twelfth aspect of the present disclosure outputs, in chronological order, second explanatory data shorter than first explanatory data indicating a work scene for each work scene in which the worker is working on equipment and for each confirmation scene in which the worker is performing confirmation work, extracted from a video taken by the worker's mobile terminal, accepts the selection of the second explanatory data, and plays back the work scene or confirmation scene corresponding to the selected second explanatory data.

[0018] A program according to a thirteenth aspect of this disclosure plays the work scene corresponding to the selected second explanatory data and simultaneously displays the first explanatory data.

[0019] The program according to the fourteenth aspect of this disclosure has a region in which the work scene to be played is displayed and a region in which the first explanatory data is displayed.

[0020] An information processing method according to a 15th aspect of the present disclosure acquires a video taken by a worker's mobile terminal, extracts a work scene in which the worker is working on equipment from the video, outputs first descriptive data of the work scene by inputting the work scene into a language model, and outputs second descriptive data which is shorter than the first descriptive data by inputting the work scene or the first descriptive data into the language model.

[0021] The information processing method according to the sixteenth aspect of this disclosure outputs, in chronological order, second explanatory data shorter than first explanatory data indicating the work scene for each work scene in which the worker is working on equipment and confirmation scenes in which the worker is performing confirmation work, extracted from a video taken by the worker's mobile terminal, accepts the selection of the second explanatory data, and plays back the work scene or confirmation scene corresponding to the selected second explanatory data.

[0022] In a program according to one embodiment of this disclosure, it is possible to add explanatory data to a video of an operator that does not contain text data.

[0023] This is a schematic diagram illustrating the outline of the information processing system according to this embodiment. This is a block diagram showing the configuration of the server device according to this embodiment. This is a flowchart showing an example of the explanatory data output processing performed by the server device according to this embodiment. This is a schematic diagram showing an example of the landmark detection result by the hand detection unit. This is a schematic diagram showing an example of the moving object detection result by the tool detection unit. This is a schematic diagram illustrating the method of detecting work points by the equipment detection unit. This is a schematic diagram illustrating the method of detecting equipment. This is a schematic diagram showing the extraction of work scenes and confirmation scenes. This is an explanatory diagram showing an example of the output of first explanatory data by the language model. This is an explanatory diagram showing an example of the output of second explanatory data for work scene videos by the language model. This is an explanatory diagram showing an example of the output of second explanatory data for confirmation scene videos by the language model. This is an explanatory diagram showing an example of mosaic processing. This is an explanatory diagram showing an example of a video database. This is an explanatory diagram showing an example of a work scene video database. This is a block diagram illustrating the configuration of the terminal device according to this embodiment. This is a flowchart showing the processing procedure for searching and displaying scene videos performed by the server device and terminal device according to this embodiment. This is an explanatory diagram showing an example of a search results screen. This is an explanatory diagram showing an example of a playback screen.

[0024] (Embodiment) Figure 1 is a schematic diagram illustrating the outline of the information processing system according to this embodiment. In the information processing system according to this embodiment, a worker 102 performing work such as installation or repair of air conditioning equipment 101 has a camera and photographs the work using a wearable terminal 103, including a headset or the like, worn on the worker's head. In this embodiment, the photography is performed using a wearable terminal 103 mounted on a wearable device such as a headset worn by the worker 102, but it is not limited to this, and the work may be photographed by placing the wearable terminal 103 around the air conditioning equipment 101 and the worker 102. Furthermore, photography may be performed using a video camera, handheld camera, smartphone, or tablet terminal. In other words, some or all of the processing by the wearable terminal 103 may be performed by a portable terminal of the worker (used by the worker), such as a video camera, handheld camera, smartphone, tablet terminal, or personal computer. Portable terminals include the wearable terminal 103, video camera, handheld camera, smartphone, tablet terminal, or personal computer. Furthermore, the equipment that is the subject of the photograph is not limited to the air conditioning equipment 101, but may also be a controller, remote control (remote controller), home appliances such as a refrigerator or washing machine, industrial machinery, or vehicles such as a car, ship or aircraft.

[0025] The video footage captured by the wearable terminal 103 is provided to the server device 1. The server device 1 acquires video footage captured by one or more workers and stores the acquired video footage in a database. The method for providing video footage from the wearable terminal 103 to the server device 1 can be, for example, by directly transmitting the video footage from the wearable terminal 103 to the server device 1 via wired or wireless communication, if the wearable terminal 103 is equipped with a communication function. If the wearable terminal 103 is not equipped with a communication function, the wearable terminal 103 can record the video footage on a recording medium such as a memory card or optical disc, and provide the video footage from the wearable terminal 103 to the server device 1 via the recording medium. Alternatively, a terminal device such as a PC (personal computer) or smartphone may be interposed between the wearable terminal 103 and the server device 1, and the terminal device may acquire the video footage from the wearable terminal 103 and transmit it to the server device 1. Any method can be used to provide video footage from the wearable terminal 103 to the server device 1.

[0026] Server device 1 can communicate with language model server 2 via a network such as a LAN (Local Area Network) or the Internet. Language model server 2 stores language model M. Server device 1 extracts work scenes in which an operator is working on equipment from a video captured by the wearable terminal 103, and outputs first descriptive data for the work scene by inputting the extracted work scene into language model M of language model server 2. The first descriptive data is text data that describes the content of the work being performed by the operator in the work scene. Server device 1 also outputs second descriptive data, which is shorter than the first descriptive data, by inputting the first descriptive data into language model M of language model server 2. The second descriptive data is, for example, data of a group of words (captions) indicating the name of the object and the name of the work being performed in the work scene. Server device 1 stores the video of the work scene (work scene video), the first descriptive data, and the second descriptive data in association. Note that language model M may also be stored in server device 1. Furthermore, some or all of the processing performed by the server device 1 may be performed by the wearable terminal 103 or the terminal device 3. In this case, the language model M may be stored in the wearable terminal 103 or the terminal device 3.

[0027] Server device 1 can communicate with one or more terminal devices 3 via a network such as a LAN or the Internet. The terminal devices are general-purpose information processing devices such as PCs or smartphones, and in this embodiment, they are used, for example, by unskilled users learning work such as the installation or repair of air conditioning equipment 101 to view videos (work scene videos) of skilled workers performing the work. Based on a request from terminal device 3, server device 1 retrieves a desired video from among multiple work scene videos stored in the database and transmits it to terminal device 3. Terminal device 3 displays (plays) the video received from server device 1.

[0028] The server device 1 according to this embodiment provides the user with a video search system to assist the user in selecting and viewing a desired work scene video from a large number of work scene videos stored in a database. The server device 1 extracts work scene videos from the work scene videos stored in the database whose captions match the search criteria entered by the user into the terminal device 3, and sends the captions of the extracted work scene videos to the terminal device 3 as search results. The terminal device 3 displays the captions of the search results received from the server device 1 and accepts the user's selection of the work scene to be played.

[0029] Terminal device 3 sends a transmission request to server device 1 regarding the work scene selected by the user. Server device 1 reads the requested work scene video and first explanatory data from the database and transmits them to terminal device 3. Terminal device 3 receives the work scene video and first explanatory data transmitted from server device 1, plays the received work scene video, and displays the first explanatory data. The playback of the work scene video may be performed in parallel with the transmission and playback of the work scene video, a so-called streaming method.

[0030] Figure 2 is a block diagram showing the configuration of the server device 1 according to this embodiment. The server device 1 according to this embodiment is configured to include a control unit 11, a storage unit 12, a communication unit (transceiver) 13, etc. In this embodiment, the explanation is given assuming that processing is performed by one server device, but processing may be performed by multiple server devices in a distributed manner. For example, the server device 1 may consist of two devices: a server device that extracts work scenes from videos captured by the wearable terminal 103 and generates first explanatory data and second explanatory data for the work scenes, and a server device that provides the work scene videos stored in the database to the user's terminal device 3.

[0031] The storage unit 12 is configured using a large-capacity storage device such as a hard disk. The storage unit 12 stores various programs executed by the control unit 11, and various data necessary for processing by the control unit 11. In this embodiment, the storage unit 12 stores the server program 12a executed by the control unit 11. The storage unit 12 is also provided with a video DB 12b for storing videos taken by the wearable terminal 103, and a work scene video DB 12c for storing work scenes extracted from videos taken by the wearable terminal 103.

[0032] In this embodiment, the server program (program product) 12a is provided in a form recorded on a recording medium 99 such as a memory card or optical disc, and the server device 1 reads the server program 12a from the recording medium 99 and stores it in the storage unit 12. However, the server program 12a may also be written to the storage unit 12 during the manufacturing stage of the server device 1, for example. Alternatively, the server program 12a may be obtained by the server device 1 via communication from another remote server device, for example. For example, the server program 12a may be read from the recording medium 99 by a writing device and written to the storage unit 12 of the server device 1. The server program 12a may be provided by distribution via a network, or it may be provided in a form recorded on the recording medium 99.

[0033] The video database 12b is a database that stores information related to videos taken by the wearable terminal 103 and videos of work scenes extracted from those videos (work scene videos). The work scene video database 12c is a database that stores information related to each work scene video extracted from the videos. Details of the video database 12b and the work scene video database 12c will be described later.

[0034] The communication unit 13 communicates with various devices via a network N, which includes a mobile phone communication network, a wireless LAN (Local Area Network), and the Internet. In this embodiment, the communication unit 13 communicates with one or more terminal devices 3 and a wearable terminal 103 via the network N. The communication unit 13 transmits data received from the control unit 11 to other devices and provides data received from other devices to the control unit 11.

[0035] The storage unit 12 may be an external storage device connected to the server device 1. The server device 1 may be a multicomputer comprising multiple computers, or it may be a virtual machine virtually constructed by software. Furthermore, the server device 1 is not limited to the above configuration and may include, for example, a reading unit that reads information stored on a portable storage medium, an input unit that accepts operation input, or a display unit that displays images.

[0036] Furthermore, in the server device 1 according to this embodiment, the control unit 11 reads and executes the server program 12a stored in the storage unit 12, thereby realizing the hand detection unit 11a, tool detection unit 11b, equipment detection unit 11c, work scene extraction unit 11d, first explanatory data output unit 11e, second explanatory data output unit 11f, mosaic processing unit 11g, and DB processing unit 11h, etc., as software-based functional units in the control unit 11. In this figure, the functional units of the control unit 11 that relate to moving images are shown, and functional units related to other processing are omitted from the illustration.

[0037] The hand detection unit 11a performs the process of detecting a person's hand in a video captured by the wearable terminal 103. The video footage obtained by the wearable terminal 103 consists of several dozen frames (still images) per second, and the hand detection unit 11a performs hand detection for each frame that makes up the video footage. However, the hand detection unit 11a does not perform the process of detecting a hand for all frames that make up the video footage, but may reduce the frequency to, for example, about once per second and select frames to detect a hand. The hand detection unit 11a performs the process of detecting a hand captured in a frame that makes up the video footage using, for example, a pre-trained machine learning model, so-called AI (Artificial Intelligence). The learning model of the hand detection unit 11a may include, for example, DNN (Deep Neural Network), CNN (Convolutional Neural Network), FCN (Fully Convolutional Network), SegNet, U-Net, or BlazePalm. In this embodiment, the learning model used by the hand detection unit 11a is pre-trained to, for example, accept a still image as input and output information indicating whether or not what is depicted in each pixel constituting the image is a human hand. In this embodiment, the hand detection unit 11a performs the process of detecting a human hand from an image using a pre-trained learning model, but it is not limited to this, and a human hand may be detected from an image using a method that does not use a learning model.

[0038] In this embodiment, the hand detection unit 11a performs a process to detect specific parts of the detected hand, such as fingertips and joints, so-called landmarks. The hand detection unit 11a performs the detection of specific parts of the hand using a pre-trained model. For example, the learning model may be a DNN, CNN, mediapipe, or OpenPose. This learning model, for example, receives a still image as input, detects a predetermined number of specific parts of the hand depicted in the image, and outputs the coordinates of the detected specific parts. The information regarding the hand landmarks detected by the hand detection unit 11a is used in the detection process of the work object described later. The detailed configuration and learning method of the learning model for hand detection are based on existing technology, so a detailed explanation is omitted. The hand detection unit 11a may also be an image-to-text model. In this case, the hand detection unit 11a outputs text that describes the hand movements based on the input video.

[0039] The tool detection unit 11b performs a process to detect tools shown in a video captured by the wearable terminal 103, based on the hand detection result by the hand detection unit 11a. The tool detection unit 11b detects image regions in which moving objects are shown by comparing two frames that are sequentially preceding or following each other in the multiple frames that make up the video. Next, the tool detection unit 11b defines the portion of the image region where the moving object is shown, excluding the image region corresponding to the hand detected by the hand detection unit 11a, as the image region in which the tool is shown. For each frame, the tool detection unit 11b detects the image region in which the tool is shown and outputs (detects) information indicating whether the object shown in each pixel that makes up the frame is a tool, and the type of tool shown. Alternatively, the tool detection unit 11b may use an image-to-text model. In this case, the tool detection unit 11b outputs text that describes the type or operation of the tool shown, based on the input video.

[0040] The equipment detection unit 11c performs a process to detect equipment that is the target of work such as construction or repair from the video captured by the wearable terminal 103. Unlike the hand detection unit 11a and the tool detection unit 11b, which detect the target of work for each frame of the video, the equipment detection unit 11c detects the target equipment for the entire video. First, based on the detection results of the hand detection unit 11a and the tool detection unit 11b, the equipment detection unit 11c selects frames in which a hand or tool is detected and detects the work points made by the hand or tool in each frame. The equipment detection unit 11c aggregates the information of the work points detected from multiple frames for the entire video or for each scene contained in the video, and identifies (detects) the image region containing the equipment to be worked on and the type of equipment. Alternatively, the equipment detection unit 11c may be an image-to-text model. In this case, the equipment detection unit 11c outputs text that describes the equipment shown based on the input video.

[0041] The work scene extraction unit 11d extracts work scenes from the video recorded by the wearable terminal 103, based on the worker's hands, tools, or equipment that are captured and detected by the wearable terminal 103, in which the worker is working on the equipment. The work scene extraction unit 11d extracts multiple work scenes for each task from the video based on the presence or absence of hand movements of the worker, the type of tool, or the type of equipment shown in each frame of the video. If the hand detection unit 11a detects a pointing motion of the worker's hand, the consecutive frames in which the pointing motion is detected are extracted as a video of a confirmation scene in which the worker is performing a confirmation task (confirmation scene video). In addition, a predetermined number of consecutive frames in which the worker's hands, tools, or equipment are not shown are not included in the work scenes and are treated as a video of a non-work scene (non-work scene video). The work scene extraction unit may also delete the non-work scenes from the video recorded by the wearable terminal 103 and create a digest video by linking the work scenes and confirmation scenes.

[0042] The first explanatory data output unit 11e outputs first explanatory data by inputting each work scene video extracted by the work scene extraction unit 11d into the language model M of the language model server 2 together with a first prompt. The first prompt input to the language model M by the first explanatory data output unit 11e is stored in advance in the storage unit 12, and includes an instruction to output a description of the work shown in the work scene video (first explanatory data), and an instruction to output a plurality of pieces of first explanatory data corresponding to the skill level of the viewer (for example, two levels: beginner and advanced).

[0043] The second explanatory data output unit 11f outputs second explanatory data by inputting the first explanatory data output by the first explanatory data output unit 11e into the language model M of the language model server 2 together with a second prompt. The second prompt input to the language model M by the second explanatory data output unit 11f is stored in advance in the storage unit, and includes an instruction to output a caption (second explanatory data) shorter than the first explanatory data that includes nouns indicating equipment or parts of the equipment and verbs or nouns indicating work content from the first explanatory data. Note that the second explanatory data output unit 11f may output the second explanatory data by inputting the work scene video into the language model M.

[0044] A mosaic processing unit 11g executes mosaic processing on portions other than the photographed and detected worker's hands, tools, and equipment in each frame of the work scene video and the confirmation scene. By applying mosaic processing to portions other than the photographed worker's hands, tools, and equipment, it is possible to protect the worker's personal information or information related to the location of the equipment. Note that the mosaic processing unit 11g may detect an area where privacy information including information related to the worker, information related to the location where the equipment is installed, or information related to the owner of the equipment is captured, and execute mosaic processing on the detected area.

[0045] The DB processing unit 11h stores the work scene video extracted by the work scene extraction unit 11d in the video DB 12b of the storage unit 12, associating it with the original video. It also stores the work scene video in the work scene video DB 12c of the storage unit 12, associating it with the first explanatory data output based on the work scene video and the second explanatory data output based on the first explanatory data. The DB processing unit 11h also obtains search conditions entered by the user from the terminal device 3, extracts work scene videos corresponding to the search conditions from the work scene video DB 12c, and sends a list of the extracted work scene videos as search results to the terminal device 3. The DB processing unit 11h receives a request to play a video from the terminal device 3, reads the data of the work scene video to be played from the work scene video DB 12c, and sends the read video data to the requesting terminal device 3. The DB processing unit 11h may also perform a search for the original video from which the work scene video is extracted and send the original video to the terminal device 3.

[0046] Figure 3 is a flowchart showing an example of the explanatory data output processing performed by the server device 1 according to this embodiment. The control unit 11 of the server device 1 according to this embodiment communicates with the wearable terminal 103, for example, using the communication unit 13, and acquires the video captured by the wearable terminal 103 by receiving the video transmitted from the wearable terminal 103 (S1). A video ID is assigned to each video acquired by the server device 1 to identify it.

[0047] The control unit 11 performs, for each frame constituting the acquired video, a process of detecting the hand of a person (worker) captured in the video by the hand detection unit 11a (S2). At this time, the hand detection unit 11a of the control unit 11 detects a human hand from each frame of the video using a pre-trained learning model that has been subjected to machine learning in advance. In the present embodiment, the hand detection unit 11a detects an image region where the hand is captured in each frame and landmarks such as the fingertips and joints of the hand. FIG. 4 is a schematic diagram showing an example of landmark detection results obtained by the hand detection unit 11a. For a human hand captured in a frame, the hand detection unit 11a estimates specific positions such as the fingertips and joints of the hand as landmarks, and outputs the estimated positions as two-dimensional coordinates in the still image of the frame. The number of landmarks detected by the hand detection unit 11a is approximately 20 to 30 per hand. In FIG. 4, the landmarks detected by the hand detection unit 11a are indicated by black dots, and the landmarks are connected by solid lines to show the shape of the hand and the like. The hand detection unit 11a detects the shape or motion of the hand such as a pointing motion based on the positional relationship of the landmarks.

[0048] The control unit 11 performs, for each frame constituting the acquired video, a process of detecting a tool captured in the video by the tool detection unit 11b (S3). At this time, the tool detection unit 11b of the control unit 11 detects, as a tool, an object gripped by the hand of a person performing work such as construction or repair, and excludes tools not used for work, such as tools placed on a table, from detection targets. Furthermore, in the present embodiment, an object that a worker grips and moves by hand, such as a component removed from an air conditioning equipment 101, is regarded as a tool to be detected by the tool detection unit 11b. That is, the tool detection unit 11b according to the present embodiment does not treat "tools" as general industrial products, but treats any object gripped by a worker's hand as a tool.

[0049] The tool detection unit 11b compares two frames that are sequentially preceding or following each other in a video. The tool detection unit 11b can determine, for example, that the area where the difference in pixel values ​​between two frames exceeds a threshold is the region in which an object that moved between the two frames is captured. Figure 5 is a schematic diagram showing an example of the detection result of a moving object by the tool detection unit 11b. The upper part of Figure 5 shows one frame (still image) of a video captured by the wearable terminal 103, and the lower part of Figure 5 shows the region of the moving object detected by the tool detection unit 11b from this frame as a hatched region. The tool detection unit 11b removes the image region of the hand detected by the hand detection unit 11a from the image region of the moving object detected from the frame of the video, and uses this as the image region in which the tool is captured. For each frame, the tool detection unit 11b outputs information indicating the image region in which the tool is captured, for example, information indicating whether each pixel in the frame is a tool or not, and the type of tool captured in the image. In the example shown in Figure 5, a washing machine is detected.

[0050] The control unit 11 performs a process to detect equipment shown in the video for each frame that makes up the acquired video using the equipment detection unit 11c (S4). The equipment detection unit 11c detects what is at the end of the detected hand or tool as the work target. For frames in which a human hand or tool is detected, the equipment detection unit 11c performs a process to detect the position (work point) where work is being performed by the hand or tool. Figure 6 is a schematic diagram illustrating the method of detecting work points by the equipment detection unit 11c.

[0051] In this embodiment, the equipment detection unit 11c obtains the fingertip positions from landmarks detected by the hand detection unit 11a for a hand detected from the frame, and calculates the center position from the positions of multiple fingertips. In the example shown in Figure 6, the fingertip positions are shown as black circles, and their center positions are shown as white circles. The equipment detection unit 11c identifies the point furthest from the center position of the fingertips from the image area of ​​the tool detected from the frame, and stores the coordinate information of this identified point as a work point. In the example shown in Figure 6, the point furthest from the center position of the fingertips is shown as a star.

[0052] However, for frames where no tool is detected but a hand is detected, the equipment detection unit 11c does not identify the work point using the method described above, but rather identifies the work point using a different method. For frames where no tool is detected, the equipment detection unit 11c obtains the position of the fingertips of the hand, calculates the center position of the multiple fingertips obtained, and uses the calculated center position as the work point.

[0053] The equipment detection unit 11c performs a process to identify work points for all frames in which a hand or tool is detected. The equipment detection unit 11c aggregates the work points identified in each frame for the entire video and identifies areas with a high frequency of identified work points as equipment. Figure 7 is a schematic diagram illustrating the equipment detection method. The example shown is an image in which the frequency of identified work points in one video is color-coded. Lighter colors (white) indicate areas with a low frequency of identified work points, while darker colors (black) indicate areas with a high frequency of identified work points. The equipment detection unit 11c identifies rectangular areas surrounding areas where the frequency exceeds a predetermined threshold as equipment. In the example shown, two areas are detected as equipment. The equipment detection unit 11c also performs image recognition on the equipment areas and outputs the type of equipment (see Figure 6).

[0054] The control unit 11 extracts work scenes from the video captured by the wearable terminal 103 using the work scene extraction unit 11d (S5). In this embodiment, the work scene extraction unit 11d extracts a series of frames in which the same equipment or equipment parts and tools are continuously shown, along with the hands of a worker (person), and makes this a work scene video. The work scene extraction unit 11d may also extract a series of frames in which the same equipment or equipment parts and the hands of a worker are continuously shown, even if different tools are shown, and make this a work scene video, or it may extract a series of frames in which the hands of a worker (person) are continuously shown and make this a work scene video. If the number of frames in which the hands of a worker are not continuously shown is less than a predetermined number, the work scene extraction unit may extract the frames in which the hands of the worker are not continuously shown, and the preceding and following series of frames in which the hands are continuously shown, as a work scene. In other words, the work scene video may include frames in which the hands are not shown. Each work scene video extracted from the video is assigned a scene ID to identify the individual work scene video and the confirmation scene video.

[0055] Furthermore, the control unit 11 uses the work scene extraction unit 11d to extract confirmation scenes from the video captured by the wearable terminal 103 (S6). In this embodiment, when the hand detection unit 11a detects a worker's pointing motion, the work scene extraction unit 11d extracts the consecutive frames in which the pointing motion was detected as a video of the confirmation scene in which the worker performs the confirmation work (confirmation scene video). Each work scene video extracted from the video is assigned a scene ID to identify each work scene video and confirmation scene video.

[0056] Figure 8 is a schematic diagram illustrating the extraction of work scenes and confirmation scenes. The work scene extraction unit 11d of the control unit 11 extracts and separates work scenes, confirmation screens, and non-work scenes based on the worker's hands, tools, or equipment shown (filmed) in the video. The separation of each scene is indicated by the time elapsed from the start of the original video (original video) filmed by the wearable terminal 103. The work scene extraction unit may also delete the non-work scenes from the extracted and separated scenes and create a digest video by concatenating the work scenes and confirmation scenes. In this case, the separation of each scene in the digest video is indicated by the time elapsed from the start of the digest video. That is, the work scene video and confirmation scene video are accompanied by information relating to the start and end times in the original video, as well as the start and end times in the digest video.

[0057] The control unit 11 inputs the extracted work scene video and the first prompt to the language model M via the first explanatory data output unit 11e (S7), and outputs the first explanatory data (S8). The first explanatory data output unit 11e of the control unit 11 reads the first prompt stored in, for example, the storage unit 12, and inputs it together with the work scene video to the language model M of the language model server 2. Figure 9 is an explanatory diagram showing an example of the output of the first explanatory data by the language model M. The language model is composed of a Large Language Model (LLM), such as GPT (Generative Pretrained Transformer) (registered trademark). The language model M also has an aspect as a Vision Language Model (VLM) that outputs text based on the input image. The language model M receives a first prompt that includes a command, output format, and specification of the work scene video, and the specified work scene video as input. The command included in the first prompt includes a command that outputs multiple sentences explaining the work shown in the work scene video as first explanatory data, according to the viewer's skill level.

[0058] The output format included in the first prompt specifies the level of detail and length of the text (first explanatory data) according to the respective skill levels of beginner and advanced. In the example shown in Figure 9, the output format of the first explanatory data for skill level "beginner" is "a text of approximately 30 characters that describes the work procedure in detail," and the output format of the first explanatory data for skill level "advanced" is "a text of approximately 15 characters that describes the outline of the work procedure."

[0059] The work scene video included in the first prompt is represented by a work scene video ID. The language model M receives the work scene video corresponding to the specified work scene video ID.

[0060] The language model M, which receives the aforementioned work scene video and the first prompt, outputs first explanatory data corresponding to the beginner and advanced skill levels. As shown in Figure 9, the first explanatory data for the beginner level is a detailed sentence that reads, "Remove the six screws on the shut-off valve cover, hold the shut-off valve cover with both hands, and remove it," while the first explanatory data for the advanced level is a simple sentence that reads, "Remove the shut-off valve cover."

[0061] The control unit 11 inputs the first explanatory data and the second prompt to the language model via the second explanatory data output unit 11f (S9), and outputs the second explanatory data (S10). Figure 10 is an explanatory diagram showing an example of the output of the second explanatory data for a work scene video by the language model M. The second explanatory data output unit 11f of the control unit 11 reads the second prompt stored in the storage unit 12, for example, and inputs it to the language model M along with the first explanatory data corresponding to the lowest skill level (in this embodiment, the beginner level first explanatory data). The second explanatory data output unit 11f may also input all the first explanatory data related to a single work scene video to the language model M. The second prompt includes an instruction, output format, reference list, and first explanatory data.

[0062] The commands included in the second prompt include commands that output captions as the second explanatory data, which are extracted from the first explanatory data by referring to a reference list and selecting from the words included in the reference list, and which contain nouns indicating the equipment or equipment parts that are the object of the work, and verbs or nouns indicating the work content.

[0063] The output format included in the second prompt includes a specification of the format of the second descriptive data. In the example shown in Figure 10, the output format of the second descriptive data is a "caption phrase combining the name of the equipment or equipment part with a verb or noun indicating the work performed."

[0064] The reference list included in the second prompt preferably includes a list of nouns indicating equipment or equipment parts, and verbs or nouns indicating work procedures, which should be included in the second descriptive data. The reference list may also include associations between the names of equipment or equipment parts and the work procedures that may be performed on said equipment or equipment parts.

[0065] The first explanatory data included in the second prompt is the first explanatory data corresponding to the lowest skill level, output by the first explanatory data output unit 11e. In the example shown in Figure 10, the sentence "Remove the six screws on the shut-off valve cover, hold the shut-off valve cover with both hands, and remove it" output from the language model M in Figure 9 is included in the second prompt as the first explanatory data.

[0066] When the language model M receives the second prompt described above, it outputs a collocation as the second descriptive data, which is a caption extracted from the first descriptive data and combines the name of the equipment or equipment part with a verb or noun indicating the work performed. The second descriptive data shown in Figure 10 is the collocation "removing the shut-off valve cover," which is shorter than the first descriptive data. The second descriptive data may also include synonyms of the words extracted from the first descriptive data. For example, if the language model is input with the first descriptive data "removing the six screws on the shut-off valve cover, holding the shut-off valve cover with both hands, and removing it," then "shut-off valve cover" is extracted as the noun indicating the equipment part, and "removed" is extracted as the verb indicating the work performed. However, in the second descriptive data, "removed" may be converted to the synonym "dismantled," and "shut-off valve cover dismantled" may be output as the second descriptive data. When words extracted from the first descriptive data are converted to synonyms and included in the second descriptive data, it is desirable that these synonyms are words included in the reference list. Furthermore, the second explanatory data is not limited to collocations, but may also consist of sentences, words, or icons. In addition, the second explanatory data output unit 11f may output the second explanatory data by inputting a work scene video to the language model M. In this case, the command included in the second prompt includes a command to output captions as the second explanatory data, which are extracted from the work scene video and consist of nouns indicating the equipment or equipment parts that are the object of the work, and verbs or nouns indicating the work content. The language model M may be pre-trained to output second explanatory data containing words included in the reference list when a work scene is input, using words indicating the names of equipment or equipment parts included in the reference list as training data.

[0067] Furthermore, the control unit 11 inputs the confirmation scene video and the third prompt to the language model via the second explanatory data output unit 11f (S11), and outputs the second explanatory data (S12). Figure 11 is an explanatory diagram showing an example of the output of the second explanatory data for the confirmation scene video by the language model M. The second explanatory data output unit 11f of the control unit 11 reads, for example, the third prompt stored in the storage unit 12 and inputs it to the language model M together with the confirmation scene video. The third prompt includes the instruction, output format, reference list, and specification of the confirmation scene video.

[0068] The commands included in the third prompt include commands that, by referring to a reference list and selecting from the words included in the reference list, extract nouns that indicate the equipment or equipment parts that are the subject of verification from the verification scene video, and output a caption with "verification" attached to the noun as the second explanatory data.

[0069] The output format included in the third prompt includes a specification of the format of the second explanatory data. In the example shown in Figure 11, the output format of the second explanatory data is a "phrase that serves as a caption combining the name of the equipment or equipment part with the word "confirmation"".

[0070] The reference list included in the third prompt preferably includes a list of nouns that indicate equipment or equipment parts, which should be included in the second descriptive data. The reference list may also include a correspondence between the name of the equipment or equipment part and the work that may be performed on the equipment or equipment part.

[0071] The confirmation scene specified in the third prompt is represented by the confirmation scene video ID. The language model M is input with the confirmation scene video corresponding to the specified confirmation scene video ID.

[0072] When the third prompt described above is input to the language model M, it outputs a compound word as second explanatory data, which is a caption combining the name of the equipment or equipment part with the word "confirmation," extracted from the first explanatory data. The second explanatory data shown in Figure 11 is a compound word that serves as the caption for the confirmation scene video, "Confirmation of electrical equipment box." The second explanatory data for the confirmation scene may also be output when the confirmation scene is input to the object detection model. In this case, the object detection model may be composed of the equipment detection unit 11c of the control unit 11, and the second explanatory data may be a compound word combining the name of the equipment or equipment part output by the object detection model with the word "confirmation." The language model M may also be pre-trained to output second explanatory data containing words from the reference list when a confirmation scene is input, using words indicating the names of equipment or equipment parts included in the reference list as training data.

[0073] The control unit 11, using the mosaic processing unit 11g, performs mosaic processing on parts other than the detected worker's hands, tools, and equipment in each frame of the work scene video and the confirmation scene (S13). Figure 12 is an explanatory diagram showing an example of mosaic processing. The mosaic processing unit 11g of the control unit 11 identifies areas (parts) other than the areas of the hands detected by the hand detection unit 11a, the tools detected by the tool detection unit 11b, and the equipment detected by the equipment detection unit 11c, and performs mosaic processing on the identified areas.

[0074] The control unit 11 stores video information, including the video captured by the wearable terminal 103, tags, and IDs of work scene videos or confirmation scene videos extracted from the video, in the video DB 12b via the DB processing unit 11h (S14). Figure 13 is an explanatory diagram showing an example of the video DB 12b. The management items (fields) of the video DB 12b include a video ID field, a video field, a tag field, and a scene ID field. The video ID field stores the video ID assigned to the video captured by the wearable terminal 103. The video field stores the video captured by the wearable terminal 103, for example, in file format. The tag field stores tags that indicate the content of the video. The video tags are, for example, short sentences or phrases that indicate the content of the video, entered by the worker in the wearable terminal 103. The tags may be entered in the terminal device 3 described later, or obtained by inputting the video into the language model M. The scene ID field stores multiple work scene IDs assigned to work scene videos extracted from the video by the work scene extraction unit 11d, or confirmation scene IDs assigned to confirmation scenes, in chronological order as they appear in the video. The video DB 12b may also store the start and end times of the original video for each work scene video or confirmation scene video.

[0075] The control unit 11 stores the work scene video and confirmation screen video extracted by the work scene extraction unit 11d, as well as information related to the work scene video and confirmation screen video (scene video information), in the work scene video DB 12c via the DB processing unit 11h (S15). Figure 14 is an explanatory diagram showing an example of the work scene video DB 12c. The management items (fields) of the work scene video DB 12c include a scene ID field, a scene video field, a video ID field, a beginner-level first explanation data field, an advanced-level first explanation data field, and a second explanation data field. The scene ID field stores the work scene ID assigned to the work scene video, or the confirmation scene ID assigned to the confirmation scene. The scene video field stores the work scene video or confirmation screen video, for example, in file format. The beginner-level first explanation data field stores the beginner-level first explanation data output by the language model M via the first explanation data output unit 11e. The Advanced First Explanation Data Field stores the Advanced First Explanation Data output by the First Explanation Data Output Unit 11e using the Language Model M. In a record where a confirmation scene is recorded, the Beginner First Explanation Data Field and the Advanced First Explanation Data Field store the First Explanation Data output by the First Explanation Data Output Unit 11e using the Language Model M, and the Second Explanation Data Field stores the Second Explanation Data output by the Second Explanation Data Output Unit 11f using the Language Model M. The Work Scene Video DB 12c may store the start and end points of the original video or digest video of the work scene video and the confirmation scene video. After executing S15, the control unit 11 terminates processing.

[0076] Figure 15 is a block diagram showing the configuration of the terminal device 3 according to this embodiment. The terminal device 3 according to this embodiment is configured to include a terminal control unit 31, a storage unit 32, a communication unit (transceiver) 33, a display unit 34, and an operation unit 35, etc. The terminal device 3 is a device used by, for example, an unskilled user learning techniques such as the installation or repair of air conditioning equipment 101, and can be configured using an information processing device such as a smartphone, a tablet terminal device, or a personal computer.

[0077] The terminal control unit 31 is configured using a processing unit such as a CPU or MPU, ROM, etc. The terminal control unit 31 reads and executes a program 32a stored in the storage unit 32, thereby performing processing such as searching for work scene videos or confirmation scene videos stored in the work scene video DB 12c of the server device 1, and displaying (playing) these work scene videos or confirmation scene videos.

[0078] The storage unit 32 is configured using, for example, a non-volatile memory element such as flash memory or a storage device such as a hard disk. The storage unit 32 stores various programs executed by the terminal control unit 31 and various data necessary for processing by the terminal control unit 31. In this embodiment, the storage unit 32 stores the program 32a executed by the terminal control unit 31. In this embodiment, the program 32a is distributed by a remote server device or the like, acquired by the terminal device 3 via communication, and stored in the storage unit 32. However, the program 32a may also be written to the storage unit 32 during the manufacturing stage of the terminal device 3, for example. For example, the program 32a may be read by the terminal device 3 from a recording medium 98 such as a memory card or optical disc and stored in the storage unit 32. For example, the program 32a may be read from the recording medium 98 by a writing device and written to the storage unit 32 of the terminal device 3. The program 32a may be provided by distribution via a network, or it may be provided in the form of being recorded on the recording medium 98.

[0079] The communication unit 33 communicates with various devices via a network N, which includes a mobile phone network, a wireless LAN, and the internet. In this embodiment, the communication unit 33 communicates with the server device 1 via the network N. The communication unit 33 transmits data received from the terminal control unit 31 to other devices and provides data received from other devices to the terminal control unit 31.

[0080] The display unit 34 is configured using a liquid crystal display or the like, and displays various images and characters based on processing by the terminal control unit 31. The operation unit 35 receives user operations and notifies the terminal control unit 31 of the received operations. For example, the operation unit 35 receives user operations via mechanical buttons or an input device such as a touch panel provided on the surface of the display unit 34. Alternatively, the operation unit 35 may be an input device such as a mouse and a keyboard, and these input devices may be configured to be detachable from the terminal device 3.

[0081] Furthermore, in the terminal device 3 according to this embodiment, the terminal control unit 31 reads and executes the program 32a stored in the storage unit 32, thereby realizing the search processing unit 31a, the display processing unit 31b, and other functions as software-based functional units in the terminal control unit 31. The program 32a may be a program dedicated to the information processing system according to this embodiment, or it may be a general-purpose program such as an internet browser or web browser.

[0082] The search processing unit 31a accepts input from the user of words contained in the captions as conditions for searching for work scene videos or confirmation scene videos (scene videos). The search processing unit 31a transmits the input search conditions to the server device 1 and receives information about scene videos extracted from the work scene video DB 12c according to the search conditions from the server device 1. The search processing unit 31a also transmits, for example, a request to play a scene video that matches the search conditions to the server device 1.

[0083] The display processing unit 31b performs display processing such as displaying a screen for accepting input of search conditions and playing (displaying) scene videos.

[0084] Figure 16 is a flowchart illustrating the processing procedure for searching and displaying scene videos performed by the server device 1 and terminal device 3 according to this embodiment. The terminal control unit 31 of the terminal device 3 according to this embodiment displays a search screen for searching scene videos on the display unit 34 using the display processing unit 31b (S21). Although not shown in the figure, the search screen is provided with a text box for accepting text input, and the user can use the operation unit 35 of the terminal device 3 to input a string of search conditions, such as words included in the scene video caption (second explanatory data), into the text box. Note that search conditions for video tags may also be entered on the search screen. The terminal control unit 31 of the terminal device 3 receives input of search conditions for the second explanatory data (caption) from the user using the search processing unit 31a (S22). The control unit 11 provides the server device 1 with a search request including information about the search conditions received by the search processing unit 31a (S23).

[0085] The control unit 11 of the server device 1 acquires a search request that includes the caption search criteria (S24). The control unit 11 of the server device 1 uses the DB processing unit 11h to extract from the work scene video DB 12c work scene video videos or confirmation scene video videos (scene videos) whose captions match the acquired search criteria (the entered words are included in the captions) (S25). The control unit 11 uses the DB processing unit 11h to read the work scene ID of the extracted work scene video or the confirmation scene ID of the confirmation scene video (scene video ID), the second explanatory data, and the video ID of the original video from the work scene video DB 12c (S26), and transmits the search results including the read work scene ID or confirmation scene ID, the second explanatory data, and the video ID of the original video to the terminal device 3 (S27). If tag search criteria are entered on the search screen, the control unit 11 may also extract videos from the video DB 12b whose tags match the search criteria and transmit them to the terminal device 3.

[0086] The terminal control unit 31 of the terminal device 3 acquires search results from the server device 1 (S28). The terminal control unit 31, using the display processing unit 31b, displays a search results screen including the work scene ID or confirmation scene ID of the extracted scene video, second explanatory data, and the video ID of the original video (S29). Figure 17 is an explanatory diagram showing an example of the search results screen. The search results screen display includes multiple search results fields. Each search results field contains information relating to one work scene video or confirmation scene video. The search results field contains the type of video (work scene or confirmation scene), the work scene video ID or confirmation scene ID, the second explanatory data (caption), and the video ID of the original video. The search results field may also contain the first explanatory data, the start or end time in the original video of the work scene video or confirmation scene video, or the start or end time in the digest video of the work scene video or confirmation scene video. In addition, a function button (selection button) for accepting the selection of a work scene video or confirmation scene video is displayed in the search results field. The operation unit 35 can accept the selection of a work scene video or a confirmation scene by receiving the press of a selection button. The terminal control unit 31 accepts the selection of a work scene video or a confirmation scene video (scene video) by the operation unit 35 (S30). Furthermore, the search results screen displays multiple function buttons (skill level buttons) that accept the selection of the user's skill level when viewing the work scene video. In the example shown in Figure 17, two skill level buttons, a beginner button and an advanced button, are displayed. By receiving the press of either skill level button, the operation unit 35 can accept the selection of the skill level corresponding to the first explanatory data to be displayed when the work scene video is played. The terminal control unit 31 accepts the skill level selection by the operation unit 35 (S31). The terminal control unit 31 provides the server device 1 with a request to transmit the scene video, including information about the selected scene video and the skill level (S32).

[0087] The control unit 11 of the server device 1 receives a request to transmit scene videos from the terminal device 3 (S33). The control unit 11, using the DB processing unit 11h, reads the video file of the selected scene video and the first explanatory data corresponding to the selected skill level from the work scene video DB 12c (S34). The control unit 11, using the DB processing unit 11h, also reads the work scene ID of other work scene videos or the confirmation scene ID of confirmation scene videos (IDs of other scene videos), the second explanatory data, and the start time in the original video from the work scene video DB 12c by referring to the video DB 12b (S35). The control unit 11 transmits the read scene video file and first explanatory data, as well as the read IDs of other scene videos, the second explanatory data, and the start time in the original video to the terminal device 3 (S36).

[0088] The terminal control unit 31 of the terminal device 3 acquires the video file and first explanatory data of the selected scene video from the server device 1, as well as the IDs and second explanatory data of other scene videos, and the start time in the original video (S37). The terminal control unit 31 uses the display processing unit 31b to display a playback screen on the display unit 34 that includes playback of the acquired scene video, the first explanatory data, as well as the IDs and second explanatory data of other scene videos included in the original video of the scene video, and the start time in the original video (S38). Figure 18 is an explanatory diagram showing an example of a playback screen. The playback screen includes a video playback field for playing the selected work scene video or confirmation scene video, a first explanatory data display field where the selected proficiency level and the first explanatory data (explanatory text) corresponding to the selected proficiency level are displayed, and an original video field where the IDs and second explanatory data of multiple other scene videos included in the original video of the scene video being played, and the start time in the original video are displayed (output) in chronological order in the original video. As shown in Figure 18, the video playback area and the first explanatory data display area do not overlap on the playback screen, and the area where the work scene to be played is displayed is different from the area where the first explanatory data is displayed. This prevents text from overlapping the work scene and prevents a decrease in the visibility of the work scene video. The original video area displays a function button (play button) for receiving instructions to play the confirmed scene video. When any play button is pressed, the terminal control unit 31 can receive an instruction to play the scene video corresponding to the pressed play button. The terminal control unit 31 accepts the selection of the scene video to be played by receiving the press of the play button via the operation unit 35 (S39). The terminal control unit 31 returns the process to S32 and provides the server device 1 with a request to transmit the scene video, including information about the selected scene video and the proficiency level. Note that in S32, after the process has been returned, the transmission request does not need to include information about the proficiency level. The server device 1 and terminal device 3 repeat the process from S32 onward until they receive input indicating that playback of the scene video should be terminated.

[0089] With the above configuration and processing, even if a work scene video does not contain text data, it is possible to add data that explains the work scene video (first explanatory data). Furthermore, by associating a second explanatory data, which is shorter than the first explanatory data, with the work scene video and storing it, it is possible to facilitate the user's search for work scene videos.

[0090] The embodiments disclosed herein should be considered in all respects as illustrative and not restrictive. The technical features described in each embodiment can be combined with each other, and the scope of the present invention is intended to include all modifications within the claims and scope equivalent to the claims. Furthermore, the independent and dependent claims described in the claims can be combined with each other in any combination, regardless of the form of reference. Moreover, the claims use a multi-claim format in which claims refer to two or more other claims (multi-claim format), but are not limited to this. They may also be described using a multi-claim format in which at least one multi-claim refers to another multi-claim (multi-multi-claim format).

[0091] 1: Server device 11: Control unit 11a: Hand detection unit 11b: Tool detection unit 11c: Equipment detection unit 11d: Work scene extraction unit 11e: First explanatory data output unit 11f: Second explanatory data output unit 11g: Mosaic processing unit 11h: DB processing unit 12: Storage unit 12a: Server program 12b: Video DB 12c: Work scene video DB 13: Communication unit 2: Language model server 3: Terminal device 31: Terminal control unit 31a: Search processing unit 31b: Display processing unit 32: Storage unit 32a: Program 33: Communication unit 34: Display unit 35: Operation unit 98: Recording medium 99: Recording medium 101: Air conditioning equipment 102: Worker 103: Wearable terminal M : Language model N: Network

Claims

1. A program that causes a computer to perform the following processes: acquire video footage taken by a worker's mobile device; extract work scenes from the video showing the worker performing tasks on equipment; input the work scenes into a language model to output first descriptive data for the work scenes; and input either the work scenes or the first descriptive data into the language model to output second descriptive data that is shorter than the first descriptive data.

2. The program according to claim 1, which outputs the first explanatory data according to the skill level of the viewer watching the video.

3. The program according to claim 1 or 2, which extracts the work scene based on the hands, tools, or equipment of the worker that have been photographed.

4. The program according to any one of claims 1 to 3, wherein the second descriptive data includes a noun indicating the equipment or a part of the equipment, and a verb or noun indicating the work content.

5. The program according to claim 4, which outputs the second descriptive data by inputting a list of nouns indicating the equipment or parts of the equipment, and a list of verbs or nouns indicating the work content, into a language model.

6. The program according to any one of claims 1 to 5, which applies a mosaic effect to parts of the video other than the worker's hands, tools, or equipment that have been filmed.

7. A program according to any one of claims 1 to 6, which extracts a confirmation scene in which the worker performs a confirmation task from the video, and outputs second explanatory data of the confirmation scene by inputting the confirmation scene into a language model or an object detection model.

8. The program according to claim 7, which recognizes the pointing motion of the worker in the video and extracts the confirmation scene based on the pointing motion.

9. The program according to any one of claims 1 to 8, which stores the video with tags indicating the content of the video.

10. A program according to any one of claims 1 to 9, which accepts the selection of the work scene to be played back based on the second explanatory data, and plays back the work scene while simultaneously displaying the first explanatory data.

11. The program according to claim 10, wherein the area in which the work scene to be reproduced is displayed is different from the area in which the first explanatory data is displayed.

12. A program that outputs, in chronological order, second explanatory data shorter than the first explanatory data representing the work scene for each work scene in which the worker is working on equipment and for each confirmation scene in which the worker is performing confirmation work, extracted from a video taken by the worker's mobile terminal; accepts the selection of the second explanatory data; and causes a computer to execute a process to play back the work scene or confirmation scene corresponding to the selected second explanatory data.

13. The program according to claim 12, which plays back the work scene corresponding to the selected second explanatory data and simultaneously displays the first explanatory data.

14. The program according to claim 13, wherein the area in which the work scene to be reproduced is displayed is different from the area in which the first explanatory data is displayed.

15. An information processing method that acquires a video taken by a worker's mobile device, extracts a work scene from the video in which the worker is working on equipment, outputs first explanatory data for the work scene by inputting the work scene into a language model, and outputs second explanatory data shorter than the first explanatory data by inputting either the work scene or the first explanatory data into the language model.

16. An information processing method that outputs, in chronological order, second explanatory data shorter than first explanatory data indicating the work scene for each work scene in which the worker is working on equipment and confirmation scenes in which the worker is performing confirmation work, extracted from a video taken by the worker's mobile terminal; accepts the selection of the second explanatory data; and plays back the work scene or confirmation scene corresponding to the selected second explanatory data.