Computer program, information processing method, and information processing system
The system uses a wearable device and task classification model to determine work completion by identifying restoration scenes, addressing the challenge of inaccurate task completion detection and minimizing redundant data uploads.
Patent Information
- Application Number
- JP2024171421
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2044-09-30
AI Technical Summary
Existing technologies fail to accurately determine the completion of a work task based on image analysis, particularly when the work determination result for the preceding frame differs.
A computer program and information processing system that utilizes a wearable device to capture work videos, analyzes the scenes using a task classification model, and determines the completion of work by identifying scenes where the work is restored to its original state, with the ability to notify users of any work omissions or complete the task.
Accurately determines the completion of work tasks and reduces unnecessary video uploads by detecting the restoration of work scenes, thereby enhancing efficiency and reducing data redundancy.
Smart Images

Figure 0007825241000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a computer program, an information processing method, and an information processing system. [Background technology]
[0002] Patent Document 1 discloses a technology that uses image analysis to determine the start and end frames of a job from captured frames, inspects the target frame group after the job is completed, and notifies the user if any defects are found. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-191117 Summary of the Invention [Problem to be solved by the invention]
[0004] In Patent Document 1, image analysis is performed on the captured frame, and if the work determination result for the immediately preceding frame is different, it is determined that one work task has been completed; it is not possible to determine the completion of a work task based on the detection result of the work scene.
[0005] An object of the present disclosure is to provide a computer program, an information processing method, and an information processing system that can determine the end of work based on the detection result of the work scene. [Means for solving the problem]
[0006] A computer program according to a first aspect of the present disclosure acquires a work video obtained by filming work performed by a worker, detects a work scene to be restored to its original state based on the acquired work video, and, if the work scene is detected, causes at least one computer to execute a process of determining that the work has been completed.
[0007] A computer program according to a second aspect of the present disclosure causes the computer to execute a process of transmitting the work video to an external device when it is determined that the work has been completed in the computer program according to the first aspect.
[0008] A computer program according to a third aspect of the present disclosure is a computer program according to the first or second aspect, which determines whether or not there is any work omission for each work content based on the work video, and if it determines that there is any work omission, causes the computer to execute a process to notify the user of the work omission.
[0009] A computer program according to a fourth aspect of the present disclosure is a computer program according to any one of the first to third aspects, wherein the work scene is a scene in which assembly work, putting away tools, checking the operation of equipment, or removing protective coverings is being carried out.
[0010] A computer program according to a fifth aspect of the present disclosure is the computer program according to any one of the first to fourth aspects, wherein the work video is a video obtained by capturing an image of repair or inspection work on facility equipment.
[0011] A computer program according to a sixth aspect of the present disclosure is the computer program according to any one of the first to fifth aspects, wherein the work video is video captured by a wearable device worn by the worker.
[0012] A computer program according to a seventh aspect of the present disclosure is the computer program according to any one of the first to sixth aspects, which notifies the user to end imaging when the work scene is detected.
[0013] A computer program according to an eighth aspect of the present disclosure is a computer program according to any one of the first to seventh aspects, in which, when the work scene is not detected and an instruction to end imaging is given, a notification is output inquiring as to whether work to restore the original state has been carried out.
[0014] A computer program according to a ninth aspect of the present disclosure is a computer program according to any one of the first to eighth aspects, which identifies a plurality of work items performed by the worker based on an acquired work video, acquires a list of a plurality of work items to be performed in relation to the work, and determines whether the work has been completed by comparing the identified plurality of work items with the plurality of work items included in the acquired list.
[0015] A computer program according to a tenth aspect of the present disclosure is the computer program according to the ninth aspect, wherein the identified plurality of work items are compared in no particular order with the plurality of work items included in the acquired list.
[0016] A computer program according to an eleventh aspect of the present disclosure is the computer program according to the ninth aspect, which determines whether or not there is any missed work by comparing work items, and if it determines that there is any missed work, notifies the user of the missed work.
[0017] An information processing method according to a twelfth aspect of the present disclosure includes acquiring a work video obtained by filming work performed by a worker, detecting a work scene to be restored to its original state based on the acquired work video, and, if the work scene is detected, executing a process by at least one computer to determine that the work has been completed.
[0018] An information processing system according to a thirteenth aspect of the present disclosure includes at least one processing unit, which acquires a work video obtained by filming work performed by a worker, detects a work scene to be restored to its original state based on the acquired work video, and determines that the work has been completed when the work scene is detected. [Effects of the Invention]
[0019] According to the present disclosure, the end of work can be determined based on the detection result of the work scene. [Brief explanation of the drawings]
[0020] [Figure 1]FIG. 2 is an explanatory diagram illustrating an outline of processing executed by the information processing system according to the first embodiment. [Figure 2] FIG. 2 is a block diagram showing the internal configuration of the wearable device. [Figure 3] FIG. 2 is a block diagram showing the internal configuration of an analysis server. [Figure 4] FIG. 10 is an explanatory diagram illustrating a method for generating a task classification model. [Figure 5] FIG. 10 is an explanatory diagram illustrating an annotation method for a work video. [Figure 6] FIG. 10 is an explanatory diagram illustrating a scene recognition method using a task classification model. [Figure 7] 4 is a flowchart illustrating a procedure of a process executed by an analysis server according to the first embodiment. [Figure 8] FIG. 10 is a conceptual diagram illustrating an example of a work item list. [Figure 9] 10 is a flowchart illustrating a procedure of a process executed by an analysis server according to the second embodiment. [Figure 10] 11 is a flowchart illustrating a procedure of processing executed by a wearable device according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0021] Hereinafter, an information processing system according to an embodiment will be specifically described with reference to the drawings. (Embodiment 1) FIG. 1 is an explanatory diagram illustrating an overview of processing executed by an information processing system 1 according to a first embodiment. The information processing system 1 is a system for capturing images of work performed by workers, acquiring work videos, and analyzing the acquired work videos. The workers in this embodiment are those who perform repair work, inspection work, etc. on facility equipment such as air conditioning equipment. Upon completion of work, such workers perform assembly work, put away tools, check the operation of the facility equipment, remove protective coverings, etc.
[0022] In the embodiment, a wearable device 10 with an imaging function is worn by a worker to capture images of the work being performed by the worker. In one example, the wearable device 10 is worn around the worker's neck. Alternatively, the wearable device 10 may be worn on the worker's head or on another body part such as the shoulder or arm. Furthermore, the wearable device 10 may be a goggle-type camera device, and any device with an imaging function such as a smartphone or an action camera may be used instead of the wearable device 10.
[0023] The wearable device 10 captures images of the work being performed by the worker and generates a video (work video). Alternatively, the wearable device 10 may generate a plurality of still images obtained by capturing images at predetermined time intervals.
[0024] In this embodiment, the work video is a video from the worker's perspective (first-person perspective video), but does not mean a video that completely reproduces the worker's field of view. In other words, the range (field of view) captured by wearable device 10 does not need to strictly match the range visible to the worker, and the capturing direction of wearable device 10 does not need to strictly match the line of sight of the worker. Wearable device 10 only needs to be worn by the worker so that it faces generally forward of the worker, so that at least a portion of the worker's field of view can be captured.
[0025] The information processing system 1 includes an analysis server 20 that analyzes a work video obtained by the wearable device 10. The analysis server 20 is communicatively connected to the wearable device 10 via a communication network NW such as the Internet. The analysis server 20 acquires the work video from the wearable device 10 by communicating with the wearable device 10 via the communication network NW. In the embodiment, the analysis server 20 acquires the work video from the wearable device 10 in real time. The analysis server 20 may acquire the work video frame by frame, or may acquire the work video frame by frame at predetermined time intervals.
[0026] The route by which the analysis server 20 acquires the work video is not limited to a route by which the work video is acquired by directly communicating with the wearable device 10. For example, the work video captured by the wearable device 10 may be transmitted to the analysis server 20 via another terminal device (for example, a smartphone of the worker).
[0027] The analysis server 20 executes a process for distinguishing and recognizing a plurality of types of work scenes based on the acquired work video. The analysis server 20 uses a work classification model MD1 (see FIG. 4), which will be described later, to classify the work of the workers, and is able to distinguish and recognize a plurality of types of work scenes based on the classification results. In this embodiment, the analysis server 20 divides the work video into predetermined time units (for example, 6-second units), and performs scene recognition for each work video in each time unit. The work scenes recognized by the analysis server 20 include work scenes for restoring work to its original state through tasks including assembly work, putting away tools, checking the operation of facility equipment, and removing protective coverings.
[0028] If the analysis server 20 detects a work scene to be restored to its original state as a result of scene recognition, it determines that the worker's work has ended. If it determines that the worker's work has ended, the analysis server 20 may issue an instruction to end image capture to the wearable device 10. In response to the instruction to end image capture from the analysis server 20, the wearable device 10 ends capturing the work video. Since the worker does not need to perform an operation to end the work, there is no need to upload unnecessary videos due to forgetting to perform the operation.
[0029] When the analysis server 20 determines that the worker has completed the work, it determines whether any work has been omitted based on the work video captured during the work, and if it determines that any work has been omitted, it may notify the worker. The notification to the worker may be sent via the wearable device 10 or via another terminal device.
[0030] 2 is a block diagram showing the internal configuration of wearable device 10. Wearable device 10 includes a processing unit 11, a storage unit 12, a communication unit 13, an imaging unit 14, a sound input unit 15, a sound output unit 16, a sensor unit 17, an operation unit 18, and the like.
[0031] The processing unit 11 includes a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The ROM included in the processing unit 11 stores control programs and the like that control the operation of each hardware unit included in the wearable device 10. The CPU in the processing unit 11 reads and executes the control programs and the like stored in the ROM, and controls the operation of each hardware unit, thereby causing the entire device to function as the wearable device 10 of the present disclosure. The RAM included in the processing unit 11 temporarily stores data used during the execution of various processes.
[0032] The storage unit 12 includes an auxiliary storage device and stores the work video and the like generated by the imaging unit 14. The storage unit 12 may also be installed with an application program to be executed by the processing unit 11. The application program may be installed in advance or may be installed after use has begun.
[0033] The communication unit 13 includes a communication module for wireless communication with an external device such as the analysis server 20. The communication module may be a communication module for wireless communication using a known mobile communication standard such as 3G, 4G, or 5G, or a wireless LAN system such as WiFi (registered trademark). The communication unit 13 communicates with an external device such as the analysis server 20 via a communication network NW, transmitting necessary data such as a work video, and receiving appropriate data transmitted from the external device. The communication unit 13 may also include a communication module for short-range wireless communication such as Bluetooth (registered trademark) or ZigBee (registered trademark) to communicate with a device such as a smartphone carried by the worker.
[0034] The imaging unit 14 includes an optical lens, an imaging element, a driver circuit, etc. A wide-angle lens is preferably used as the optical lens. The imaging element is a CMOS (Complementary Metal Oxide Semiconductor), a CCD (Charge-Coupled Device), etc., and generates an electrical signal according to the intensity of light imaged through the optical lens. The driver circuit includes a timing generator (TG), etc., and sequentially reads out the electrical signals from the imaging element in synchronization with a clock signal output from the TG to generate video data. The video data generated by the imaging unit 14 is sent to the processing unit 11 and stored in the memory unit 12. Alternatively, the video data generated by the imaging unit 14 is transmitted to the analysis server 20 via the communication unit 13.
[0035] The sound input unit 15 includes a microphone for collecting sound, a processing circuit for converting the collected sound into a digital signal (sound data), etc. The sound data generated by the sound input unit 15 is sent to the processing unit 11, where appropriate processing such as noise removal is performed. The sound data generated by the sound input unit 15 is stored in the memory unit 12, or transmitted to the analysis server 20 via the communication unit 13.
[0036] The sound output unit 16 includes a speaker that outputs sound. The sound output unit 16 outputs sound based on the acoustic data provided by the processing unit 11.
[0037] The sensor unit 17 includes a non-contact sensor for detecting the worker's fingers and the like. The sensors included in the sensor unit 17 include a proximity sensor, a gesture sensor, and the like. The proximity sensor detects, for example, when the worker's fingers have approached within a predetermined range. The gesture sensor detects, for example, the movement of the worker's fingers. The detection result by the sensor unit 17 is notified to the processing unit 11. The processing unit 11 may issue an instruction to start capturing or stop capturing to the imaging unit 14 based on the detection result of the sensor unit 17.
[0038] Operation unit 18 includes various operation buttons, operation switches, etc., and receives operations from the worker. Operation information corresponding to the operation of operation unit 18 is input to processing unit 11. Processing unit 11 executes appropriate processing based on the operation information input from operation unit 18. The worker may operate operation unit 18 to give wearable device 10 an instruction to start imaging or an instruction to stop imaging.
[0039] 3 is a block diagram showing the internal configuration of the analysis server 20. The analysis server 20 is a dedicated or general-purpose server device, and includes a processing unit 21, a storage unit 22, a communication unit 23, an operation unit 24, and a display unit 25, for example.
[0040] The processing unit 21 includes a CPU, a ROM, a RAM, etc. The ROM included in the processing unit 21 stores a control program and the like that controls the operation of each hardware unit included in the analysis server 20. The CPU in the processing unit 21 reads and executes the control program stored in the ROM and a computer program (described below) stored in the storage unit 22, and executes processing to control the operation of each hardware unit, thereby causing the entire device to function as the analysis server 20 of the present disclosure. The RAM included in the processing unit 21 temporarily stores data used during the execution of various processes.
[0041] In the embodiment, the processing unit 21 is configured to include a CPU, a ROM, and a RAM, but the configuration of the processing unit 21 is not limited to the above. The processing unit 21 may be one or more processing circuits including, for example, a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), a quantum processor, volatile or non-volatile memory, etc. Furthermore, the processing unit 21 may also have functions such as a clock that outputs date and time information, a timer that measures the elapsed time from when a measurement start instruction is given until when a measurement end instruction is given, and a counter that counts numbers.
[0042] The storage unit 22 includes a storage device such as a hard disk drive (HDD), a solid state drive (SSD), etc. The storage unit 22 stores various computer programs executed by the processing unit 21 and various data acquired via the communication unit 23.
[0043] The computer programs (program products) stored in the storage unit 22 include an analysis processing program PG1 for analyzing work videos. The analysis processing program PG1 is a computer program for causing a computer to execute processing to acquire work videos obtained by capturing images of work performed by a worker, detect work scenes to be restored to their original state based on the acquired work videos, and, when such work scenes are detected, determine that the work performed by the worker has been completed.
[0044] The analysis processing program PG1 may be a single computer program or a group of programs made up of multiple computer programs. The analysis processing program PG1 may be executed by a single computer or may be executed by multiple computers (for example, the wearable device 10 and the analysis server 20) working together.
[0045] A computer program including the analysis processing program PG1 is provided by a non-transitory recording medium RM on which the computer program is readably recorded. The recording medium RM is a portable memory such as a CD-ROM, a USB memory, an SD card, a micro SD card, or a CompactFlash (registered trademark). The processing unit 21 reads various computer programs from the recording medium RM using a reading device (not shown) and stores the read various computer programs in the storage unit 22. The computer programs stored in the storage unit 22 may also be provided by communication. In this case, the processing unit 21 acquires the computer programs by communication via the communication unit 23 and stores the acquired computer programs in the storage unit 22.
[0046] The storage unit 22 also stores a task classification model MD1 used in the above-mentioned analysis processing program PG1. The task classification model MD1 is a learning model that has been trained to output information about the actions of the cameraman (i.e., the worker himself) in response to the input of a task video. The task classification model MD1 is described by its definition information. The definition information of the task classification model MD1 includes information about the layers that make up the model, information about the nodes that make up each layer, and parameters such as weighting and bias between nodes. These parameters are obtained by training using, as training data, a dataset that includes a task video and annotation data that indicates the actions of the cameraman shown in the task video. In this embodiment, it is assumed that the trained task classification model MD1 is stored in the storage unit 22.
[0047] The communication unit 23 includes a communication module for wireless communication with an external device such as the wearable device 10. The communication module may be a communication module for wireless communication using a known mobile communication standard such as 3G, 4G, or 5G, or a wireless LAN system such as WiFi (registered trademark). The communication unit 23 communicates with an external device such as the wearable device 10 via the communication network NW, receiving data such as a work video, and transmitting appropriate data to be transmitted to the external device. The communication unit 23 may further include a communication module for short-range wireless communication such as Bluetooth (registered trademark) or ZigBee (registered trademark).
[0048] The operation unit 24 is equipped with operation devices such as a touch panel, a keyboard, and switches, and receives various inputs and operations from a work manager, etc. The processing unit 21 acquires information input through the operation unit 24 and performs appropriate control based on various operation information provided by the operation unit 24.
[0049] The display unit 25 includes a display device such as a liquid crystal monitor or an organic EL (Electro-Luminescence) monitor, and displays information to be notified to a work manager or the like in response to an instruction from the processing unit 21.
[0050] The analysis server 20 may be a single computer, or may be a computer system configured with multiple computers and peripheral devices, etc. Furthermore, the analysis server 20 may be a virtual machine whose entity is virtualized, or may be a cloud.
[0051] The task classification model MD1 used by the analysis server 20 will be described below. FIG. 4 is an explanatory diagram illustrating a method for generating a task classification model MD1. In this embodiment, a tuned model of EgoVLPv2 based on Transformer is used as the task classification model MD1. In EgoVLP (Egocentric Video-Language pre-training), a VLP model MD is constructed by pre-training using a dataset containing first-person perspective video clips and annotations (text narration) for the video clips as training data. EgoClip created by Ego4D is used as the training data for pre-training. Ego4D records the movements (annotations) of the camera wearer in the videos and timestamps for approximately 10,000 videos captured from a first-person perspective. By segmenting the videos based on the timestamps, a dataset (EgoClip) DS consisting of a large number of pairs of video clips and text is obtained.
[0052] The VLP model MD comprises a video encoder En1 that extracts features from input video clips and a text encoder En2 that extracts features from input text. In EgoVLP, the video encoder En1 and the text encoder En2 are trained by contrastive learning using EgoNCE (NCE: Noise Constastive Estimation) as the loss function. In contrastive learning, a self-supervised learning mechanism that compares data is used to learn features so that similar data are close and dissimilar data are far apart.
[0053] The second-generation EgoVLPv2 has a gating mechanism to turn on / off the cross-attention fusion between the video encoder En1 and the text encoder En2, and is characterized by the ability to flexibly switch between dual encoders and fusion encoders. Furthermore, compared to stacking transform layers and shared encoders specialized for fusion, it requires fewer fusion parameters, less GPU memory, less computational resources, and less training time.
[0054] In this embodiment, a task classification model MD1 is generated by tuning EgoVLPv2, an existing model, using a dataset containing work videos captured at a work site using a wearable device 10 and annotations added to the work videos as training data.
[0055] In this embodiment, a method for generating a task classification model MD1 using EgoVLPv2 has been described. However, other learning models, such as VideoMaev2 and CAST, may be used instead of EgoVLPv2, and learning models such as LSTM (Long-Term Memory) and 3D-CNN (Convolutional Neural Network) may also be used as the base model. The training data used to generate the learning model is not limited to the above examples and may be designed appropriately depending on the type of learning model used. Furthermore, instead of tuning the base model, the task classification model MD1 may be generated by training a model with initial parameters set from scratch.
[0056] Furthermore, in this embodiment, a task classification model MD1 is shown that is configured to output task classification results in response to input of a task video, but the task classification model MD1 may also be configured to input audio (acoustic data) recorded together with the task video, or text generated by voice recognition, and output task classification results in response to input of the task video and audio (or task video and text).
[0057] The task classification model MD1 may be generated in the analysis server 20 or in an external server. The generated task classification model MD1 may be stored in the storage unit 22 of the analysis server 20, or may be stored in a storage unit of an external server accessible from the analysis server 20.
[0058] FIG. 5 is an explanatory diagram illustrating a method for annotating a work video. FIG. 5 shows an example in which a label is assigned every two seconds. In this embodiment, the accuracy of labels for periods shorter than two seconds is not an issue, and the minimum unit for assigning labels as annotations is two seconds. In the example of FIG. 5, if the first two-second section is a scene of work A, the work video for that section is labeled "Work A." If the next two-second section includes scenes of work A and work B, it is determined which scene accounts for the majority of the video. If it is determined that scenes of work B account for the majority of the video, the video is labeled "Work B." Similarly, labels are assigned for each two-second section according to the work performed by the worker. In this embodiment, labels such as "work preparation," "visual confirmation," "checking measuring instruments," "disassembly," "repair / replacement," "assembly work," "tool cleaning," "equipment operation check," and "removal of protective covering" are assigned as labels according to the work content.
[0059] In this embodiment, in order to tune EgoVLPv2, a label was assigned to every two seconds of work video using the method illustrated in FIG. 5 to prepare training data. Due to the specifications of the base model, only six seconds of video can be input as training data. Depending on how the video is cut, two or more labels may be assigned to a six-second video. In this embodiment, in order to assign one label to a six-second video, if two or more labels are included, a process of rounding to the majority label was performed.
[0060] The above process results in a dataset containing a large number of videos (videos every 6 seconds) extracted from a series of task videos and labels (text) assigned to each video. The task classification model MD1 according to this embodiment is generated by tuning the trained EgoVLPv2 using the obtained dataset as training data. The task classification model MD1 can be used for various tasks that require a dual encoder and a fusion encoder due to the switching capability of EgoVLPv2. The analysis server 20 according to this embodiment uses the task classification model MD1 to perform a task of recognizing tasks performed by workers. The analysis server 20 recognizes task scenes based on the results of task classification using the task classification model MD1.
[0061] FIG. 6 is an explanatory diagram illustrating a scene recognition method using task classification model MD1. In the operation phase after task classification model MD1 is generated, a worker wears wearable device 10 and uses wearable device 10 to capture images of the worker's work. Analysis server 20 acquires task videos captured by wearable device 10 in real time and inputs the task videos in 6-second increments (180-frame increments for 30 fps video) into task classification model MD1, thereby executing task classification tasks using task classification model MD1 for each 6-second task video. Analysis server 20 obtains classification results as the execution results of the task classification tasks, including "work preparation," "visual confirmation," "measuring instrument confirmation," "dismantling," "repair / replacement," "assembly work," "tool cleaning," "equipment operation confirmation," and "protection removal."
[0062] Based on the results of the work classification task, the analysis server 20 distinguishes between work continuation scenes, which indicate scenes in which work is continuing, and restoration scenes, which indicate scenes in which the work is being restored to its original state, as scenes shown in the work video in 6-second increments. That is, if the work performed by the worker classified by the work classification task is any of "work preparation," "visual confirmation," "instrument confirmation," "dismantling," and "repair / replacement," the analysis server 20 recognizes the scene in the work video as a work continuation scene. Furthermore, if the work performed by the worker classified by the work classification task is "assembly work," "tool cleaning," "equipment operation check," or "protection removal," the analysis server 20 recognizes the scene in the work video as a restoration scene.
[0063] The processing executed by the analysis server 20 will be described below. FIG. 7 is a flowchart illustrating the procedure of processing executed by the analysis server 20 according to the first embodiment. A worker who starts work at a work site wears the wearable device 10 on his or her body. The wearable device 10 automatically starts capturing images at an appropriate timing after the worker wears the wearable device 10. Alternatively, the wearable device 10 starts capturing images when an instruction is given from the worker or the analysis server 20. The wearable device 10 transmits a work video captured by the imaging unit 14 to the analysis server 20 via the communication unit 13. In this embodiment, the wearable device 10 transmits the work video to the analysis server 20 in real time while the work is being performed.
[0064] The processing unit 21 of the analysis server 20 reads out and executes the analysis processing program PG1 from the storage unit 22, and performs the analysis processing according to the following procedure.
[0065] Processing unit 21 acquires the work video transmitted from wearable device 10 via communication unit 23 (step S101). It is assumed that processing unit 21 sequentially acquires the work video continuously transmitted from wearable device 10.
[0066] The processing unit 21 inputs the acquired task video into the task classification model MD1 in predetermined units (for example, every 6 seconds) and executes calculations (task classification tasks) using the task classification model MD1 (step S102). The processing unit 21 acquires task classification results as the execution results of the calculations using the task classification model MD1 (step S103).
[0067] The processing unit 21 recognizes a work scene based on the classification result by the work classification model MD1 (step S104). In this embodiment, it is sufficient to be able to distinguish and recognize work continuation scenes and restoration scenes. If the classification result by the work classification model MD1 is any of "work preparation," "visual confirmation," "measuring instrument confirmation," "dismantling," and "repair / replacement," the processing unit 21 recognizes it as a work continuation scene, and if the classification result by the work classification model MD1 is any of "assembly work," "tool putting away," "operation confirmation of facility equipment," and "removal of protection," the processing unit 21 recognizes it as a restoration scene.
[0068] The processing unit 21 determines whether the work scene recognized in step S104 is a restoration scene (step S105). If it is determined that the work scene is not a restoration scene (S105: NO), the processing unit 21 returns the process to step S102 because the worker is continuing work.
[0069] If it is determined that the scene is a restoration scene (S105: YES), processing unit 21 determines that the work by the worker has finished (step S106) and instructs wearable device 10 to end capturing the work video (step S107). The instruction to end capturing the work video is transmitted to wearable device 10 via communication unit 23. Upon receiving the instruction to end capturing from analysis server 20, wearable device 10 ends capturing by imaging unit 14. Furthermore, analysis server 20 may instruct wearable device 10 to end capturing and also instruct wearable device 10 to notify the user that capturing has ended. In this case, information that capturing has ended is output as sound from sound output unit 16 to notify the worker. Analysis server 20 may transmit a system shutdown instruction to wearable device 10 instead of the instruction to end capturing.
[0070] As described above, the analysis server 20 according to the first embodiment determines that the worker has completed the work when it detects a scene of restoring the work to its original state based on the work video acquired from the wearable device 10. When it determines that the worker has completed the work, the analysis server 20 can instruct the wearable device 10 to stop capturing images, thereby eliminating unnecessary video uploads due to forgetting to perform an operation.
[0071] (Embodiment 2) In the second embodiment, a configuration will be described in which a list of work items to be performed by a worker is prepared in advance, and whether or not the work by the worker has been completed is determined by checking against the work item list.
[0072] FIG. 8 is a conceptual diagram showing an example of a work item list. The work item list lists the work items to be performed by a worker in the order in which they are performed, and records, for example, the work order, work content, and confirmation flag in association with each other. Here, the confirmation flag is a flag that indicates whether or not the worker has performed the work for each work item. FIG. 8 shows an example of a work list in which the work begins with work preparation, progresses through visual inspection, disassembly, repair / replacement work, and ends with assembly work. Furthermore, by referring to the confirmation flag column, it can be seen that the worker has performed the work from work preparation to repair / replacement work.
[0073] The work item list is created, for example, by the terminal of the work manager. Before the worker starts work, the analysis server 20 acquires the work item list from the terminal of the work manager and stores it in the storage unit 22.
[0074] After acquiring the work item list, the analysis server 20 executes the following process. 9 is a flowchart illustrating the procedure of processing executed by the analysis server 20 according to embodiment 2. In the same procedure as in embodiment 1, the work performed by the worker is captured by the wearable device 10, and the work video is transmitted to the analysis server 20.
[0075] Processing unit 21 acquires the work video transmitted from wearable device 10 via communication unit 23 (step S201). It is assumed that processing unit 21 sequentially acquires the work video continuously transmitted from wearable device 10.
[0076] The processing unit 21 inputs the acquired task video into the task classification model MD1 in predetermined units (for example, every 6 seconds) and executes calculations (task classification tasks) using the task classification model MD1 (step S202). The processing unit 21 acquires task classification results as the execution results of the calculations using the task classification model MD1 (step S203). The processing unit 21 identifies the tasks performed by the worker based on the classification results using the task classification model MD1 and checks the confirmation flags in the task item list (step S204).
[0077] The processing unit 21 checks the confirmation flag column of the work item list to identify multiple work items performed by the worker, and determines whether the work by the worker has been completed by comparing them with the work items included in the work item list (step S205). In this embodiment, the processing unit 21 determines that the work by the worker has been completed when it can confirm from the confirmation flag that the worker has performed the work to restore the original state (assembly work in the work item list of FIG. 8).
[0078] If it is determined that the work has not been completed (S205: NO), the processing unit 21 returns the process to step S202.
[0079] If it is determined that the work has been completed (S205: YES), the processing unit 21 determines whether any work has been omitted (step S206). If the processing unit 21 determines that the work has been completed but that there is a work item in the work item list whose confirmation flag is not checked, it determines that there is any work that has been omitted.
[0080] If it is determined that there is an omitted task (S206: YES), processing unit 21 notifies the worker of the omitted task (step S207). Processing unit 21 transmits information that there is an omitted task to wearable device 10 via communication unit 23. At this time, processing unit 21 may also transmit information about the task item for which the omitted task has occurred to wearable device 10. Based on the information transmitted from analysis server 20, wearable device 10 notifies the worker of the omitted task information by voice from sound output unit 16, for example.
[0081] If it is determined that there is no work omission (S206: NO), the processing unit 21 ends the processing according to this flowchart.
[0082] If the tasks are being performed in a predetermined order, processing unit 21 may check the task items in the task order defined in the task item list when checking them against the task item list in step S205. In this case, processing unit 21 can determine whether any tasks have been omitted each time it checks the task items. On the other hand, if the tasks are being performed in an order different from the predetermined order, processing unit 21 may check the task items in any order when checking them against the task item list in step S205. In this case, processing unit 21 may determine whether any tasks have been omitted after determining that the tasks have been completed.
[0083] As described above, in the second embodiment, it is possible to determine whether a task has been completed and whether any task has been omitted by using a task item list prepared in advance.
[0084] (Embodiment 3) In the third embodiment, a configuration for determining whether or not the wearable device 10 has completed the task will be described.
[0085] The analysis processing program PG1 and task classification model MD1 described in embodiment 1 are stored in the storage unit 12 of the wearable device 10. The processing unit 11 of the wearable device 10 reads out and executes the analysis processing program PG1 from the storage unit 12, thereby performing the following processing.
[0086] 10 is a flowchart illustrating the procedure of processing executed by wearable device 10 according to embodiment 3. In the same procedure as in embodiment 1, the work performed by the worker is captured by wearable device 10. Processing unit 11 of wearable device 10 acquires a work video from imaging unit 14 (step S301). It is assumed that processing unit 11 sequentially acquires the work video continuously output from imaging unit 14.
[0087] The processing unit 11 inputs the acquired task video into the task classification model MD1 in predetermined units (for example, every 6 seconds) and executes calculations (task classification tasks) using the task classification model MD1 (step S302). The processing unit 11 acquires task classification results as the execution results of the calculations using the task classification model MD1 (step S303).
[0088] The processing unit 11 recognizes a work scene based on the classification result by the work classification model MD1 (step S304). If the classification result by the work classification model MD1 is any of "work preparation," "visual confirmation," "measuring instrument confirmation," "dismantling," and "repair / replacement," the processing unit 11 recognizes it as a work continuation scene, and if the classification result is any of "assembly work," "tool cleaning up," "operational confirmation of facility equipment," and "removal of protection," the processing unit 11 recognizes it as a restoration scene.
[0089] The processing unit 11 determines whether the work scene recognized in step S304 is a restoration scene (step S305). If it is determined that the work scene is not a restoration scene (S305: NO), the processing unit 11 returns the process to step S302 because the worker is continuing work.
[0090] If it is determined that the scene is a restoration scene (S305: YES), the processing unit 21 determines that the work by the worker has finished (step S306) and ends image capture by the image capture unit 14 (step S307). At this time, the processing unit 11 may notify the worker by outputting a sound from the sound output unit 16 that the image capture will end. Furthermore, if the processing unit 11 receives an image capture end instruction through the operation unit 18 without detecting a restoration scene in step 305, it may output a sound from the sound output unit 16 inquiring as to whether the work of restoring the scene to the original state has been performed.
[0091] After the shooting is completed, the processing unit 11 uploads the work video shot from the start to the end of the work to a server device (step S308). In the third embodiment, since analysis by the analysis server 20 is not performed, there is no need to upload the work video to the analysis server 20, and the work video can be uploaded to a predetermined server device that accumulates the work video as evidence.
[0092] As described above, wearable device 10 according to the third embodiment determines that the worker has completed the work when it detects a scene of restoring the work to its original state based on the work video captured by imaging unit 14. When wearable device 10 determines that the worker has completed the work, it can upload the work videos acquired from the start to the end of the work to a server device.
[0093] Furthermore, as an application example of the information processing system according to this embodiment, a work scene involving air conditioning equipment has been envisaged, but it goes without saying that the system can be applied to a variety of work scenes, not limited to work scenes involving air conditioning equipment, such as work scenes relating to elevators, and work scenes relating to maintenance and inspection at chemical plants and power supply facilities.
[0094] The embodiments disclosed herein should be considered in all respects as illustrative and not restrictive. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]
[0095] 10. Wearable Devices 20 Analysis Server 21 Processing section 22 Memory section 23 Communications Department 24 Control section 25 Display section PG1 Analysis and Processing Program MD1 Task Classification Model
Claims
1. Acquire a work video obtained by capturing the work performed by the worker, Based on the acquired work video, scenes in which assembly work, checking the operation of equipment or removing protective coverings are being carried out are detected as restoration scenes, distinguishing them from work continuation scenes in which the worker continues work; When the scene of restoring the original state is detected, it is determined that the work related to the scene of continuing the work has been completed. A computer program for causing at least one computer to execute a process.
2. When it is determined that the work related to the work continuation scene has been completed, the work video is transmitted to an external device.
2. The computer program product according to claim 1, for causing the computer to execute a process.
3. Determine whether or not there is any work omission for each work content based on the work video, If it is determined that there is work omission, it will notify you of the work omission.
2. The computer program product according to claim 1, for causing the computer to execute a process.
4. The work video is a video obtained by capturing images of repair or inspection work on equipment.
2. The computer program of claim 1.
5. The work video is a video captured by a wearable device worn by the worker.
2. The computer program of claim 1.
6. When the scene of restoration to the original state is detected, a notice to end the imaging is sent.
2. The computer program product according to claim 1, for causing the computer to execute a process.
7. If the restoration scene is not detected and an instruction to end the image capture is given, a notification is output inquiring whether the restoration work has been carried out.
2. The computer program product according to claim 1, for causing the computer to execute a process.
8. Identifying a plurality of work items performed by the worker in the work continuation scene based on the acquired work video; obtaining a list of a plurality of work items to be performed regarding the work related to the work continuation scene; By comparing the identified work items with the work items included in the acquired list, it is determined whether the work items included in the list have been completed.
2. The computer program product according to claim 1, for causing the computer to execute a process.
9. The identified plurality of work items are compared with the plurality of work items included in the acquired list in no particular order.
9. A computer program product according to claim 8, for causing the computer to execute a process.
10. By checking the work items, we can determine whether any work has been missed, If it is determined that there is work omission, it will notify you of the work omission.
9. A computer program product according to claim 8, for causing the computer to execute a process.
11. Acquire a work video obtained by capturing the work performed by the worker, Based on the acquired work video, scenes in which assembly work, checking the operation of equipment or removing protective coverings are being carried out are detected as restoration scenes, distinguishing them from work continuation scenes in which the worker continues work; When the scene of restoring the original state is detected, it is determined that the work related to the scene of continuing the work has been completed. An information processing method in which processing is carried out by at least one computer.
12. At least one processing unit; The processing unit Acquire a work video obtained by capturing the work performed by the worker, Based on the acquired work video, scenes in which assembly work, checking the operation of equipment or removing protective coverings are being carried out are detected as restoration scenes, distinguishing them from work continuation scenes in which the worker continues work; When the scene of restoring the original state is detected, it is determined that the work related to the scene of continuing the work has been completed. Information processing system.
Citation Information
Patent Citations
Production result collection system
JP2017120454A
Imaging system, imaging device and imaging method
JP2019032593A
Image processing device, image processing method, and program
JP2019191117A
Work state analysis system and work state analysis method
JP2023019009A