Video aggregation device, video aggregation method, and video aggregation program

The video aggregation device analyzes worker actions and conditions to efficiently generate and view scenes with varying reliability, addressing the inefficiencies in existing systems by focusing on work cycle time, specific actions, and error logs.

JP7775753B2Active Publication Date: 2025-11-26OMRON CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022038606
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-11
Publication Date
2025-11-26
Estimated Expiration
2042-03-11

AI Technical Summary

Technical Problem

Existing systems fail to efficiently aggregate video images of multiple desired scenes with varying reliability thresholds, making it difficult for work managers to view relevant work scenarios.

Method used

A video aggregation device and method that includes an acquisition unit, detection unit, determination unit, and generation unit to analyze time series data of a worker's skeleton or body parts, determine scene conditions, and generate aggregated videos based on set criteria such as work cycle time, specific actions, and error logs.

Benefits of technology

Enables efficient generation and viewing of multiple types of scenes, allowing managers to focus on critical work activities and errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007775753000001
    Figure 0007775753000001
  • Figure 0007775753000002
    Figure 0007775753000002
  • Figure 0007775753000003
    Figure 0007775753000003
Patent Text Reader

Abstract

To improve work recognition accuracy.SOLUTION: A moving image integration device includes: an acquisition unit configured to acquire moving images in which work by a worker is imaged; a detection unit configured to detect time-series data of detection information relating to the skeleton or parts of the worker based on the moving images; a determination unit configured to determine, for each of a plurality of types of object scenes being extracted, based on the detected time-series data of the detection information, whether the work meets a condition corresponding to the object scenes being extracted; and a generation unit configured to generate a moving image obtained by integrating partial moving images including points in time when it was determined that the work met the condition corresponding to the object scenes being extracted, based on an extraction count or extraction time that is set for each of the plurality of types of object scenes being extracted.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosed technology relates to a video aggregating device, a video aggregating method, and a video aggregating program. [Background technology]

[0002] Patent Document 1 discloses a work analysis system comprising: a time information acquisition unit that acquires time information; a work video acquisition unit that films the work state of a worker and acquires a work video; a work information acquisition unit that acquires work information for estimating the work of the worker; a work estimation unit that estimates the work of the worker based on the work information, calculates a reliability indicating the certainty of the estimated work, and calculates a start time and an end time for each estimated work based on the time information; a work linking unit that divides the work video at the start time and end time of the estimated work and links the estimated work with the reliability of the work; a confirmation information output unit that outputs confirmation information to allow a user to determine whether the reliability is below a threshold; an input unit that accepts instruction input by a user; and a video playback unit that plays back section videos for which the reliability is below the threshold based on the instruction input by the input unit. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2020-91801 Summary of the Invention [Problem to be solved by the invention]

[0004] When a work manager views video images of a work, he or she may want to check video images of a plurality of different scenes.

[0005] However, in the technology described in Patent Document 1, video images of sections where the reliability of the work is low are displayed, and therefore it is not possible to efficiently view video images of multiple types of desired scenes.

[0006] The disclosed technology has been developed in consideration of the above points, and aims to provide a video aggregation device, a video aggregation method, and a video aggregation program that can generate video images for efficiently viewing multiple types of scenes. [Means for solving the problem]

[0007] A first aspect of the disclosure is a video aggregation device that includes an acquisition unit that acquires video images of a worker working, a detection unit that detects time series data of detection information related to the worker's skeleton or body parts based on the video images, a determination unit that determines, for each of a plurality of types of scenes to be cut out, based on the time series data of the detected detection information, whether the work satisfies the conditions corresponding to the scenes to be cut out, and a generation unit that generates a video that aggregates video images of portions including the time points at which the work is determined to satisfy the conditions corresponding to the scenes to be cut out, based on the number of cuts or cut-out times set for each of the plurality of types of scenes to be cut out.

[0008] In the first aspect, the scenes to be extracted may include scenes in which the work cycle time is equal to or greater than a threshold value, and the determination unit may analyze the work cycle time for each work cycle for scenes in which the work cycle time is equal to or greater than the threshold value based on the time series data of the detected detection information, and determine that the work satisfies the condition if the work cycle time is equal to or greater than the threshold value.

[0009] In the first aspect, the scene to be extracted may include a scene in which a worker performs a specific action, and the determination unit may determine that the scene in which the worker performs the specific action satisfies the condition when the worker moves to a position corresponding to the location where the specific action is performed, based on time series data of the detection information detected.

[0010] In the first aspect, the specific action may be to place the defective product in a defective product storage area.

[0011] In the first aspect, the scene to be extracted may include a scene in which an error log related to equipment used in the work has occurred, and the determination unit may further determine that the work satisfies the condition if, for a scene in which an error log has occurred, the log related to equipment used in the work is an error log.

[0012] A second aspect of the disclosure is a video aggregation method, in which an acquisition unit acquires video images of a worker working, a detection unit detects time series data of detection information related to the worker's skeleton or body parts based on the video images, a determination unit determines, for each of a plurality of types of scenes to be cut out, based on the time series data of the detected detection information, whether the work satisfies the conditions corresponding to the scenes to be cut out, and a generation unit generates a video that aggregates video images of portions including the time points at which the work is determined to satisfy the conditions corresponding to the scenes to be cut out, based on the number of cut-outs or cut-out times set for each of the plurality of types of scenes to be cut out.

[0013] A third aspect of the disclosure is a video aggregation program that causes a computer to acquire video footage of a worker working, detect time series data of detection information related to the worker's skeleton or body parts based on the video footage, determine, for each of a plurality of types of scenes to be cut out, based on the detected time series data of the detection information, whether the work satisfies the conditions corresponding to the scenes to be cut out, and generate a video that aggregates video footage of portions including the time points at which the work is determined to satisfy the conditions corresponding to the scenes to be cut out, based on the number of cuts or cut-out times set for each of the plurality of types of scenes to be cut out. [Effects of the Invention]

[0014] According to the disclosed technology, it is possible to generate moving images for efficiently viewing a plurality of types of scenes. [Brief explanation of the drawings]

[0015] [Figure 1] FIG. 1 is a configuration diagram of a video aggregation system. [Figure 2] FIG. 2 is a configuration diagram showing a hardware configuration of the video aggregation device. [Figure 3] FIG. 2 is a functional block diagram of the video aggregation device. [Figure 4] FIG. 2 is a functional block diagram of a determination unit of the video aggregation device. [Figure 5] FIG. 10 is a diagram for explaining a method for detecting a work cycle time. [Figure 6] 10A and 10B are diagrams for explaining a method for recognizing an action of a worker placing a defective product in a defective product storage area. [Figure 7] FIG. 2 is a functional block diagram of a generation unit of the video aggregation device. [Figure 8] 10 is a flowchart of a video aggregation process. [Figure 9] 10 is a flowchart of a video aggregation process. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of the present invention will be described below with reference to the drawings. The same or equivalent components and parts in each drawing are designated by the same reference numerals. The dimensional proportions of the drawings may be exaggerated for the sake of explanation and may differ from the actual proportions.

[0017] 1 shows the configuration of a video aggregation system 10. The video aggregation system 10 includes a video aggregation device 20 and a camera 30.

[0018] The video image aggregation device 20 aggregates videos showing the work performed by the worker W based on the videos captured by the camera 30.

[0019] As an example, a worker W performs a predetermined task using equipment M on a workbench T. The workbench T is installed in a location with sufficient brightness to allow the worker's actions to be recognized. If a defective product is produced as a result of the work, the worker W places the defective product in a defective product storage area S.

[0020] The camera 30 captures, for example, RGB color images and outputs them to the video image aggregation device 20. The camera 30 is installed in a position where the work being done by the worker W can be easily recognized. Specifically, the camera 30 is installed in a position that satisfies conditions such as a position where the work being done by the worker W is not hidden by a workbench T or the like, and a position where the worker W who has moved in front of the defective product storage area S is not hidden by other objects or the like. In this embodiment, as an example, a case will be described where the camera 30 is installed in a position where at least the upper body of the worker W can be viewed from diagonally above.

[0021] In this embodiment, a case where one camera 30 is provided will be described, but a configuration may be adopted in which multiple cameras 30 are provided. In addition, in this embodiment, a case where there is one worker W will be described, but there may be two or more workers W.

[0022] The device M used for the work outputs a log related to the use of the device M, including an error log, to the video aggregation device 20. When an error occurs, the device M outputs the error log to the video aggregation device 20.

[0023] Fig. 2 is a block diagram showing the hardware configuration of the video aggregation device 20 according to this embodiment. As shown in Fig. 2, the video aggregation device 20 includes a controller 21. The controller 21 is configured as a device including a general computer.

[0024] 2, the controller 21 includes a central processing unit (CPU) 21A, a read-only memory (ROM) 21B, a random access memory (RAM) 21C, and an input / output interface (I / O) 21D. The CPU 21A, the ROM 21B, the RAM 21C, and the I / O 21D are connected to each other via a bus 21E. The bus 21E includes a control bus, an address bus, and a data bus.

[0025] Furthermore, an operation unit 22, a display unit 23, a communication unit 24, and a storage unit 25 are connected to the I / O 21D.

[0026] The operation unit 22 includes, for example, a mouse and a keyboard.

[0027] The display unit 23 is configured by, for example, a liquid crystal display.

[0028] The communication unit 24 is an interface for performing data communication with an external device such as a camera 30 .

[0029] The storage unit 25 is configured with a non-volatile external storage device such as a hard disk, etc. As shown in Fig. 2, the storage unit 25 stores a video aggregation program 25A, video data 25B which is video images captured by the camera 30 and video images extracted from the camera 30, a log 25C output from the device M, etc.

[0030] The CPU 21A is an example of a computer. The term "computer" here refers to a processor in a broad sense, and includes a general-purpose processor (e.g., a CPU) or a dedicated processor (e.g., a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device, etc.).

[0031] The video aggregation program 25A may be realized by being stored in a non-volatile non-transitory recording medium or distributed via a network and installed in the video aggregation device 20 as appropriate.

[0032] Examples of non-volatile non-transitory recording media include CD-ROMs (Compact Disc Read Only Memory), magneto-optical disks, HDDs (Hard Disk Drives), DVD-ROMs (Digital Versatile Disc Read Only Memory), flash memory, memory cards, etc.

[0033] Fig. 3 is a block diagram showing the functional configuration of the CPU 21A of the video aggregation device 20. As shown in Fig. 3, the CPU 21A functionally includes a setting unit 40, an acquisition unit 41, a detection unit 42, a determination unit 43, a generation unit 44, and an output unit 45. The CPU 21A functions as each functional unit by reading and executing a video aggregation program 25A stored in the storage unit 25.

[0034] The setting unit 40 receives settings for the number of cutouts or the cutout time for each of a plurality of types of cutout target scenes.

[0035] For example, the multiple types of scenes to be extracted include a scene in which the work cycle time is equal to or exceeds a threshold value that is a standard work time, a scene in which a defective product is placed in a defective product storage area, and a scene in which an error log for device M is generated.

[0036] In addition, on the setting screen displayed on the display unit 23, by operating the operation unit 22, settings are accepted for each of the multiple types of scenes to be cut out, such as the number of cut-outs indicating how many times the scene to be cut out is to be cut out, or the cut-out time indicating how many minutes the scene to be cut out is to be cut out, and the standard working time.

[0037] Furthermore, on the setting screen displayed on the display unit 23, settings are accepted by operating the operation unit 22 as to whether a new video is to be acquired from the camera 30 and whether a video is to be cut out from a saved video.

[0038] The acquisition unit 41 acquires from the camera 30 moving images of the worker W working, and stores the moving image data 25B in the storage unit 25.

[0039] Furthermore, the acquisition unit 41 acquires a log from the device M and stores it in the log 25C of the storage unit 25.

[0040] The detector 42 detects time-series data of detection information relating to the worker W's body parts or skeleton based on the moving images acquired from the camera 30.

[0041] Specifically, the detection information regarding the body part includes, for example, the coordinates of the four corners of a bounding box that represents an area including a specific body part (at least one of the right and left hands). Here, a bounding box refers to a rectangular shape such as a rectangle or a square that circumscribes the object to be detected. Specifically, the reliability of the object to be detected is calculated for each anchor box (rectangular region) of multiple sizes. Then, the coordinates of the four corners of the anchor box with the highest reliability are set as the coordinates of the four corners of the bounding box. Such a bounding box detection method can be, for example, a known method such as Faster R-CNN (Regions with Convolutional Neural Networks), and for example, the method described in Reference 1 below can be used.

[0042] (Reference 1) "Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks", Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun.

[0043] As a method for detecting detection information related to a body part based on a moving image, a learning model that uses an image as input and detects information related to a body part as output can be used, which is a trained detection model that has been trained using a large number of images as training data. As a training method for obtaining such a trained detection model, a known method such as CNN can be used, and for example, the method described in Reference 2 below can be used.

[0044] (Reference 2) "Understanding Human Hands in Contact at Internet Scale", pp.9869-9878, Dandan Shan1, Jiaqi Geng, Michelle Shu, David F. Fouhey, University of Michigan, Johns Hopkins University, CVPR2020.

[0045] The detected information about the skeleton also includes coordinates of feature points such as body parts and joints of the worker W, and link information that defines links connecting each feature point. For example, the feature points include facial parts such as the eyes and nose of the worker W, and joints such as the neck, shoulders, elbows, wrists, waist, knees, and ankles.

[0046] As a method for detecting detection information related to the skeleton based on an image, a learning model that uses an image as input and skeletal detection information as output can be used, which is a trained detection model trained using a large number of images as training data. As a training method for obtaining such a trained detection model, a known method such as CNN (Regions with Convolutional Neural Networks) can be used, and for example, the method described in Reference 3 below can be used.

[0047] (Reference 3) "OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields", Zhe Cao, Student Member, IEEE, Gines Hidalgo, Student Member, IEEE, Tomas Simon, Shih-En Wei, and Yaser Sheikh, IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE.

[0048] The determination unit 43 determines, for each of a plurality of types of scenes to be extracted, whether the work satisfies the conditions corresponding to the scene to be extracted, based on the time series data of the detected detection information and the time series data of the acquired log.

[0049] Specifically, as shown in FIG. 4, the determination unit 43 includes a period detection unit 50, a time determination unit 51, an action recognition unit 52, an action determination unit 53, and a log determination unit .

[0050] The cycle detection unit 50 analyzes the start time and end time of each work cycle based on the time series data of the detection information related to the detected body parts and the time series data of the detection information related to the skeleton, and detects the duration of the work cycle.

[0051] Specifically, as shown in Figure 5, periodically appearing movements (signals) are automatically detected using DTW (Dynamic Time Warping) from time series data of movement features extracted based on time series data of detection information related to body parts and time series data of detection information related to skeletons, thereby detecting the start and end times of a work cycle and detecting the duration of the work cycle. Figure 5 above shows an example in which the start time (2 minutes 4 seconds) and end time (2 minutes 54 seconds) of a work cycle are detected, and the duration of the work cycle (50 seconds) is detected.

[0052] As for the period estimation method using DTW, the same method as in Reference 4 can be used, and therefore a detailed description thereof will be omitted.

[0053] (Reference 4) Yasuo Namioka et al., "Method for automatic measurement of cycle time for repetitive tasks using wearable sensors," Internet search <URL:https: / / www.global.toshiba / content / dam / toshiba / migration / corp / techReviewAssets / tech / review / 2018 / 03 / 73_03pdf / a12.pdf>

[0054] In the above description, a case has been described in which periodically occurring movements are automatically detected using DTW from time-series data of movement features extracted based on time-series data of detection information related to body parts and time-series data of detection information related to skeletons, but the present invention is not limited to this. Periodically occurring movements may also be automatically detected using DTW from time-series data of detection information related to body parts or time-series data of detection information related to skeletons.

[0055] Based on the time of the work cycle detected for each work cycle, if the time of the work cycle is equal to or greater than a threshold value, the time determination unit 51 determines that the work satisfies the conditions corresponding to the scene in which the time of the work cycle is equal to or greater than the threshold value, and records the start time and end time of the work cycle.

[0056] The action recognition unit 52 recognizes the action of the worker W placing the defective product in the defective product storage area S based on the time series data of the detection information related to the detected body part or the time series data of the detection information related to the skeleton.

[0057] Specifically, the action of the worker W placing the defective product in the defective product storage area S is recognized based on whether or not the worker W has moved to a position corresponding to the defective product storage area S.

[0058] For example, as shown in Figure 6, if the coordinates (x, y) = (50, 50) of either the right or left hand are within the rectangular range defined by the upper left coordinates (x, y) = (20, 20) and the lower right coordinates (x, y) = (150, 150) of the area of ​​the defective product storage area S, it is recognized that the hand is in the defective product storage area S and that worker W is performing the action of placing a defective product in the defective product storage area S.

[0059] Alternatively, if the head coordinates (x, y) = (250, 300) are within the rectangular range defined by the upper left coordinates (x, y) = (200, 200) and the lower right coordinates (x, y) = (500, 500) of the area in front of the defective product storage area S, it is determined that a worker is in front of the defective product storage area S, and it is recognized that worker W is placing a defective product in the defective product storage area S.

[0060] In addition, a pre-trained model may be used based on time series data of detection information regarding the detected body part or time series data of detection information regarding the skeleton to recognize the action of worker W placing the defective product in the defective product storage area S.

[0061] When the action determination unit 53 recognizes the action of worker W placing a defective product in the defective product storage area S, it determines that the action satisfies the conditions corresponding to the scene in which worker W performs the action of placing a defective product in the defective product storage area S, and records the time.

[0062] The log determination unit 54 determines whether or not the log is an error log based on the time series data of the log related to the device M, and if it is an error log, determines that it is an operation that satisfies the conditions corresponding to the scene in which the error log occurred and records the time.

[0063] The generation unit 44 cuts out the video image of the portion including the point in time when it is determined that the work satisfies the conditions corresponding to the scene to be cut out, based on the number of cut-outs or the cut-out time set for each of the multiple types of scenes to be cut out, and generates a video image that aggregates the cut-out videos.

[0064] Specifically, the generation unit 44 includes a video clipping unit 60 and a video selecting unit 61, as shown in FIG.

[0065] For each of the plurality of types of scenes to be extracted, the video extracting unit 60 extracts a portion of the video including a point in time when it is determined that the activity satisfies the condition corresponding to the scene to be extracted.

[0066] The video selection unit 61 selects videos cut out for the scenes to be cut out based on the number of cut-outs or cut-out times set for each of the multiple types of scenes to be cut out, and generates a video that aggregates the selected videos.

[0067] For example, if the number of cutouts is set to four for a scene in which the work cycle time is equal to or greater than a threshold, four cycles of video cut out from the start time to the end time of the work cycle that is determined to be work that satisfies the conditions corresponding to the scene in which the work cycle time is equal to or greater than the threshold are selected, and an aggregated video is generated by combining the four selected cycles of video.

[0068] In addition, if the cut-out time is set to 4 minutes for a scene in which the work cycle time is equal to or greater than the threshold, video cut out from the start time to the end time of the work cycle that is determined to be work that satisfies the conditions corresponding to the scene in which the work cycle time is equal to or greater than the threshold is selected within a range not exceeding 4 minutes, and an aggregated video is generated by combining the selected video.

[0069] The output unit 45 outputs the collected moving images by displaying them on the display unit 23 or storing them in the storage unit 25.

[0070] Next, the video aggregation process executed by the CPU 21A of the video aggregation device 20 will be described with reference to the flowchart shown in FIG.

[0071] In step S100, CPU 21A accepts, for each of a plurality of types of cut-out target scenes, a cut-out count indicating how many times the cut-out target scene is to be cut out, or a cut-out time indicating how many minutes the cut-out target scene is to be cut out, and a standard working time setting, on the setting screen displayed on display unit 23. CPU 21A also accepts, on the setting screen displayed on display unit 23, a setting regarding whether to acquire new moving images from camera 30 and whether to cut out from stored moving images. Note that the setting in step S100 does not have to be accepted every time a moving image aggregation process is performed, and may be accepted periodically (for example, once a month).

[0072] In step S102, CPU 21A determines whether or not to acquire new moving images. If it is set that new moving images are to be acquired from camera 30, the process proceeds to step S104. On the other hand, if it is set that new moving images are not to be acquired from camera 30, the process proceeds to step S108.

[0073] In step S104, the CPU 21A acquires a moving image of the work of the worker W from the camera 30, and also acquires time-series data of the log related to the device M.

[0074] In step S106, the CPU 21A stores the acquired time-series data of the video and the log in the storage unit 25.

[0075] In step S108, CPU 21A determines whether to cut out from the stored video. If it is set to cut out from the stored video, the process proceeds to step S110. On the other hand, if it is set not to cut out from the stored video, the process proceeds to step S126.

[0076] In step S110, the CPU 21A acquires from the storage unit 25 moving images that were captured in the past.

[0077] In step S111, the CPU 21A detects time-series data of detection information relating to the body parts or skeleton of the worker W based on the moving images acquired in step S104 or step S110.

[0078] In step S112, the CPU 21A analyzes the start time and end time of each work cycle based on the time series data of the detection information related to the detected body parts and the time series data of the detection information related to the skeleton, and detects the duration of the work cycle.

[0079] In step S114, the CPU 21A recognizes the action of the worker W placing the defective product in the defective product storage area S based on the time series data of the detection information related to the detected body part or the time series data of the detection information related to the skeleton.

[0080] In step S116, the CPU 21A acquires time-series data of the log related to the device M from the storage unit 25, and determines whether or not it is an error log.

[0081] In step S118, CPU 21A determines, for each of the multiple types of scenes to be extracted, whether the work satisfies the conditions corresponding to the scene to be extracted. Specifically, based on the time of the work cycle detected for each work cycle, if the time of the work cycle is equal to or greater than a threshold, CPU 21A determines that the work satisfies the conditions corresponding to the scene in which the time of the work cycle is equal to or greater than a threshold, and records the start time and end time of the work cycle. If CPU 21A recognizes the action of worker W placing a defective product in the defective product storage area S, CPU 21A determines that the work satisfies the conditions corresponding to the scene in which worker W places a defective product in the defective product storage area S, and records the time. If the log is an error log, CPU 21A determines that the work satisfies the conditions corresponding to the scene in which the error log occurred, and records the time.

[0082] In step S120, CPU 21A cuts out, for each of the plurality of types of cut-out target scenes, a portion of the moving image including a point in time at which it is determined that the activity satisfies the condition corresponding to the cut-out target scene.

[0083] In step S122, CPU 21A stores in storage unit 25 the moving images cut out for each of the plurality of types of scenes to be cut out.

[0084] In step S124, CPU 21A selects moving images cut out for the plurality of types of cut-out target scenes based on the cut-out count or cut-out time set for each of the plurality of types of cut-out target scenes, and generates a moving image that aggregates the selected moving images.

[0085] In step S126, CPU 21A outputs the moving image generated in step S124 by displaying it on display unit 23 or storing it in storage unit 25.

[0086] In this manner, in this embodiment, a video is generated by aggregating the video of the portion including the time point at which the activity is determined to satisfy the conditions corresponding to the scene to be extracted, based on the number of extractions or the extraction time set for each of the multiple types of scenes to be extracted. This makes it possible to generate a video for efficiently viewing multiple types of scenes.

[0087] The above-described embodiment merely exemplifies the configuration of the present invention, and the present invention is not limited to the specific embodiment described above, and various modifications are possible within the scope of the technical concept thereof.

[0088] For example, the multiple types of scenes to be extracted include a scene in which the work cycle time is equal to or exceeds a threshold value that is a standard work time, a scene in which a defective product is placed in a defective product storage area, and a scene in which an error log is generated for device M, but the present invention is not limited to this. The scenes to be extracted may be other types of scenes. The scenes to be extracted may also be scenes related to good work.

[0089] Furthermore, although the specific action of the scene to be extracted is the action of placing a defective product in a defective product storage area, the present invention is not limited to this example. Actions other than the action of placing a defective product in a defective product storage area may also be the specific action of the scene to be extracted.

[0090] Furthermore, the video aggregation process executed by the CPU in the above embodiment by reading software (programs) may be executed by various processors other than the CPU. Examples of processors in this case include dedicated electrical circuits, such as programmable logic devices (PLDs) (such as field-programmable gate arrays (FPGAs)) whose circuit configuration can be changed after manufacture, and application-specific integrated circuits (ASICs) that are processors with circuit configurations specifically designed to execute recognition processes. The video aggregation process may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (e.g., multiple FPGAs, or a combination of a CPU and an FPGA). The hardware structure of these various processors is, more specifically, an electrical circuit that combines circuit elements such as semiconductor devices. [Explanation of symbols]

[0091] 10 Video Aggregation System 20 Video Aggregation Device 22 Control section 23 Display section 24 Communications Department 25 Memory section 25A Video Aggregation Program 25B Video data 25C Log 30 Camera 40 Setting section 41 Acquisition Department 42 Detection unit 43 Judgment section 44 Generation part 45 Output section 50 Period detection unit 51 Time determination section 52 Motion recognition section 53 Operation judgment section 54 Log determination unit 60 Video clipping section 61 Video selection section M equipment S Defective goods storage area W Worker

Claims

1. an acquisition unit that acquires video images of a worker's work; a detection unit that detects time-series data of detection information related to a skeleton or a body part of the worker based on the moving image; a determination unit that determines, for each of a plurality of types of scenes to be extracted, whether the scene satisfies a condition corresponding to the scene to be extracted, based on time-series data of the detected detection information; a generation unit that generates a video that aggregates video images of portions including time points determined to be activities that satisfy conditions corresponding to the scenes to be extracted, based on the number of extractions or the extraction times set for each of the plurality of types of scenes to be extracted; A video aggregation device including:

2. The scene to be extracted includes a scene in which the time of a work cycle is equal to or greater than a threshold value, The video aggregation device of claim 1, wherein the determination unit analyzes the work cycle time for each work cycle based on the time series data of the detected detection information for scenes in which the work cycle time is greater than or equal to a threshold, and determines that the work satisfies the condition if the work cycle time is greater than or equal to the threshold.

3. the scene to be extracted includes a scene in which a worker performs a specific action; The video aggregation device of claim 1 or 2, wherein the determination unit determines that the work satisfies the condition when the worker moves to a position corresponding to the location where the specific action is performed, based on the time series data of the detected detection information for a scene in which the worker performs the specific action.

4. 4. The video aggregating device according to claim 3, wherein the specific action is to place the defective product in a defective product storage area.

5. the scene to be extracted includes a scene in which an error log related to a device used in the work occurs, The video aggregation device according to any one of claims 1 to 4, wherein the determination unit further determines that the work satisfies the condition when a log related to equipment used in the work for a scene in which an error log occurred is an error log.

6. The acquisition unit acquires video images of the worker's work, a detection unit that detects time-series data of detection information related to a skeleton or a body part of the worker based on the moving image; a determination unit, based on the time-series data of the detected detection information, determining whether or not each of a plurality of types of scenes to be extracted satisfies a condition corresponding to the scene to be extracted; A generating unit generates a video that aggregates video images of portions including time points determined to be activities that satisfy conditions corresponding to the scenes to be extracted, based on the number of extractions or the extraction times set for each of the plurality of types of scenes to be extracted. Video aggregation method.

7. Acquire video footage of the worker's work, Detecting time-series data of detection information related to the skeleton or body part of the worker based on the moving image; determining whether or not a task satisfies a condition corresponding to each of a plurality of types of scenes to be extracted based on the time-series data of the detected detection information; A moving image is generated by aggregating moving images of portions including time points determined to be activities that satisfy conditions corresponding to the scenes to be extracted, based on the number of times of extraction or the extraction time set for each of the plurality of types of scenes to be extracted. A video aggregation program that allows a computer to perform the following:

Citation Information

Patent Citations

  • Work operation analysis system and work operation analysis method

    JP2019152802A

  • Work analysis system and work analysis method

    JP2020091801A