Video analysis system, video analysis method, and video analysis program

JPWO2024201632A5Pending Publication Date: 2025-12-04
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025509265
Authority / Receiving Office
JP · JP
Patent Type
Applications
Priority Date
2023-03-27
Filing Date
2023-03-27
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing video analysis systems in manufacturing environments face inefficiencies in identifying the timing of work performance due to redundant processes in recognizing work timing from video shots, as described in Patent Documents 1 and 2.

Method used

A video analysis system that includes an acquisition unit to capture video of a work area, a detection unit using a detection model to identify the work in each frame, and an output unit to specify the start and end timing of the work by analyzing the detection results in chronological order, reducing redundant frame detection and improving efficiency.

Benefits of technology

The system efficiently identifies the timing of work performance by narrowing down frames using a two-step detection process, thereby enhancing the accuracy and speed of work analysis and improving manufacturing efficiency.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present invention provides a video analysis system comprising an acquisition unit, a detection unit, an identification unit, and an output unit. The acquisition unit acquires a video taken of a work area. The detection unit uses a detection model for detecting work to be analyzed appearing in frames of video to detect, on the basis of the acquired video, whether the work to be analyzed appears at intervals of a set number of frames in a time series. On the basis of the result of using the detection model to detect whether the work to be analyzed appears, the identification unit identifies the start and the end of the work to be analyzed for each of predetermined frames based on the frames in which the work to be analyzed is detected. The output unit outputs the identified timings.
Need to check novelty before this filing date? Find Prior Art

Description

Video analysis system, video analysis method, and recording medium

[0001] The present disclosure relates to a video analysis system and the like.

[0002] In factories, for example, analysis of work within a manufacturing process is performed to improve work efficiency, quality, and safety. Work analysis is performed, for example, by analyzing video footage of the work. The results of the work analysis are used, for example, by a manager who manages the manufacturing process to improve work procedures. In addition, for example, the time required to perform a work specified in a work procedure may be used to analyze whether the work is being performed correctly. Analysis of the time required to perform a work is performed, for example, by identifying the start and end timings of the work based on the recognition results of video footage of the work.

[0003] The workload analysis device of Patent Document 1 determines whether a worker is working based on an image of the worker taken by a camera.

[0004] The information processing device of Patent Document 2 detects a range in which an event may occur from a time-series image, and detects the time when the event occurs from the event occurrence range.

[0005] JP 2021-125183 A JP 2021-56808 A

[0006] In the techniques described in Patent Documents 1 and 2, the process of identifying the timing at which work was performed in the video of the work area may be redundant.

[0007] In order to solve the above-mentioned problems, the present disclosure aims to provide a video analysis system etc. that can efficiently identify the timing of tasks.

[0008] In order to solve the above problems, the video analysis system of the present disclosure comprises an acquisition means for acquiring video footage of a work area, a detection means for detecting whether the work to be analyzed is shown in each of a set number of frames in chronological order based on the acquired video using a detection model that detects the work to be analyzed that is shown in the frames included in the video, an identification means for identifying at least one of the timings of the start and end of the work to be analyzed based on the results of detection using the detection model to determine whether the work to be analyzed is shown in each of a specified number of frames, using the frame in which the work to be analyzed is detected by the detection means as a reference, and an output means for outputting the identified timings.

[0009] The video analysis method disclosed herein acquires video of a work area, and based on the acquired video, uses a detection model that detects the work to be analyzed that is shown in frames included in the video to detect whether the work to be analyzed is shown in each of a set number of frames in chronological order, and, based on the results of the detection using the detection model to detect whether the work to be analyzed is shown in each of a specified number of frames using the frame in which the work to be analyzed is detected as a reference, identifies at least one of the timings of the start and end of the work to be analyzed, and outputs the identified timings.

[0010] The recording medium of the present disclosure non-temporarily records a video analysis program that causes a computer to execute the following processes: acquiring video footage of a work area; detecting whether the work to be analyzed is shown in a set number of frames in chronological order based on the acquired video using a detection model that detects the work to be analyzed that is shown in frames included in the video; determining the timing of at least one of the start and end of the work to be analyzed based on the results of detecting whether the work to be analyzed is shown in each of a specified number of frames using the frame in which the work to be analyzed was detected as a reference; and outputting the determined timing.

[0011] According to the present disclosure, the timing at which the work was performed can be efficiently identified.

[0012] FIG. 1 is a diagram illustrating an example of the configuration of a video analysis system according to an embodiment of the present disclosure; FIG. 2 is a diagram illustrating an example of a shooting form of a work area according to an embodiment of the present disclosure; FIG. 3 is a diagram illustrating an example of a detection result by a detection model according to an embodiment of the present disclosure; FIG. 4 is a diagram illustrating an example of a shooting form of a work area according to an embodiment of the present disclosure; FIG. 5 is a diagram illustrating an example of an operation flow of a video analysis system according to an embodiment of the present disclosure; and FIG. 6 is a diagram illustrating an example of the hardware configuration of a video analysis system according to an embodiment of the present disclosure.

[0013] An embodiment of the present disclosure will be described in detail with reference to the drawings. Fig. 1 is a diagram illustrating an example of the configuration of a video analysis system 10. The video analysis system 10 basically includes an acquisition unit 11, a detection unit 12, an identification unit 13, and an output unit 15. The video analysis system 10 also includes, for example, an extraction unit 14 and a storage unit 16.

[0014] The video analysis system 10 is a device that analyzes the work being performed by a worker in a work area based on, for example, a video of the work area. The analysis of the work being performed by the worker based on the video is performed by, for example, using image recognition to analyze whether the work is being performed in accordance with the work procedure.

[0015] When performing such an analysis, for example, by using image recognition to identify the timing of each task, it may be possible to analyze whether the task is being performed according to the correct procedure. Furthermore, for example, if the task is being performed correctly, the time required for a series of tasks specified in the work procedure will be within a certain range. Therefore, for example, by using image recognition to identify the timing of the start and end of a task, it may be possible to analyze whether the task is being performed correctly based on the time required for the task.

[0016] The video analysis system 10 identifies the timing at which the task to be analyzed was performed, for example, based on video footage of the work area. The video analysis system 10 identifies the timing at which the task started and finished, for example, based on video footage of the work area. Here, the timing at which the task started and finished is the timing at which frames corresponding to the task started and the task finished were captured. In other words, the video analysis system 10 identifies the timing at which frames corresponding to the task started and the task finished by identifying the timing at which frames corresponding to the task started and the task finished were captured.

[0017] The video analysis system 10 identifies frames captured at the start or end of a task from among multiple frames capturing the same action, for example, at the start or end of the task. The same action is, for example, an action that is identified as the same action through image recognition. The time required for a worker's action is usually longer than the interval between captures of each frame included in the video. Therefore, when identifying the start of a task, the same action performed at the start of the task is captured in multiple frames in the video captured of the task in the work area. For example, when a worker presses a button at the start of the task, the action that is identified as the worker pressing the button through image recognition is captured in multiple frames, such as the moment the button is touched and the moment the button is pressed.

[0018] When determining the start timing of a task, the video analysis system 10 determines, for example, the timing at which the first frame image in chronological order was captured among a plurality of frames showing the same action performed at the start of the task as the start timing of the task. Furthermore, when determining the end timing of a task, the video analysis system 10 determines, for example, the timing at which the last frame image in chronological order was captured among a plurality of frames showing the same action performed at the end of the task as the start timing of the task. The frames used to determine the start and end timings of a task are not limited to those described above.

[0019] Furthermore, a work area is, for example, an area in a factory where one or more workers perform predetermined tasks. The predetermined tasks are, for example, tasks specified in a work procedure. The work procedure specifies, for example, the name of the product to be worked on, the materials to be used, the tools to be used, and the order of the tasks. A combination of multiple tasks whose order is specified in a work procedure is also called a work flow. The items specified in a work procedure are not limited to those mentioned above. Work procedures are created, for example, by a production engineer or a manufacturing process manager.

[0020] FIG. 2 shows an example of how work performed in a work area is captured. In the example of FIG. 2, a worker is working in a position where he or she can work on a product placed on a workbench. Also, in the example of FIG. 2, the worker is, for example, using tools held in his or her hand to assemble the product placed on the workbench. In the example of FIG. 2, the worker performs work, for example, according to a work procedure. Also, in the example of FIG. 2, a camera is installed on the workbench and to capture images of the worker. In this case, the work area is, for example, the workbench and the location where the worker is standing. The location where the worker is standing includes the range in which the worker can move while working. The camera outputs the captured images to the video analysis system 10, for example, via a network. The video analysis system 10, for example, identifies the start and end timings of work based on the video of the work area captured by the camera.

[0021] The video analysis system 10 detects whether the task to be analyzed is captured in each frame at a first interval in a time series of video of the work area. Furthermore, the video analysis system 10 detects whether the task to be analyzed is captured in each frame at a second interval shorter than the first interval, for frames within a predetermined time range from the frame in which the task to be analyzed is detected. The video analysis system 10 then determines whether the first or last frame in the time series of multiple frames in which the task to be analyzed is detected is the start or end of the task. In other words, the video analysis system 10 uses a detection model to detect frames capturing the task to be analyzed in two stages from the video of the work area, thereby determining the start and end of the task.

[0022] FIG. 3 is a diagram schematically illustrating the relationship between a target frame for which the detection model detects the work to be analyzed and a frame in which the work to be analyzed is captured. In the example of FIG. 3, the horizontal axis is the time axis. In the example of FIG. 3, the line shown above the horizontal axis as a "detection peak" indicates the time point at which the frame in which the work to be analyzed is captured is captured. In addition, in the example of FIG. 3, the upward arrow indicates the time point at which the target frame for which the detection model detects the work to be analyzed is captured is captured.

[0023] In the example of FIG. 3 , in step 1 shown on the left, detection is performed every five frames to determine whether the work to be analyzed is shown. Step 1 is a step for detecting whether the work to be analyzed is shown for each frame of a first interval. In step 1 of the example of FIG. 3 , in the frame indicated as "A," the frame in which the work to be analyzed is shown matches the frame to be detected by the detection model. Therefore, the detection model detects that the work to be analyzed is being performed in the frame indicated as "A."

[0024] In the example of FIG. 3 , in step 2 shown on the right, whether the task to be analyzed is captured is detected for each frame. Step 2 is a step for detecting whether the task to be analyzed is captured for each frame at a second interval, which is shorter than the first interval. In the example of FIG. 3 , the frame indicated by the bold arrow in step 2 corresponds to the frame to be analyzed in step 1. In step 2 of the example of FIG. 3 , the detection model detects whether the task to be analyzed is captured for four frames before and after the frame indicated as "A" in chronological order. In step 2, the detection model detects whether the task to be analyzed is being captured in five consecutive frames. In this case, when determining the start timing of the task, the video analysis system 10 determines that the first frame of the five consecutive frames is the start timing of the task. In this way, by narrowing down the frames that are likely to capture the task to be analyzed in step 1 and performing more detailed detection in step 2, the number of frames to be detected by the detection model can be reduced. This makes it possible to efficiently determine, for example, the start and end timing of the task.

[0025] A specific example of the configuration of the video analysis system 10 will now be described.

[0026] The acquisition unit 11 acquires video footage of the work area. Furthermore, the video footage of work is, for example, video footage captured so as to capture work performed by a worker in the work area. The work performed by the worker is, for example, work related to product manufacturing. Work related to product manufacturing includes, for example, product assembly, product processing, product inspection, product packaging, product transportation, equipment operation, or storage shelf organization. The work performed by the worker in product manufacturing may also include equipment inspection, equipment repair, or equipment assembly. Work related to product manufacturing may also include food preparation. Work related to product manufacturing is not limited to the above. Furthermore, the work performed by the worker may include cleaning, driving, exercise, instructing others, guiding, security, medical procedures, dispensing medicine, and equipment operation. The work performed by the worker is not limited to the above.

[0027] The acquisition unit 11 acquires, for example, video of the work area captured by a camera installed in a position where the work area can be captured. If a monitoring system for monitoring the manufacturing process is present, the acquisition unit 11 may acquire the video of the work area via a server of the monitoring system. Alternatively, the acquisition unit 11 may acquire the video of the work area via a recording medium.

[0028] The detection unit 12 uses a detection model that detects the work of the analysis target that is shown in frames included in the video, based on the video acquired by the acquisition unit 11, to detect whether the work of the analysis target is shown for each set number of frames in time series. For example, the detection unit 12 uses the detection model based on the video acquired by the acquisition unit 11 to detect whether the work of the analysis target is shown for each frame at a first interval in time series. The detection model is a learning model that detects the work of the analysis target that is shown in frames included in the video by image recognition. For example, the detection model detects whether the action of the detection target is shown in each frame.

[0029] The detection unit 12 uses the detection model to detect whether the task to be analyzed is captured in frame images at intervals of 10 frames, which are set as a first interval. For example, if the first interval is set to 10 frames, the detection unit 12 detects whether the task to be analyzed is captured in frame images at intervals of 10 frames, such as the first, eleventh, and twenty-first frames in a time series. The first interval is set, for example, so that the task to be analyzed can be detected in any of the frames to be detected when the detection model thins out the frames to be detected. The first interval is set, for example, based on the time required for the task to be analyzed. The first interval is set, for example, to an interval at which the task to be analyzed can be detected. The first interval is set, for example, to a time interval shorter than the time required for the task to be analyzed. For example, if the task to be analyzed takes one minute, the first interval is set to an interval shorter than the number of frames contained in one minute of video. The first interval is set, for example, by the operator of the video analysis system 10 or the person analyzing the task.

[0030] The detection model is, for example, a machine learning model using a neural network. The detection model is generated, for example, by learning the relationship between an image showing the target motion, the area showing the target motion, and the type of target motion. The detection model is generated, for example, in a system external to the video analysis system 10.

[0031] The detection unit 12 may detect a starting task and a finishing task using multiple detection models. For example, the detection unit 12 may use a detection model that detects the position of the worker's body from video to identify the position of the worker's hand in the frame image. Furthermore, the detection unit 12 may detect the tool held by the worker at the identified position using a detection model that detects the type of tool. The detection unit 12 may then detect the start of a task, for example, when the worker's hand is in a tool storage area and the tool being held is the tool designated for use in the first task of the workflow. Furthermore, the detection unit 12 may detect the end of a task, for example, when, after the task has started, the worker's hand is located in a tool storage area and the tool being held is the tool designated for use in the last task of the workflow. Furthermore, the detection unit 12 may detect the starting task and the finishing task using different detection models. Furthermore, the detection unit 12 may detect the task motion and the task motion using different detection models for each task included in the workflow.

[0032] The detection unit 12 may use the detection model to detect at least one of the worker's clothing and / or accessories. The detection unit 12 may use the detection model to detect the position where the worker is standing. The detection unit 12 may also use the detection model to detect the position where the worker is sitting. The detection unit 12 may also use the detection model to detect the worker's posture. The detection unit 12 may also use the detection model to detect the number of workers in the work area. The detection unit 12 may also use the detection model to detect whether a lamp indicating that work is in progress is on or off. The detection targets using the detection model are not limited to those described above.

[0033] The detection unit 12 detects whether the task to be analyzed is shown in each frame of a first interval, for example, with frames included in the video of the section extracted by the extraction unit 14 as the detection target. That is, the detection unit 12 detects whether the task to be analyzed is shown in each frame of a first interval from the frame of the video of the section narrowed down by the extraction unit 14. Details of the extraction unit 14 will be described later.

[0034] The identification unit 13 identifies at least one of the start and end timings of the work to be analyzed for each of predetermined frames, based on the results of detecting whether the work to be analyzed is captured using the detection model, using the frame in which the work to be analyzed is detected by the detection unit 12 as a reference. The identification unit 13 identifies at least one of the start and end timings of the work to be analyzed for each of predetermined frames, based on the results of detecting frames within a predetermined range using the detection model for each frame at a second interval. The frames included in the predetermined range are, for example, frames included within a predetermined range from a reference point on the time series, where the reference point is the position on the time series of the frame in which the work to be analyzed is detected by the detection unit 12. The predetermined range is set, for example, based on a first interval. The predetermined range is set, for example, to include up to the frame immediately preceding the frame separated by the first interval from the reference point on the time series. For example, if the first interval is 10 frames, the predetermined range is set to within 9 frames from the reference point on the time series. The predetermined range may also be set to include up to a frame intermediate between the reference point and the frame separated by the first interval on the time series. For example, if the first interval is 10 frames, the predetermined range is set to within 5 frames from the reference point on the time series. The predetermined range can be set as appropriate.

[0035] The second interval is a time interval shorter than the first interval. The second interval is set, for example, based on the frame rate of the video of the work area and the accuracy required for timing identification. For example, when it is necessary to accurately identify the start or end timing of a task, the second interval is set so that it is possible to detect whether the task to be analyzed is captured in all frames within a predetermined range. For example, when a one-second delay in the detection of the start and end timing of a task does not pose a problem for analysis, the second interval is set based on the number of frames captured per second. For example, when a one-second delay in the detection of the start and end timing of a task does not pose a problem for analysis, the second interval is set to the value of the number of frames captured per second. The second interval and the predetermined range are set, for example, by the operator of the video analysis system 10 or the person analyzing the tasks.

[0036] The identification unit 13 detects whether the task to be analyzed is shown in each frame of the second interval, for example, using the same detection model as the detection unit 12. That is, the identification unit 13 detects whether the task to be analyzed is shown in each frame of the second interval, which is shorter than the first interval at which the detection unit 12 detects whether the task to be analyzed is shown, for example, using the same detection model as the detection unit 12. Furthermore, the identification unit 13 may detect whether the task to be analyzed is shown in each frame of the second interval, using a detection model that differs from the detection model used by the detection unit 12 in at least one of learning data and algorithm.

[0037] Furthermore, the identification unit 13 may set at least one of the second interval and the predetermined range based on the detection result by the detection unit 12. For example, the identification unit 13 may set the second interval so that the higher the detection frequency of the frames detected by the detection result of the detection unit 12, the shorter the second interval. Furthermore, the identification unit 13 may set the second interval so that the longer the time required for the task detected by the detection model. For example, the identification unit 13 sets the predetermined range in accordance with the set second interval. For example, the identification unit 13 sets the predetermined range so that frames included within the predetermined range do not overlap with frames of adjacent second intervals in the time series. The criteria for setting the second interval and the predetermined range are not limited to those described above.

[0038] The identification unit 13, for example, identifies frames captured at the start of the work and at the start of the work from consecutive frames in which the detection model detects the work to be analyzed, and then identifies the time points in the time series at which the identified frames were captured as the start of the work and at the start of the work, respectively.

[0039] When the detection model detects consecutive tasks performed by a worker from the start to the end of the task as a single task, the identification unit 13 identifies, for example, the timing at which the first frame of the consecutive frames in which the task to be analyzed is detected is captured as the start timing of the task. Here, consecutive frames refer to frames that fall within the second interval and are the target of detection when the identification unit 13 detects the task using the detection model. In other words, consecutive frames refer to frames that are consecutive when only the frames that fall within the second interval are considered. Furthermore, the identification unit 13 identifies, for example, the timing at which the last frame of the consecutive frames in which the task to be analyzed is detected is captured as the end timing of the task.

[0040] Furthermore, for example, when the detection model detects each of multiple tasks included in a workflow as an independent task, the identification unit 13 identifies, for example, the timing at which the first frame of a series of consecutive frames in which the first task in the workflow is detected is captured as the start timing of the task. Furthermore, the identification unit 13 identifies, for example, the timing at which the last frame of a series of consecutive frames in which the last task in the workflow is detected is captured as the end timing of the task. Furthermore, the identification unit 13 may identify the start and end timings of each of multiple tasks included in the workflow. The identification unit 13 may also identify the start and end timings of tasks other than the start and end of the workflow, among the multiple tasks included in the workflow.

[0041] When the task to be analyzed is a task that is performed repeatedly, the identification unit 13 identifies, for example, the start and end of the task for each repetition. When the task to be analyzed is a task that is performed repeatedly, the identification unit 13 may identify the start of the task in the first repetition and the end of the task in the last repetition.

[0042] Furthermore, when the start and end of multiple tasks included in a workflow are detected as the same task by the detection model, the identification unit 13 determines, for example, that the first consecutive frames detected in time series among the consecutive frames detected by the detection model correspond to the start of the task. The identification unit 13 then determines, using the detection model, that the next consecutive frames detected in time series correspond to the end of the task. The identification unit 13 also determines that the next consecutive frames detected in time series after the consecutive frames determined by the detection model to correspond to the end of the task correspond to the start of the task. For example, when the start and end tasks of tasks included in a workflow are button pressing actions, the identification unit 13 determines, after starting analysis of the video, that the first consecutive frames detected in time series as the button pressing action correspond to the start of the task. The identification unit 13 then determines that the next consecutive frames detected in time series as the button pressing action correspond to the end of the task.

[0043] When the analysis target work is detected in all frames of the second interval in a section sandwiched between frames of the first interval that are adjacent on the time series in which the analysis target work is detected, the identification unit 13 determines, for example, that the analysis target work continues to be detected between adjacent frames of the first interval. In such a case, the identification unit 13 determines, for example, that a frame of the second interval that corresponds to a frame of the first interval that is earlier on the time series and a frame of the second interval that corresponds to a frame of the first interval that is later on the time series are consecutive frames.

[0044] If the task to be analyzed is detected in consecutive frames in a time series, and then the task to be analyzed is not detected in one frame, and then the task to be analyzed is again detected in consecutive frames, the identification unit 13 may determine that the movement is continuing before and after the frame in which the task to be analyzed is not detected. That is, when detecting the task to be analyzed, if the task to be analyzed is not detected in a predetermined number of frames or less in which the task to be analyzed is detected before and after the frame, the identification unit 13 may determine that the task to be analyzed is continuing. When there is a frame in which the task to be analyzed is not detected, the number of frames for determining whether the movement is continuing in the previous and next frames in the time series is not limited to the above example of one frame and can be set appropriately. Furthermore, when there is a frame in which the task to be analyzed is not detected, the identification unit 13 may use the detection model to detect whether the task to be analyzed is shown in the previous and next frames in the time series at intervals shorter than the second interval.

[0045] The extraction unit 14 extracts sections that satisfy set conditions from, for example, video footage of the work area. The extraction unit 14 extracts sections where work is likely to be taking place from, for example, video footage of the work area. The set condition, for example, is that the section is a time period during which work is being performed in the work area. For example, break times are not included in the sections from which the extraction unit 14 extracts video. When detecting work during a time period when work is not scheduled to be performed in the work area, such as during a break time, the set condition may be a time period during which work is not being performed in the work area. The extraction unit 14 acquires information about time periods during which work is being performed in the work area or break times from, for example, a production management server.

[0046] The set condition may be that a worker is shown in the frame. Alternatively, the set condition may be that a person is shown in the frame in a posture appropriate for the task being analyzed. For example, the extraction unit 14 uses a detection model to detect whether a person is shown in a frame included in a video of the work area in a posture appropriate for the task being analyzed. Then, the extraction unit 14 extracts, from the video of the work area, a section of video in which a person is shown in a frame in a posture appropriate for the task being analyzed. For example, if the task being analyzed is performed while sitting in a chair, the extraction unit 14 extracts a section of video in which a person sitting in a chair is shown. Alternatively, the posture appropriate for the task may include the person holding a tool used for the task being analyzed. For example, if the task being analyzed is cleaning, the extraction unit 14 may extract a section of video in which a person holding cleaning tools is shown.

[0047] FIG. 4 is a diagram showing an example of a shooting format of a work area. In the example of FIG. 4, a worker is standing in the work area holding cleaning tools. For example, if the work to be analyzed is cleaning, in the example of FIG. 4, the extraction unit 14 detects the worker holding cleaning tools, for example, using a detection model. For example, the extraction unit 14 extracts, from the video of the work area, a section of consecutive frames in which the worker holding cleaning tools is captured, as a section that satisfies a set condition.

[0048] The set condition may be set based on the number of people working in the work area. For example, the set condition may be set as the number of people working in the work area being different from the usual number of people. For example, if the work procedure stipulates that work should be performed in pairs for safety reasons, the extraction unit 14 extracts video of sections in the work area where a single person is working from the video of the work area. The extraction unit 14 may also extract video of sections from the video of the work area that exclude sections in which multiple people are working on tasks that are normally performed by one person. For example, if two or more people are performing a single-person task, this is because a supervisor is present and the task does not need to be analyzed. The set condition may also be set based on the attributes of the people in the work area. For example, the extraction unit 14 extracts sections in the video of the work area where people with a rank below a standard based on the proficiency level of the task are working. In this case, the detection model detects the rank of the people in the work area based on, for example, the color or marking of the work uniforms that are assigned to each rank.

[0049] The set condition may be that a worker is standing at a predetermined position or that a worker is sitting at a predetermined position. The set condition may also be that a worker is wearing predetermined clothing. The set condition may also be that a worker is wearing predetermined accessories. For example, the extraction unit 14 extracts, from a video of the work area, a section in which a worker wearing predetermined accessories is captured, as a section that satisfies the set condition. The set condition may also be that a lamp indicating that work is in progress is on or off. The set condition may also be that a display indicating that work is in progress is being displayed. The set conditions are not limited to the above.

[0050] In this way, by the extraction unit 14 extracting sections that satisfy the set conditions, the number of frames that the detection unit 12 and the identification unit 13 target for detection can be reduced. The extraction unit 14 may also extract sections that satisfy the set conditions from a video of a work area with a reduced number of pixels. By extracting sections that satisfy the set conditions in a state where the number of pixels of images included in the video of the work area has been reduced, the amount of processing required to extract sections that satisfy the set conditions can be reduced. In this way, by the extraction unit 14 extracting sections that satisfy the set conditions, the time and computer resources required for the process of extracting sections that satisfy the set conditions from the video of the work area can be reduced.

[0051] When extracting sections that satisfy the set conditions using a detection model, the extraction unit 14 extracts sections that satisfy the set conditions from the video of the work area using, for example, a detection model that differs from the detection model used by the detection unit 12 in at least one of learning data and algorithm. When detecting sections in which people appear in the video, the detection model may not require as high image recognition accuracy as a detection model that detects the work being analyzed. In such cases, by using a detection model that cannot detect the details of the work but can detect whether a person is appearing, the time and computer resources required for the process of extracting sections that satisfy the set conditions from the video of the work area can be further reduced.

[0052] The extraction unit 14 may, for example, use the same detection model as the detection unit 12 to extract frames from footage of the work area in which a person is shown in a posture related to the work being analyzed.

[0053] The output unit 15 outputs the timing identified by the identification unit 13. The output unit 15 outputs at least one of the timings of the start and end of the work identified by the identification unit 13, for example, to a terminal device operated by a person analyzing the work. The output unit 15 may output a video recording the timings of the start and end of the work. The output unit 15 may output the timing identified by the identification unit 13 to a display device (not shown) connected to the video analysis system 10. The output unit 15 may output the timing identified by the identification unit 13 to a storage means connected to the video analysis system 10 via a network. Furthermore, the output unit 15 may output the timing identified by the identification unit 13 to a manufacturing process management system.

[0054] The output unit 15 may output, from among the images captured of the work area, images between the start and end of work. Furthermore, the output unit 15 may output a work time calculated from the start and end of work. Furthermore, the output unit 15 may output images of frames identified as the start and end of work, and information indicating that the images of the frames identified as the start and end of work, together with images of frames preceding and following them in chronological order.

[0055] The storage unit 16 stores, for example, video footage of the work area. The storage unit 16 also stores a detection model. The detection model may be stored in a storage means external to the video analysis system 10. The storage unit 16 stores, for example, the results of identifying the timing of the start and end of work. The storage unit 16 may store, from the video footage of the work area, video footage between the start and end of work. The storage unit 16 may store images of frames identified as the timings of the start and end of work, as well as images of frames preceding and following them in chronological order.

[0056] The following describes the operation of the video analysis system 10 to identify the start of a task and the timing of the task start. Fig. 5 shows an example of the operation flow when the video analysis system 10 identifies the start of a task and the timing of the task start.

[0057] The acquisition unit 11 acquires a video of the work area (step S11). For example, the acquisition unit 11 acquires a video of the work area captured by a camera installed so as to be able to capture the work area.

[0058] When a video of the work area is acquired, the detection unit 12 uses a detection model based on the acquired video to detect whether the analysis target work is shown in each frame at a first interval in time series (step S12). The detection model detects the analysis target work shown in the frames included in the video.

[0059] When the task to be analyzed is detected, the identification unit 13 identifies at least one of the start and end timings of the task to be analyzed based on the results of detection for each frame at a second interval using the detection model for frames included within a predetermined range (step S13). The frames included within the predetermined range are frames included within a predetermined range on the time series, with the position on the time series of the frame at which the task to be analyzed was detected by the detection unit 12 as the reference point. The second interval is an interval on the time series that is shorter than the first interval.

[0060] In step S13, if the timing of the start and end of the work being analyzed for all of the repetitive work has been identified (Yes in step S14), the output unit 15 outputs the timing identified by the identification unit 13 (step S15).

[0061] In step S13, if the timing of the start and end of all repetitive tasks has not been determined (No in step S15), the process returns to step S12, and the detection unit 12 uses a detection model based on the video acquired by the acquisition unit 11 to detect whether the task to be analyzed is shown in each frame of the first interval in chronological order.

[0062] The video analysis system 10 uses a detection model to detect whether a task to be analyzed is captured in frames captured at a first interval in a video of the work area. Furthermore, the video analysis system 10 uses a detection model to detect whether a task to be analyzed is captured in frames captured at a second interval shorter than the first interval for frames within a predetermined range of frames in which the task to be analyzed is detected. Based on the results of the detection, the video analysis system 10 identifies the start and end timings of the task to be analyzed. That is, the video analysis system 10 narrows down frames that are likely to represent the start and end timings of the task using the first interval, and then performs more detailed detection using the second interval. In this way, by identifying the start and end timings of the task to be analyzed in two stages, the video analysis system 10 can reduce the number of frames included in the video of the work area that are subject to detection by the detection model for the task to be analyzed. This allows the video analysis system 10 to efficiently identify the start and end timings of the task from the video of the work area.

[0063] Furthermore, after extracting a section that satisfies the set conditions, the video analysis system 10 can more efficiently identify the timing of the start and end of the work by detecting whether the work to be analyzed is shown in each frame of the first interval and then each frame of the second interval in that section.

[0064] The processes in the video analysis system 10 may be distributed and executed among multiple information processing devices connected via a network. For example, the processes in the detection unit 12 and the identification unit 13 and the process in the extraction unit 14 may be executed in different information processing devices. It can be set as appropriate which of the multiple information processing devices executes each process in the video analysis system 10.

[0065] Each process in the video analysis system 10 can be realized by executing a computer program on a computer. Fig. 6 shows an example of the configuration of a computer 100 that executes a computer program that performs each process in the video analysis system 10. The computer 100 includes a CPU (Central Processing Unit) 101, a memory 102, a storage device 103, an input / output I / F (Interface) 104, and a communication I / F 105.

[0066] The CPU 101 reads and executes computer programs for each process from the storage device 103. The CPU 101 may be configured as a combination of multiple CPUs. The CPU 101 may also be configured as a combination of a CPU and another type of processor. For example, the CPU 101 may be configured as a combination of a CPU and a graphics processing unit (GPU). The memory 102 is configured with a dynamic random access memory (DRAM) or the like, and temporarily stores computer programs executed by the CPU 101 and data being processed. The storage device 103 stores computer programs executed by the CPU 101. The storage device 103 is configured with, for example, a non-volatile semiconductor storage device. Other storage devices such as a hard disk drive may also be used for the storage device 103. The input / output I / F 104 is an interface that accepts input from an operator and outputs display data, etc. The communication I / F 105 is an interface that transmits and receives data to and from other information processing devices. The computer programs used to execute each process can also be distributed by storing them on a computer-readable storage medium that non-temporarily stores data. The recording medium may be, for example, a magnetic disk such as a magnetic tape for recording data or a hard disk. Alternatively, the recording medium may be an optical disk such as a CD-ROM (Compact Disc Read Only Memory). A non-volatile semiconductor memory device may also be used as the recording medium.

[0067] The present disclosure has been described above using the above-described embodiments as examples. However, the present disclosure is not limited to the above-described embodiments. That is, the present disclosure can be applied in various aspects that can be understood by a person skilled in the art within the scope of the present disclosure.

[0068] REFERENCE SIGNS LIST 10 Video analysis system 11 Acquisition unit 12 Detection unit 13 Identification unit 14 Extraction unit 15 Output unit 16 Storage unit 100 Computer 101 CPU 102 Memory 103 Storage device 104 Input / output I / F 105 Communication I / F

Claims

1. An acquisition means for acquiring an image of a work area; a detection means for detecting whether the analysis target task is shown in each of a set number of frames in chronological order based on the acquired video, using a detection model that detects the analysis target task shown in frames included in the video; and an identification means for identifying at least one of the start and end timings of the work to be analyzed based on the results of detecting whether the work to be analyzed is shown in each of predetermined frames using the detection model, with the frame in which the work to be analyzed is detected by the detection means as a reference; an output means for outputting the specified timing; A video analysis system comprising:

2. the detection means uses the detection model to detect whether the task to be analyzed is captured in each frame at a first interval along a time series, based on the acquired video; the specifying means uses the detection model to detect whether the work to be analyzed is captured in each frame at a second interval shorter than the first interval, for frames included within a predetermined range on the time series, with the position on the time series of the frame in which the work to be analyzed is detected by the detection means as a reference point, and specifies at least one of the timings of the start and end of the work to be analyzed based on the results of the detection. The video analysis system of claim 1 .

3. further comprising an extraction means for extracting a section that satisfies a set condition from the video of the work area; the detection means detects whether the work to be analyzed is shown in the frames of the extracted section. The video analysis system according to claim 1 or 2.

4. The set condition is that a person appears in the frame. The video analysis system according to claim 3 .

5. The set condition is that the person is shown in the frame in a posture corresponding to the task to be analyzed. The video analysis system according to claim 4 .

6. At least one of the second interval and the predetermined range is an interval set based on a detection result by the detection means. The video analysis system according to claim 2 .

7. the first interval is an interval set based on the time required for the task to be analyzed; The video analysis system according to claim 2 .

8. The set condition is an interval set based on the number of people working in the work area. The video analysis system according to claim 3 .

9. Obtain footage of the work area, Based on the acquired video, a detection model is used to detect the analysis target work shown in frames included in the video, and it is detected whether the analysis target work is shown for each of a set number of frames in chronological order; using the detection model to detect whether the work to be analyzed is shown in each of predetermined frames, with the frame in which the work to be analyzed is detected as a reference, and based on the result, identifying at least one of the start and end timings of the work to be analyzed; Output the specified timing, Video analysis methods.

10. Acquiring video of the work area; a process of detecting whether the analysis target task is shown in each of a set number of frames in chronological order based on the acquired video using a detection model that detects the analysis target task shown in frames included in the video; a process of identifying at least one of the start and end timings of the work to be analyzed based on the results of detecting whether the work to be analyzed is shown using the detection model for each of predetermined frames, with the frame in which the work to be analyzed is detected as a reference; Processing to output the specified timing A video analysis program that runs on a computer.