Video analysis system, video analysis method, and video analysis program

JPWO2024201631A5Pending Publication Date: 2025-10-03
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025509264
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-07-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing video analysis systems face difficulties in accurately identifying the timing of work performance in video recordings of manufacturing processes, making it challenging to analyze work efficiency, quality, and safety effectively.

Method used

A video analysis system that acquires videos of work areas, uses a detection model to identify motion within frames, and specifies the timing of work start and end based on a scoring system indicating the relationship between reference points and frame positions, allowing for precise output of these timings.

Benefits of technology

Enables easy and accurate identification of work timing, improving the analysis of work efficiency, quality, and safety by determining when tasks are performed correctly and within specified time frames.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This video analysis system is configured to comprise an acquisition unit, a detection unit, an identification unit, and an output unit. The acquisition unit acquires a captured video of a work area. The detection unit uses a detection model for detecting movement of an object of detection appearing in frames included in the video to detect movement of the object of detection from each of the frames in a time series included in the captured video of the work area. The identification unit identifies the timing of at least one of the start of work or the end of work on the basis of a score indicating the relationship between a reference point in the time series and the position, in the time series, of the frame in which movement of the object of detection is detected. The output unit outputs the identified timing.
Need to check novelty before this filing date? Find Prior Art

Description

Video analysis system, video analysis method, and recording medium

[0001] The present disclosure relates to a video analysis system and the like.

[0002] In factories, for example, analysis of work in manufacturing processes is performed to improve work efficiency, quality, and safety. Work analysis is performed, for example, by analyzing video footage of the work. The results of the work analysis are used, for example, by a manager who manages the manufacturing process to improve work procedures. In addition, for example, the time required to perform a work defined as a work procedure may be used to analyze whether the work is being performed correctly. Analysis of the time required to perform a work is performed, for example, by identifying the start and end timings of the work based on the recognition results of the video footage of the work.

[0003] The information processing device of Patent Document 1 identifies work procedures based on video footage of the work, and then determines whether a task has been omitted based on the identified work procedures.

[0004] Japanese Patent Application Laid-Open No. 2019-149154

[0005] With the technology described in Patent Document 1, it may be difficult to identify the timing at which work was performed in the video of the work area.

[0006] In order to solve the above-mentioned problems, the present disclosure aims to provide a video analysis system etc. that can easily identify the timing at which an operation was performed.

[0007] In order to solve the above problems, the video analysis system of the present disclosure comprises an acquisition means for acquiring video footage of the work area, a detection means for detecting the behavior of the target object from each of the frames in a time series contained in the video footage of the work area using a detection model that detects the behavior of the target object shown in the frames contained in the video, an identification means for identifying the timing of at least one of the start and end of the work based on a score indicating the relationship between a reference point on the time series and the position on the time series of the frame in which the behavior of the target object was detected, and an output means for outputting the identified timing.

[0008] The video analysis method disclosed herein acquires video of a work area, and uses a detection model that detects the movement of a target object shown in frames included in the video to detect the movement of a target object from each of the frames in a time series included in the video of the work area.The method then identifies at least one of the start and end times of the work based on a score indicating the relationship between a reference point on the time series and the position on the time series of the frame in which the movement of the target object was detected, and outputs the identified timing.

[0009] The recording medium of the present disclosure non-temporarily records a video analysis program that causes a computer to execute the following processes: acquiring video of the work area; detecting the motion of the target object from each of the frames in a time series contained in the video of the work area using a detection model that detects the motion of the target object shown in the frames contained in the video; identifying the timing of at least one of the start and end of the work based on a score indicating the relationship between a reference point in the time series and the position in the time series of the frame in which the motion of the target object was detected; and outputting the identified timing.

[0010] According to the present disclosure, the timing at which the work was performed can be easily identified.

[0011] FIG. 1 is a diagram illustrating an example of the configuration of a video analysis system according to an embodiment of the present disclosure; FIG. 2 is a diagram illustrating an example of a shooting form of a work area according to an embodiment of the present disclosure; FIG. 3 is a diagram illustrating an example of a worker's work according to an embodiment of the present disclosure; FIG. 4 is a diagram illustrating an example of a detection result by a detection model according to an embodiment of the present disclosure; FIG. 5 is a diagram illustrating an example of a position of a frame in which the motion of a detection target is detected according to an embodiment of the present disclosure; FIG. 6 is a diagram illustrating an example of a detection result according to an embodiment of the present disclosure; FIG. 7 is a diagram illustrating an example of an operation flow of a video analysis system according to an embodiment of the present disclosure; and FIG. 8 is a diagram illustrating an example of the hardware configuration of a video analysis system according to an embodiment of the present disclosure.

[0012] An embodiment of the present disclosure will be described in detail with reference to the drawings. Fig. 1 is a diagram illustrating an example of the configuration of a video analysis system 10. The video analysis system 10 basically includes an acquisition unit 11, a detection unit 12, an identification unit 13, and an output unit 14. The video analysis system 10 also includes, for example, a storage unit 15.

[0013] The video analysis system 10 is, for example, a device that analyzes work performed by a worker in a work area based on video captured of the work area. The analysis of work performed by a worker based on video is performed, for example, by using image recognition to analyze whether a series of tasks defined in a work procedure are being performed correctly. When performing such an analysis, for example, by using image recognition to identify the timing of each task, it may be possible to analyze whether the work is being performed according to the correct procedure. Furthermore, for example, when work is being performed correctly, the time required for the series of tasks defined in the work procedure falls within a certain range. Therefore, for example, by using image recognition to identify the timing of the start and end of a task, it may be possible to analyze whether the task is being performed correctly based on the time required for the task.

[0014] The video analysis system 10 identifies the timing at which the task to be analyzed was performed, for example, based on video footage of the work area. The video analysis system 10 identifies the timing at which the task started and finished, for example, based on video footage of the work area. Here, the timing at which the task started and finished is the timing at which frames corresponding to the task started and the task finished were captured. In other words, the video analysis system 10 identifies the timing at which frames corresponding to the task started and the task finished by identifying the timing at which frames corresponding to the task started and the task finished were captured.

[0015] For example, at the start or end of a task, the video analysis system 10 identifies which frame of images captured from multiple frames showing the same action represents the start or end of the task. The same action is, for example, an action that is identified as the same action through image recognition. For example, when identifying the start of a task, the same action performed at the start of the task is captured in multiple frames in a video of the task being performed in a work area. For example, when a worker presses a button at the start of the task, the action that is identified as the worker pressing the button through image recognition is captured in multiple frames, such as the moment the button is touched and the moment the button is pressed. The video analysis system 10 identifies which frame of images captured from multiple frames showing the same action represents the start of the task.

[0016] Furthermore, a work area is, for example, an area in a factory where one or more workers perform predetermined tasks. The predetermined tasks are, for example, tasks specified in a work procedure. The work procedure specifies, for example, the name of the product to be worked on, the materials to be used, the tools to be used, and the order of the tasks. A combination of multiple tasks whose order is specified in a work procedure is also called a work flow. The items specified in a work procedure are not limited to those mentioned above. Work procedures are created, for example, by a production engineer or a manufacturing process manager.

[0017] FIG. 2 shows an example of how work performed in a work area is captured. In the example of FIG. 2, a worker is standing in a position where he or she can work on a product placed on a workbench. Also, in the example of FIG. 2, the worker is, for example, using tools held in his or her hand to assemble the product placed on the workbench. In the example of FIG. 2, the worker performs work, for example, according to a work procedure. Also, in the example of FIG. 2, a camera is installed on the workbench and to capture images of the worker. In this case, the work area is, for example, the workbench and the location where the worker is standing. The location where the worker is standing includes the range in which the worker can move while working. The camera outputs the captured images to the video analysis system 10, for example, via a network. The video analysis system 10, for example, identifies the start and end timings of work based on the video of the work area captured by the camera.

[0018] A specific example of the configuration of the video analysis system 10 will now be described.

[0019] The acquisition unit 11 acquires video of the work area. The video of the work area is, for example, video captured so as to capture the actions of the worker from the start to the end of the work performed by the worker. The work performed by the worker is, for example, work related to the manufacture of a product. Examples of work related to the manufacture of a product include product assembly, product processing, product inspection, product packaging, product transportation, equipment operation, and storage shelf organization. The work performed by the worker related to the manufacture of a product may also include equipment inspection, equipment repair, or equipment assembly. Work related to the manufacture of a product may also include food preparation. Work related to the manufacture of a product is not limited to the above. Work performed by the worker may also include cleaning, driving, exercise, instructing others, guiding, security, medical procedures, dispensing medicine, and equipment operation. The work performed by the worker is not limited to the above.

[0020] The acquisition unit 11 acquires, for example, video of the work area captured by a camera installed in a position where the work area can be captured. If a monitoring system for monitoring the manufacturing process is present, the acquisition unit 11 may acquire the video of the work area via a server of the monitoring system. Alternatively, the acquisition unit 11 may acquire the video of the work area via a recording medium.

[0021] The detection unit 12 detects the behavior of the detection target from each of the time-series frames included in the video using a detection model that detects the behavior of the detection target that is captured in the frames included in the video. The detection model is a learning model that uses image recognition technology to identify the work that is captured in each of the time-series frames included in the video. The detection model identifies, for example, whether the behavior of the detection target is captured in each frame.

[0022] The motions to be detected are, for example, motions related to tasks performed at the start and end of multiple tasks included in a workflow. For example, if the first task in a workflow is defined as picking up a tool, the detection model detects the motion of a worker picking up a tool, which is performed as the starting task. Furthermore, if the last task in a workflow is defined as putting down a tool, the detection model detects the motion of a worker putting down a tool, which is performed as the ending task. The starting and ending tasks may be, for example, tasks related to reporting the start and end of a task. For example, if a worker reports the start and end of a task by pressing a reporting button, the detection model detects the motion of the worker pressing the reporting button. Furthermore, the motions to be detected may be the starting and ending motions of individual tasks included in a workflow. Furthermore, the motions to be detected may be the starting and ending motions of some of the tasks included in a workflow. For example, the motions to be detected may be the starting motion of the first task and the ending motion of the last task of three consecutive tasks out of ten tasks included in a workflow. The motions to be detected are not limited to the above.

[0023] FIG. 3 shows an example of work performed by a worker according to a work procedure. "Start report" in the example of FIG. 3 indicates the work of a worker pressing a reporting button at the start of a work flow. "Work" in the example of FIG. 3 indicates, for example, a state in which a worker is performing work according to a work procedure. "End report" in the example of FIG. 3 indicates the work of a worker pressing a reporting button at the end of a work flow.

[0024] When the worker presses the reporting button for "reporting start" in the example of Fig. 3, the action of the worker pressing the button is captured in multiple frames. The detection model detects the button pressing action from each frame image that captures the worker's button pressing action in, for example, a video of the work area.

[0025] FIG. 4 shows an example of the detection results of a worker's button pressing action using a detection model. In the example of the detection results in FIG. 4 , the horizontal axis represents time. Furthermore, in the example of the detection results in FIG. 4 , the vertical lines representing detection peaks indicate that the button pressing action was detected by the detection model in the frame corresponding to each time on the time series. The example of the detection results in FIG. 4 illustrates, for example, an example in which a button pressing action is detected when the certainty of the identification by the detection model exceeds a reference value. The detection result by the detection model may also be a value of the certainty of the identification. In the example of the detection results in FIG. 4 , the reference point (start) and the reference point (end) indicate time points that serve as references for identifying the start and end of a task. The reference points will be described later. In the example of the detection results in FIG. 4 , the button pressing action is detected from multiple frame images included in both the time period when the button was pressed at the start of the task and the time period when the button was pressed at the end of the task. In this way, the detection model detects the same action from each of the multiple frame images.

[0026] For example, when the detection model detects the action of pressing a report button while no work is being performed, the detection unit 12 detects the button action detected by the detection model as the start of work. Furthermore, for example, when the same button is used to report the start and end of work, the detection unit 12 detects the button action after a set time has elapsed since the start of work was detected, and detects the detected button action as the end of work. The elapsed time from the start to identify the end of work is set, for example, by the person analyzing the work, based on a standard work time.

[0027] The detection model is, for example, a machine learning model using a neural network. The detection model is generated, for example, by learning the relationship between an image showing the target motion, the area showing the target motion, and the type of target motion. The detection model is generated, for example, in a system external to the video analysis system 10.

[0028] The detection unit 12 may detect a starting task and a finishing task using multiple detection models. For example, the detection unit 12 may use a detection model that detects the position of the worker's body from video to identify the position of the worker's hand in the frame image. Furthermore, the detection unit 12 may detect the tool held by the worker at the identified position using a detection model that detects the type of tool. The detection unit 12 may then detect the start of a task, for example, when the worker's hand is in a tool storage area and the tool being held is the tool designated for use in the first task of the workflow. Furthermore, the detection unit 12 may detect the end of a task, for example, when, after the task has started, the worker's hand is located in a tool storage area and the tool being held is the tool designated for use in the last task of the workflow. Furthermore, the detection unit 12 may detect the starting task and the finishing task using different detection models. Furthermore, the detection unit 12 may detect the task motion and the task motion using different detection models for each task included in the workflow.

[0029] The detection unit 12 may use the detection model to detect at least one of the worker's clothing and / or accessories. The detection unit 12 may use the detection model to detect the position where the worker is standing. The detection unit 12 may also use the detection model to detect the position where the worker is sitting. The detection unit 12 may also use the detection model to detect the worker's posture. The detection unit 12 may also use the detection model to detect the number of workers in the work area. The detection unit 12 may also use the detection model to detect whether a lamp indicating that work is in progress is on or off. The detection targets using the detection model are not limited to those described above.

[0030] For example, the identification unit 13 identifies, from a plurality of frames in which the detection model detects the motion of the detection target, a frame that is most likely to have been captured at the start of the task and a frame that is most likely to have been captured at the end of the task. The identification unit 13 then identifies the timings at which the identified frames were captured as the start of the task and the end of the task, respectively. For example, the identification unit 13 identifies, from among the frames detected by the detection model, a frame that was captured at the most likely position in the time series as the frame that was captured at the start of the task and the end of the task, respectively. The identification unit 13 identifies at least one of the start of the task and the end of the task based on a score indicating the relationship between a reference point in the time series and the position in the time series of the frame in which the motion of the detection target was detected. Furthermore, the identification unit 13 identifies at least one of the start of the task and the end of the task based on, for example, a score indicating the relationship between the reference point and the position of the frame and a degree of confidence in the detection of the motion of the detection target by the detection model.

[0031] The identification unit 13, for example, determines the timings of the start and end of a task by determining that the closer a timing is to a reference point set on the timeline, the more likely it is that the timing represents the start and end of the task. The reference point is a time point on the timeline that serves as a reference for calculating a score. For example, if the reference point is set to a time point when no task is being performed, a frame captured at a time point closer to the reference point is more likely to be an appropriate frame for the start of the task. Therefore, the identification unit 13 determines the timings of the start and end of the task based on, for example, the reference point on the timeline and the position at which the frame was captured. For example, the identification unit 13 calculates a score indicating the relationship between the reference point on the timeline and the position of the frame at which the motion of the target was detected, for each frame in which the target motion was detected. For example, the identification unit 13 calculates a score indicating the relationship between the reference point and the position of the frame such that the value decreases as the position of the frame becomes farther from the reference point. The identification unit 13 determines the frame with the highest calculated score. The identification unit 13 then determines that the timing at which the frame with the highest calculated score was captured is the start or end of the task.

[0032] The score indicating the relationship between the reference point and the position in the time series at which a frame in which the detection target movement is detected is captured is calculated, for example, based on the time difference between the reference point and the time at which the frame was captured. The identification unit 13 calculates the score, for example, based on the time difference between the time at which the frame image was captured and the time of the reference point. The identification unit 13 calculates the score so that, for example, the greater the time difference between the time at which the frame in which the score is calculated was captured and the time of the reference point, the lower the score value. The identification unit 13 may calculate the score, for example, based on the number of frames between the frame in which the score is calculated and the reference point. The identification unit 13 may calculate the score so that, for example, the greater the number of frames between the frame in which the score is calculated and the reference point, the lower the score value. The identification unit 13 may also calculate the score based on the chronological order of frames in which the detection target movement is detected after the reference point. The identification unit 13 calculates the score so that, for example, the greater the number indicating the chronological order of the frame in which the detection target movement is detected, the lower the score value.

[0033] The identification unit 13 may, for example, identify the timings of the start and end of a task by determining that, among frames in which the detection model detects the target motion, frames closer to the center in the time series are more likely to be frames captured at the start and end of the task, respectively. For example, in the case of a button-pressing action, the actions from touching the button, pressing the button, and then releasing the button are detected as the same action. In such a button-pressing action, frames captured closer to the center may be more likely to represent the start or end of the task. In this case, the identification unit 13 may, for example, calculate a score indicating the relationship between the reference point and the position of a frame such that, among frames in which the target motion is detected within a predetermined time from the reference point, the closer the frame is to the center in the time series, the higher the score. As a specific example, if a target frame is detected in a nine-frame image within a predetermined time, the identification unit 13 may calculate a score indicating the relationship between the reference point and the position of the frame such that the score of the fifth frame in the middle is the highest. Furthermore, for example, if there are multiple frames with the same score because the total number of frames is an even number, the identification unit 13 identifies the timing of the frame closest to the reference point as the start or end timing of the task. The predetermined time is set by the person analyzing the task, for example, based on the times when the start and end of the task are expected to be captured in the frame images. The identification unit 13 may also identify a frame close to the center in the time series based on the time when each frame image was captured.

[0034] When it is possible to predict likely frame positions in the time series as the start and end of a task among frames in which the same task is detected, the identification unit 13 may calculate the score so that a frame located at a predetermined position in the time series from a reference point has a higher score. For example, the identification unit 13 calculates the score so that a frame in which the target motion is detected has a predetermined order in the time series has a higher score. In this case, the identification unit 13 calculates the score so that the value decreases as the order in the time series of the target frame for score calculation becomes more distant from the predetermined order. Furthermore, the identification unit 13 may calculate the score so that the value increases at a position that is a predetermined time away from the reference point. In this case, the identification unit 13 calculates the score so that the value increases as the position in the time series of the target frame for score calculation becomes closer to a position that is a predetermined time away from the reference point. For example, when detecting a button pressing task, it is more appropriate to detect the actual button pressing timing, rather than the touching of the button, as the start or end of the task. In this case, for example, if the worker waits for the target product of the task to be automatically set before pressing the button, the timing at which the button will be pressed may be predictable. In such a case, the closer the timing is to a predetermined position on the timeline predicted from the task content, the higher the score, so that the identification unit 13 can more appropriately identify the timing at which the action to be detected was performed. The predetermined time used in calculating the score is set, for example, by the person analyzing the task based on the content of the action to be detected.

[0035] In the time series, frames in which the detection model detects the detection target movement may be concentrated near the time point at which the detection target movement is performed. In such a case, the identification unit 13 calculates the score such that, for example, the greater the number of frames in which the detection confidence of the detection target movement is equal to or greater than a reference value within a predetermined number of frames from the frame for which the score is to be calculated, the higher the score of the frame for which the calculation is to be performed. For example, the identification unit 13 calculates the score indicating the relationship between the reference point and the position of the frame based on the number of frames in which the detection target movement is detected in the frame to be calculated and frames before and after it in the time series. For example, the identification unit 13 calculates the score based on the number of frames in which the detection target movement is detected within a predetermined number of frames before and after the frame to be calculated in the time series.

[0036] For example, the identification unit 13 calculates the score so that the value increases as the number of frames in which the detection target movement is detected within a predetermined number of frames before and after the frame to be calculated in the time series increases. For example, if the predetermined number of frames is set to 3, the identification unit 13 calculates the score so that the value increases as the number of frames in which the detection target movement is detected within three frames before and after the frame to be calculated in the time series increases.

[0037] The identification unit 13 may identify the start and end timings of a task using the score of the frame for which the score is to be calculated and the scores of frames within a predetermined number of frames in the time series from the frame for which the score is to be calculated. For example, the identification unit 13 may calculate a score indicating the relationship between the reference point and the position of a frame as the average value of the frame for which the score is to be calculated and frames within a predetermined number of frames before and after the frame in the time series. When the target action is performed, frames showing the target action may be detected consecutively in the time series. Therefore, by reflecting the detection status of the target action in frames adjacent to the frame for which the score is to be calculated in the score, the accuracy of identifying the start and end timings of the task may be improved. Furthermore, the criteria for calculating the score are not limited to those described above.

[0038] For example, the identification unit 13 sets a frame in a time series included in the video of the work area at which the detection unit 12 starts the process of detecting the motion of the detection target as a reference point. The identification unit 13 may set a time point in a time series included in the video of the work area that is a predetermined time after the start of the work as a reference point for identifying the end of the work. The predetermined time from the start of the work is set by the person analyzing the work based on, for example, the time required for the work. The reference point may also be set using data from a production management system. For example, the identification unit 13 sets a time point in the video that corresponds to the start and end times of the work recorded in the production management system as a reference point. The reference point may also be set appropriately depending on the analysis content and the analysis target.

[0039] When the work to be analyzed involves repeatedly performing the same task, the identification unit 13 may set the reference point for detecting the start of each repeated task to the timing at which the previous task is completed. Furthermore, when detecting the start of a repeated task, the identification unit 13 may set the reference point to a point in time when a predetermined time has elapsed since the start or end of the previous task. The predetermined time may be set based on the number of frames. The predetermined time from the start or end of the previous task is set by the person analyzing the task, for example, based on the time required for one task.

[0040] Furthermore, when determining the start and end timings of a task based on the score and the confidence level, the confidence level of the detection of the target motion is, for example, an index indicating the accuracy of the detection model's determination of the target motion. The confidence level of the detection of the target motion is, for example, the probability that the motion captured in the frame image is the target motion, calculated by the detection model. The detection model identifies the motion captured in the frame image as the target motion if, for example, the probability is equal to or greater than a standard. The standard for identifying the motion captured in the frame image as the target motion is set, for example, by the person who generates the detection model.

[0041] The identification unit 13 may identify the timings of the start and end of a task using the position of a frame in the time series and the confidence of detection by the detection model. For example, by using both the confidence of the position of the frame in the time series in which the motion of the detection target is detected and the confidence of detection by the detection model, the timings of the start and end of the task can be identified more accurately. The identification unit 13 identifies the timings of the start and end of a task based on, for example, a score indicating the relationship between the reference point and the position of the frame and an index calculated from the confidence of detection of the motion of the detection target by the detection model. The identification unit 13 identifies the timings of the start and end of a task based on, for example, an index calculated by multiplying the score indicating the relationship between the reference point and the position of the frame by the confidence of detection of the motion of the detection target by the detection model.

[0042] Here, the position score, which is a score indicating the relationship between the reference point and the position of the frame, is assumed to be Sp. The certainty of detection of the motion of the detection target is assumed to be P. The index calculated from the position score, which is a score indicating the relationship between the reference point and the position of the frame, and the certainty of detection of the motion of the detection target by the detection model is assumed to be the total score St. In this case, the identification unit 13 calculates the total score St using, for example, the formula St = Sp × P. Then, the identification unit 13 determines that the frame with the largest total score St among the frames in which the motion of the detection target is detected is the frame where the task starts and ends, and identifies the timings of the start and end of the task.

[0043] The identification unit 13 may calculate a total score St based on the detected clothing and accessories of the worker to identify the timing of the start and end of the work. For example, the identification unit 13 calculates the total score St so that the value is high when at least one of the clothing and accessories of the worker is the clothing and accessories worn when performing the work. The identification unit 13 may also calculate the total score St so that the value is high when the position where the worker is standing or sitting is a location designated as a work location. The identification unit 13 may also calculate the total score St so that the value is high when the number of workers in the work area matches the number of people performing the work. The identification unit 13 then determines the timing of the start and end of the work by determining that the frames with the largest calculated total score St are the frames where the work starts and ends.

[0044] When the detection model detects the on / off of a lamp indicating that work is being performed, the identification unit 13 may calculate the total score St so that the value of the frame in which the lamp is on is higher. For example, when identifying the timing of the start of work, the identification unit 13 calculates the total score St so that the value of the frame in which a lamp that was off in the previous frame in the time series is on is higher. Furthermore, when identifying the timing of the end of work, the identification unit 13 calculates the total score St so that the value of the frame in which a lamp that was on in the time series is off in the subsequent frame is higher.

[0045] FIG. 5 shows an example of the time-series positions of frames in which the target motion is detected among multiple frames in the time series. In the example of FIG. 5, the horizontal axis represents time. In addition, in the example of FIG. 5, peaks represented by vertical lines indicate that the target motion has been detected by the detection model in the frame corresponding to each time in the time series. In the example of FIG. 5, peak numbers are added to distinguish peaks indicating that the target motion has been detected. In the example of FIG. 5, the larger the peak number, the farther the position from the reference point, and therefore the smaller the score indicating the relationship between the reference point and the frame position.

[0046] FIG. 6 shows an example of the total score for each frame when the target motion is detected from frame images at the time-series positions shown in the example of FIG. 5 . The peak numbers in the example of FIG. 6 correspond to the peak numbers in the example of FIG. 5 . The example of FIG. 6 shows the score, certainty, and total score for the frames corresponding to each peak number. In the example of FIG. 6 , the frame with peak number 1, which is closest to the reference point, has the highest score. Also, in the example of FIG. 6 , the frame with peak number 2 has the highest certainty. Also, in the example of FIG. 6 , the frame with peak number 2 has the highest total score, which is the product of the score and the certainty. Therefore, the identification unit 13 identifies the timing when the frame with peak number 2 was captured as the start or end of the task.

[0047] The identification unit 13 may weight at least one of the position score and the certainty of detection by the detection model to identify the timing of the start and end of the work. For example, the identification unit 13 may weight at least one of the position score and the certainty of detection by the detection model to calculate a total score. Furthermore, the identification unit 13 may calculate a total score for a target frame using statistics of frames before and after the target frame.

[0048] The output unit 14 outputs the timing identified by the identification unit 13. The output unit 14 outputs at least one of the timings of the start and end of the work identified by the identification unit 13, for example, to a terminal device operated by a person analyzing the work. The output unit 14 may output video recording the timings of the start and end of the work. The output unit 14 may output the timings identified by the identification unit 13 to a display device (not shown) connected to the video analysis system. The output unit 14 may output video between the start and end of the work from among videos captured of the work area. The output unit 14 may also output a work time calculated from the timings of the start and end of the work. The output unit 14 may also output images of frames identified as the timings of the start and end of the work, and information indicating that the images of the frames identified as the timings of the start and end of the work, together with images of the frames before and after the start and end of the work in chronological order.

[0049] The storage unit 15 stores, for example, video footage of the work area. The storage unit 15 also stores a detection model. The detection model may be stored in a storage means external to the video analysis system 10. The storage unit 15 stores, for example, the results of identifying the timings of the start and end of work. The storage unit 15 may store, from the video footage of the work area, video footage between the start and end of work. The storage unit 15 may store images of frames identified as the timings of the start and end of work, as well as images of frames preceding and following them in chronological order.

[0050] The following describes the operation of the video analysis system 10 to identify the start of a task and the timing of the start of the task. Fig. 7 shows an example of the operation flow when the video analysis system 10 identifies the start of a task and the timing of the start of the task.

[0051] The acquisition unit 11 acquires a video of the work area (step S11). For example, the acquisition unit 11 acquires a video of the work area captured by a camera installed so as to be able to capture the work area.

[0052] When the video of the work area is acquired, the detection unit 12 uses a detection model that detects the movement of the detection target shown in the frames included in the video to detect the movement of the detection target from each of the time-series frames included in the video of the work area (step S12).

[0053] When the detection target movement is detected, the identification unit 13 identifies the timing of the start and end of the work based on a score indicating the relationship between a reference point on the time series and the position of a frame in which the detection target movement is detected (step S13). The identification unit 13 may identify either the timing of the start or the timing of the end of the work based on the score indicating the relationship between a reference point on the time series and the position of a frame in which the detection target movement is detected.

[0054] When the timings of the start and end of the work have been identified for all of the repeated work in step S13 (Yes in step S14), the output unit 14 outputs the timings of the start and end of the work identified by the identification unit 13 (step S15).

[0055] If the timing of the start and end of the work for all of the repeated work has not been determined in step S13 (No in step S15), the process returns to step S13, and the determination unit 13 determines at least one of the timing of the start and end of the work based on a score indicating the relationship between the reference point on the time series and the position of the frame in which the action to be detected is detected.

[0056] The video analysis system 10 uses a detection model to detect target movements from each of a time series of frames included in images of a work area. The target movements are, for example, movements performed at the start and end of a task. The video analysis system 10 identifies the start and end timings of a task based on a score indicating the relationship between a reference point in the time series and the position of the frame in which the target movement was detected. In this way, by using the video analysis system 10, the start and end timings of a task can be easily identified.

[0057] Furthermore, the video analysis system 10 can determine the start and end timings of a task based on the detection confidence of the detection model and a score related to the frame position at which the target action was detected, thereby using the confidence determined by the detection model and the confidence based on the frame position to determine the start and end timings of a task. Therefore, when the detection confidence of the detection model and the score related to the frame position at which the target action was detected are used to determine the start and end timings of a task, the video analysis system 10 can more accurately determine the start and end timings of a task.

[0058] Each process in the video analysis system 10 can be realized by executing a computer program on a computer. Fig. 8 shows an example of the configuration of a computer 100 that executes a computer program that performs each process in the video analysis system 10. The computer 100 includes a CPU (Central Processing Unit) 101, a memory 102, a storage device 103, an input / output I / F (Interface) 104, and a communication I / F 105.

[0059] The CPU 101 reads and executes computer programs for performing each process from the storage device 103. The CPU 101 may be configured as a combination of multiple CPUs. The CPU 101 may also be configured as a combination of a CPU and another type of processor. For example, the CPU 101 may be configured as a combination of a CPU and a graphics processing unit (GPU). The memory 102 is configured with a dynamic random access memory (DRAM) or the like, and temporarily stores computer programs executed by the CPU 101 and data being processed. The storage device 103 stores computer programs executed by the CPU 101. The storage device 103 is configured with, for example, a non-volatile semiconductor storage device. Other storage devices such as a hard disk drive may also be used for the storage device 103. The input / output I / F 104 is an interface that receives input from an operator and outputs display data, etc. The communication I / F 105 is an interface that transmits and receives data to and from other information processing devices.

[0060] The computer program used to execute each process can also be stored and distributed on a computer-readable recording medium that non-temporarily stores data. Examples of recording media that can be used include magnetic tapes for recording data and magnetic disks such as hard disks. Optical disks such as CD-ROMs (Compact Disc Read Only Memory) can also be used as recording media. Non-volatile semiconductor storage devices can also be used as recording media.

[0061] The present disclosure has been described above using the above-described embodiments as examples. However, the present disclosure is not limited to the above-described embodiments. That is, the present disclosure can be applied in various aspects that can be understood by a person skilled in the art within the scope of the present disclosure.

[0062] REFERENCE SIGNS LIST 10 Video analysis system 11 Acquisition unit 12 Detection unit 13 Identification unit 14 Output unit 15 Storage unit 100 Computer 101 CPU 102 Memory 103 Storage device 104 Input / output I / F 105 Communication I / F

Claims

1. An acquisition means for acquiring an image of a work area; a detection means for detecting a motion of a detection target from each of time-series frames included in the video of the work area using a detection model for detecting a motion of a detection target shown in frames included in the video; an identification means for identifying at least one of the start and end timings of a task based on a score indicating a relationship between a reference point on a time series and a position on the time series of a frame in which the motion of the detection target is detected; an output means for outputting the specified timing; A video analysis system comprising:

2. the identification means identifies the timing based on the score and a degree of certainty of detection of the motion of the detection target by the detection model. The video analysis system of claim 1 .

3. the specifying means specifies the timing based on an index calculated by multiplying the score and the confidence level. The video analysis system of claim 2 .

4. the specifying means calculates the score such that the score of the frame to be calculated increases as the number of frames in which the certainty of detection of the motion to be detected is equal to or greater than a reference value increases within a range of a predetermined number of frames from the frame to be calculated; The video analysis system according to claim 2 .

5. the specifying means specifies the timing using a score of a frame for which a score is to be calculated and scores of frames within a predetermined number of frames in time series from the frame for which a score is to be calculated; 5. The video analysis system according to claim 1.

6. The score becomes lower as the position of the frame becomes farther from the reference point.

5. The video analysis system according to claim 1.

7. The score is a score based on the chronological order of frames in which the motion of the detection target is detected after the reference point.

5. The video analysis system according to claim 1.

8. The score is based on the elapsed time from the reference point to the frame in which the motion of the detection target is detected or the number of frames in the time series. The video analysis system according to any one of claims 1 to 3.

9. Obtain footage of the work area, detecting a motion of a detection target from each of the time-series frames included in the video of the work area using a detection model that detects a motion of a detection target shown in frames included in the video; identifying at least one of the start and end timings of the task based on a score indicating a relationship between a reference point on the time series and a position on the time series of a frame in which the motion of the detection target is detected; Output the specified timing, Video analysis methods.

10. Acquiring video of the work area; a process of detecting a motion of a detection target from each of frames in a time series included in a video of the work area using a detection model that detects a motion of a detection target shown in frames included in the video; a process of identifying at least one of the start and end timings of the task based on a score indicating a relationship between a reference point on the time series and a position on the time series of a frame in which the motion of the detection target is detected; A process for outputting the specified timing. A video analysis program run on a computer.