Mobile analysis system

The mobile object analysis system enhances monitoring accuracy by combining bird's-eye and first-person views to estimate movement data, addressing positional limitations of fixed cameras and improving workplace efficiency.

JP2025166462APending Publication Date: 2025-11-06TOYOTA JIDOSHA KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024070530
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-24
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Existing monitoring systems struggle with accurately estimating dynamic data of mobile objects at work sites due to positional relationships between fixed cameras and individuals, leading to reduced monitoring accuracy and incomplete data on work activities.

Method used

A mobile object analysis system utilizing a first camera for a bird's-eye view and a second camera attached to the moving object, combining data from both perspectives to estimate movement amount and content, with machine-learned models for precise calculations.

Benefits of technology

Accurately estimates movement data of mobile objects, enabling identification of work inefficiencies such as overburden and waste, thereby improving workplace efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025166462000001_ABST
    Figure 2025166462000001_ABST
Patent Text Reader

Abstract

To provide a mobile analysis system capable of improving detection accuracy for dynamic data on a moving body in a work site etc.SOLUTION: The present invention relates to a mobile analysis system that detects a moving body included in video data and then analyzes data related to the operation of the moving body, and the mobile analysis system comprises: a first camera which acquires first video data generated by taking a bird's-eye photograph of the moving body; a second camera which is mounted on the moving body and acquires second video data on a view from the moving body; an operation quantity estimation part which estimates an operation quantity of the moving body based upon the first video data; an operation estimation part which estimates operation details of the moving body of the second video data; and an operation analysis part which analyzes data related to the operation of the moving body based upon the operation quantity of the moving body obtained when it is estimated that operation details show that the moving body is moving.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a system for analyzing the state of a moving object contained in video data. [Background technology]

[0002] Patent Document 1 discloses a monitoring system that monitors people, objects, or equipment at a manufacturing site where multiple machines are installed and people work and move around. The monitoring system in Patent Document 1 includes cameras that monitor people, objects, and equipment, an analysis server that analyzes video data acquired by the cameras, and a display computer that displays the analysis results from the analysis unit. In the monitoring system in Patent Document 1, a human monitoring camera that monitors people stores information such as the person's physique, clothing shape, and movement characteristics. Images of people contained in the video data are then identified by comparing this information with the video data. The analysis server is configured to analyze the type of work performed by the identified person, the time period during which the work was performed, the travel time, or whether the person was in a location outside the captured range, based on significant changes in the person's movement based on learned data.

[0003] Patent Document 2 discloses an action recognition device aimed at accurately estimating the actions of an actor. The action recognition device of Patent Document 2 is configured to recognize a person's actions by analyzing video data captured by a first camera placed on a ceiling, wall, or the like to capture an image of a monitored area and a second camera worn on the person's head. The action recognition device of Patent Document 2 classifies the video data into an action area that is closely related to the person's actions and a surrounding area. Then, an action label distribution is estimated based on environmental information that expresses the person's executable actions in a position overlapping the action area, environmental information in a position overlapping the surrounding area, and video feature information extracted from the video data. In other words, the action recognition device of Patent Document 2 narrows down and estimates candidates for the person's actions, thereby enabling accurate estimation of the person's actions. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent Publication No. 2021-196909 [Patent Document 2] International Publication No. 2023 / 100330 Summary of the Invention [Problem to be solved by the invention]

[0005] According to Patent Document 1, manufacturing site monitoring can be efficiently performed without an operator having to monitor the movements of people, objects, and equipment or input work status information. On the other hand, the camera that monitors the movements of people, objects, and equipment in the device of Patent Document 1 is a fixed camera installed at the work site. Therefore, depending on the distance of the person from the camera and the orientation of the person relative to the camera, it may be impossible to accurately determine the person's movements. For example, if the person has their back to the people monitoring camera, it may be impossible to accurately determine whether the person is working or walking. Thus, with the monitoring system of Patent Document 1, depending on the positional relationship between the people monitoring camera and the person, the monitoring accuracy may be reduced, and accurate data about the person's work may not be obtained.

[0006] Patent Document 2 describes a method for recognizing a person's behavior by analyzing video data captured by a first camera that captures an area to be monitored, including the person, and a second camera that captures the direction the person is looking. Since the second camera captures the direction the person is looking, it is possible to relatively accurately determine whether the person is working. However, the device in Patent Document 2 is configured to recognize a person's behavior by analyzing the video data captured by the first and second cameras. Therefore, for example, when analyzing data related to a person's walking, even if the second camera can determine that the person is walking, it may not be able to accurately estimate the distance traveled by the person. Thus, even with the device in Patent Document 2, although it can recognize a person's behavior in the video data, it may not be able to accurately estimate or analyze the dynamic data of the recognized person.

[0007] This invention has been made with an eye on the above-mentioned technical problems, and aims to provide a mobile object analysis system that can improve the accuracy of estimating dynamic data of mobile objects at work sites, etc. [Means for solving the problem]

[0008] In order to achieve the above-mentioned object, the present invention provides a mobile object analysis system that detects a moving object contained in video data and analyzes data related to the movement of the moving object, and is characterized by comprising: a first camera that acquires first video data taken from a bird's-eye view of the moving object; a second camera that is attached to the moving object and acquires second video data seen from the moving object; a movement amount estimation unit that estimates the movement amount of the moving object based on the first video data; a movement estimation unit that estimates the movement content of the moving object based on the second video data; and a movement analysis unit that analyzes data related to the movement of the moving object based on the movement amount of the moving object when the movement content is estimated to be moving.

[0009] In addition, this invention may further include a movement amount analysis unit that calculates the movement amount of the moving body quantitatively expressed based on the movement amount of the moving body, wherein the movement estimation unit is configured to estimate the movement time of the moving body based on a learning model that has been machine-learned about the moving body, and the movement analysis unit is configured to calculate the movement speed of the moving body based on the movement amount of the moving body and the movement time of the moving body.

[0010] In addition, this invention may further include a feature determination unit that identifies the moving body to which the second camera is attached from the first video data based on at least one parameter of the position, orientation, and movement of the moving body.

[0011] The motion analysis unit in the present invention may be configured to analyze a location where the data relating to the movement of the moving object was acquired, based on the first video data and the second video data. [Effects of the Invention]

[0012] A mobile object analysis system according to an embodiment of the present invention includes a first camera that captures a moving object from a bird's-eye view and a second camera attached to the moving object that acquires second video data viewed from the moving object. The amount of movement of the moving object is estimated based on the first video data captured by the first camera. Furthermore, the content of the moving object's movement is estimated based on the second video data captured by the second camera. The system is configured to analyze data related to the movement of the moving object based on the amount of movement and the content of movement of the moving object. Because the first video data captures the moving object from a bird's-eye view, the amount of movement of the moving object can be accurately estimated. Furthermore, because the second video data captures the moving object's viewpoint, the content of the moving object's movement can be accurately estimated. Therefore, by analyzing the moving object based on the amount of movement and the content of movement of the moving object estimated from each video data, highly accurate movement data can be obtained. Therefore, for example, the gait of a worker at a workplace or the like can be accurately analyzed, which can accurately identify areas for improvement, such as overstress, waste, and unevenness, at the workplace, and thereby be utilized to improve work efficiency at the workplace.

[0013] Furthermore, if the system is configured to calculate the moving speed of the moving body based on the amount of movement that quantitatively represents the amount of movement of the moving body and the moving time of the moving body estimated based on a machine-learned learning model, data regarding the movement of the moving body can be obtained with even greater accuracy. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is an explanatory diagram for explaining a state in which video data to be analyzed by a moving object analysis system according to an embodiment of the present invention is being acquired; [Figure 2] 1 is a block diagram for explaining the functional configuration of a moving object analysis system according to an embodiment of the present invention. [Figure 3] 4 is a flowchart illustrating an example of control executed by the moving body analysis system according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0015] Next, the present invention will be described based on the embodiments shown in the drawings. Note that the embodiments described below are merely examples of specific embodiments of the present invention, and are not intended to limit the present invention.

[0016] In a mobile object analysis system 1 according to an embodiment of the present invention, a worker (mobile object) 3 is filmed working or moving at a work site 2 where line work or the like is being performed, and the filmed video data is analyzed to analyze the movements of the worker 3. Fig. 1 shows a schematic diagram of an example of a work site 2 that is the subject of an embodiment of the present invention.

[0017] At a work site 2 shown in Fig. 1, a worker 3 performs a predetermined task, such as assembling parts, on a work object 5 (workpiece) that has been transported by a transport facility 4. Although not shown in the figure, the system is configured so that as the work object 5 is transported, the position where the worker 3 works also moves in the same direction as the transport direction of the work object 5. As shown in Fig. 1, the mobile object analysis system 1 mainly comprises a wearable camera 6, a fixed camera 7, and an information processing device 8.

[0018] The wearable camera 6 (head-mounted camera) is an eye-level camera attached to the worker 3, configured to capture images in the direction the worker 3 is facing. The wearable camera 6 is attached to the worker 3 so as to capture images in the direction the worker 3 is looking. For example, as shown in FIG. 1 , the camera may be attached to the ear of the worker 3 with the capture direction facing forward, or may be attached so that the lens portion is located at the temple of glasses or between the worker 3's eyebrows. The camera may be configured similarly to a conventionally known camera, such as an RGB camera. First-person video data from the worker's perspective captured by the wearable camera 6 can be output to an information processing device 8. The wearable camera 6 corresponds to the second camera in the embodiment of the present invention. The first-person video data corresponds to the first video data in the embodiment of the present invention.

[0019] Fixed camera 7 is an imaging device configured to capture an image of a predetermined area in work site 2 from above. Fixed camera 7 is fixed to, for example, the ceiling of work site 2, and is positioned so as to capture an image of the entire work site 2 from a bird's-eye view. Like wearable camera 6, fixed camera 7 may be configured similarly to a conventionally known camera, such as an RGB camera, and is configured to be able to output captured image data and video data. Third-person video data obtained by capturing an overhead image of worker 3 by fixed camera 7 is output to information processing device 8 in the same manner as first-person video data. Note that fixed camera 7 corresponds to the first camera in this embodiment of the present invention. Furthermore, third-person video data corresponds to the second video data in this embodiment of the present invention.

[0020] The information processing device 8 mainly comprises a processor, a communication unit, a storage unit, etc. The information processing device 8 is configured to perform calculations according to a predetermined program using data acquired from the outside and pre-stored data, and to output the results of the calculations as control command signals. For example, the information processing device 8 executes functions that meet predetermined purposes by having the processor load a program stored on a recording medium into a working area of ​​the storage unit and execute the program, and perform various controls through the execution of the program.

[0021] The processor is, for example, a CPU or a DSP. This processor is configured to control the information processing device 8 and perform various information processing operations. The main memory includes, for example, a RAM and a ROM. As described above, the main memory has a work area for the processor to execute programs. The auxiliary memory includes, for example, an EPROM or a hard disk drive. This auxiliary memory may also include a portable recording medium, i.e., a removable medium. The auxiliary memory also freely stores various programs, various data, and various tables in the recording medium by reading or writing them. The auxiliary memory may also store an operating system. The communication unit is a wireless communication circuit connected to an external communication device via wireless communication so as to be able to communicate data with the external communication device.

[0022] Next, a functional configuration of the information processing device 8 in an embodiment of the present invention will be described. The mobile object analysis system 1 configured as described above is configured to be able to analyze data related to the walking of the worker 3. As such a configuration, the information processing device 8 in an embodiment of the present invention includes, as shown in FIG. 2, an image acquisition unit 9, an object detection unit 10, a walking estimation unit 11, an exercise amount estimation unit 12, a movement amount analysis unit 13, a feature determination unit 14, a motion analysis unit 15, and an output unit 16.

[0023] Video acquisition unit 9 acquires image data or video data captured by wearable camera 6 and fixed camera 7. Video acquisition unit 9 acquires first-person video data captured from the first-person perspective of worker 3 by wearable camera 6, and third-person video data captured of worker 3 from the third-person perspective by fixed camera 7. Note that these video data may be configured so that arbitrary video data is input by the user, or may be configured so that they are automatically captured from each camera.

[0024] The object detection unit 10 detects the worker 3 from the acquired third-person video data based on learning data such as previously learned feature quantities of the worker 3 and objects and devices located at the work site 2. The object detection unit 10 is configured to detect the worker 3 from image frames constituting the video data using a deep learning method such as conventionally known YOLO (YOLOv8) or SSD. For example, the object detection unit 10 detects the worker 3 using methods such as skeletal estimation, object detection, or object region estimation. The object detection unit 10 also detects objects and devices included in the first-person video data. For example, it detects what item the worker 3 has grasped from the workbench, or what type of device the worker 3 has operated.

[0025] The gait estimator 11 estimates and separates the actions of the worker 3 from the first-person video data acquired from the wearable camera 6. The gait estimator 11 estimates the content of the worker's actions, such as whether the worker 3 is working or moving (walking), based on the first-person video data. For example, the gait estimator 11 estimates that the worker 3 is walking by detecting the direction the worker 3 is looking, the speed at which the scenery changes, or the shaking of the video data. Alternatively, when the gait estimator 11 detects the hand or arm of the worker 3, the gait estimator 11 detects that the worker 3's elbow is extended, then the fingertips are bent, and then the worker grasps an object and the elbow is bent again, thereby estimating that the worker 3 is working while grasping an object. The gait estimator 11 estimates or separates the actions of the worker 3 based on a learning model that has previously learned training data through machine learning. The gait estimator 11 corresponds to the motion estimator in the embodiment of the present invention.

[0026] The momentum estimation unit 12 estimates the momentum of the worker 3 from the third-person video data acquired from the fixed camera 7. The momentum estimation unit 12 estimates the momentum of the worker 3 by two-dimensional or three-dimensional skeletal estimation. For example, in skeletal estimation, body parts of the worker 3 are recognized from image frames constituting the video data, and the skeletal structure is detected in two dimensions by connecting the body parts to estimate the posture of the worker 3. For example, it is possible to estimate whether the worker 3 is walking, holding an object, or performing work using a device. Such skeletal estimation estimates that the worker 3 is walking from the movement of the worker's feet and body. The momentum estimation unit 12 estimates the momentum of the worker 3 when performing an action that is estimated to be walking. The momentum estimation unit 12 corresponds to the motion estimation unit in the embodiment of the present invention.

[0027] The movement amount analysis unit 13 quantitatively expresses the amount of movement of the worker 3 based on the amount of movement of the worker 3 acquired by the movement amount estimation unit 12. The movement amount analysis unit 13 calculates the amount of movement of the worker 3's feet based on, for example, the positions of the worker's left and right feet relative to the position of the worker's waist, the distance between the tips of the worker's left and right feet, and the amount of left-right deviation relative to the direction perpendicular to the ground. Then, based on various parameters obtained by analyzing the calculated amount of movement, the movement amount analysis unit 13 quantifies the walking time and movement amount of the worker 3.

[0028] The feature determination unit 14 determines whether the movements of the worker 3 detected or estimated from the first-person video data and the third-person video data match. For example, it determines whether the position, orientation, or movements of the worker 3 at a given time match between the first-person video data and the third-person video data. That is, in the case of first-person video data, the position, looking direction, or movement of the worker 3 can be estimated from the type, size, and changes in the objects and devices included in the video data. In addition, in the case of third-person video data, the position, looking direction, or movement of the worker 3 can be estimated from the facial orientation, standing position, and changes in the facial orientation and movements of the worker 3 included in the video data.

[0029] Therefore, feature determination unit 14 determines that the worker 3 in the first-person video data matches the worker 3 in the third-person video data when at least one feature amount among parameters such as the position, orientation, and movement of the worker 3 included in each piece of video data matches at a predetermined timing. In other words, it determines that the wearable camera 6 that acquired the first-person video data is attached to the specified worker 3 detected from the third-person video data, and matching of the workers 3 is performed. Note that if multiple workers 3 are located close to each other in the third-person video data, the system may be configured to determine that all of the above parameters match.

[0030] Conversely, if, at a predetermined timing, the feature quantities of any of the parameters, such as the position, orientation, and movement, of the worker 3 included in the first-person video data and the worker 3 included in the third-person video data differ, it is determined that the worker 3 included in the first-person video data and the worker 3 included in the third-person video data are different. In other words, it is determined that the wearable camera 6 that acquired the first-person video data is not attached to the specific worker 3 detected from the third-person video data. For example, this may occur if the worker 3 included in the first-person video data is estimated to be walking, while the worker 3 included in the third-person video data is estimated to be working. In this case, feature determination unit 14 detects another worker 3 from the third-person video data whose behavior matches the behavior of the worker 3 estimated based on the first-person video data.

[0031] The feature determination unit 14 may change which of the above-mentioned parameters to use depending on the number and positions of the workers 3 included in the third-person video data. For example, if two workers 3 are working close to each other in the third-person video data, it may be impossible to distinguish between the two workers 3. In such cases, the device may be configured to determine whether multiple parameters, such as the orientation and movement of the workers 3, match. In such cases, matching of the workers 3 may be performed by tracing back the video data and starting analysis from a point where the two workers 3 can be reliably distinguished, or by manually identifying the workers 3.

[0032] The motion analysis unit 15 analyzes data related to the walking of the worker 3 based on the data obtained by analyzing the first-person video data and the third-person video data as described above. The motion analysis unit 15 acquires the time during which the worker 3 is estimated to be walking from the walking estimation unit 11. In other words, it acquires the time during which it is determined that the worker 3 is walking based on the learning model. The motion analysis unit 15 also acquires the distance traveled by the worker 3 by walking, which is calculated by the exercise amount estimation unit 12 and the movement amount analysis unit 13.

[0033] The motion analysis unit 15 analyzes the walking speed (movement speed) of the worker 3 from the walking time and walking distance from the walking data acquired in this way. The motion analysis unit 15 also identifies the location where that walking speed occurred. That is, the motion analysis unit 15 estimates the location in the work site 2 where the worker 3 walked at that speed. In this case, the motion analysis unit 15 also analyzes the position of the worker 3 based on the data acquired from the walking estimation unit 11 and the exercise amount estimation unit 12.

[0034] The output unit 16 displays the results of the analysis by the motion analysis unit 15. The output unit 16 is configured to display, for example, symbols indicating the walking start position and walking end position of the worker 3 on a map representing the work site 2, as well as the walking speed between those two points. The output unit 16 may be configured, for example, by a PC monitor or a display of a mobile terminal.

[0035] Next, an example of control executed by the mobile object analysis system 1 configured as described above will be described with reference to Fig. 3. The flowchart shown in Fig. 3 illustrates the process of acquiring data related to the walking of worker 3 based on first-person video data captured by wearable camera 6 and third-person video data captured by fixed camera 7.

[0036] As shown in FIG. 3, in step S1, first, video data to be analyzed is input by mobile object analysis system 1. In step S1, first-person video data captured by wearable camera 6 and third-person video data captured by fixed camera 7 are acquired. The first-person video data and third-person video data may be configured to be input automatically, or may be configured to be input manually by a user or the like. When each piece of video data is input automatically, the system may be configured so that the user selects the video data to be analyzed from the video data stored in information processing device 8 and executes the analysis.

[0037] After the video data to be analyzed is input, the process proceeds to step S2. In step S2, matching of a worker 3 is performed based on the acquired first-person video data and third-person video data. In step S2, matching is performed between a specific worker 3 to be analyzed that is included in the third-person video data and the first-person video data captured by the wearable camera 6 worn by that specific worker 3.

[0038] For example, in step S2, the motions of the worker 3 estimated from the first-person video data and the third-person video data at a predetermined timing are compared and matched. Specifically, it is determined whether at least one of the facing direction, standing position, motion, etc. of the predetermined worker 3 estimated from the third-person video data matches with the facing direction, standing position, motion, etc. of the worker 3 estimated from the first-person video data. If, as a result of this comparison, any of the position or motion of the predetermined worker 3 estimated from the third-person video data does not match with the position or motion of the worker 3 estimated from the first-person video data, this flowchart is temporarily terminated without executing the subsequent control.

[0039] Conversely, if the position and movement of the specified worker 3 estimated from the third-person video data match the position and movement of the worker 3 estimated from the first-person video data, matching of the worker 3 is completed, and the process proceeds to step S3. In step S3, the position, orientation, or movement of the worker 3 estimated in step S2 is acquired. Furthermore, since step S2 only acquires the data necessary for matching the worker 3, in step S3, the data acquired in step S2 is acquired as more accurate parameters, or if there are parameters that have not been acquired, those parameters are acquired.

[0040] After acquiring data on the position, orientation, and movement of the worker 3, the process proceeds to step S4. In step S4, the behavior or work of the worker 3 is estimated based on the first-person video data. That is, in step S4, the parts where the worker 3 is walking and the other parts are estimated from the first-person video data. For example, in step S4, the behavior of the worker 3 is estimated from the first-person video data based on a learning model of the worker 3 stored in advance. This makes it possible to accurately distinguish whether the worker 3 is walking or working using the same criteria. Furthermore, by distinguishing the behavior of the worker 3, the time the worker 3 is walking can be determined.

[0041] After the behavior of the worker 3 is estimated, the process proceeds to step S5. In step S5, data related to the walking of the worker 3 is acquired from the third-person video data. In step S5, data related to the walking of the worker 3 is acquired by performing skeletal structure estimation based on the third-person video data. For example, in step S5, parameters such as the positions (coordinates) of the worker 3's hips and feet, the angle between the feet, and stride length are acquired by skeletal structure estimation. Then, the amount of exercise or movement of the worker 3 is calculated based on the acquired parameters.

[0042] After acquiring the data on the amount of exercise of the worker 3, the process proceeds to step S6. In step S6, the amount of exercise and the amount of movement of the worker 3 are quantified based on the data acquired in step S5. That is, in step S6, the distance actually traveled by the worker 3 while walking is calculated.

[0043] After the walking distance of worker 3 is calculated, the process proceeds to step S7. In step S7, the walking time of worker 3 and the corresponding walking distance are calculated based on the walking time of worker 3 calculated in step S4 and the movement distance (walking distance) of worker 3 calculated in step S6. That is, in step S7, the walking time of worker 3 obtained by analyzing the first-person video data and the walking distance of worker 3 obtained by analyzing the third-person video data are matched with each other. In this way, in step S7, when worker 3 starts and stops walking multiple times in each video data, the time required for each walk and the distance traveled at that time are stored in correspondence with each other.

[0044] After data related to the walking of worker 3 is acquired from each video data, the process proceeds to step S8. In step S8, the walking speed and walking positions of worker 3 are determined based on the acquired data. In step S8, the walking speed of worker 3 is determined based on the walking time and walking distance of worker 3 acquired in step S7. In addition, in step S8, the places or points where worker 3 walked at that walking speed are acquired. That is, in step S8, the determined walking speed is matched to the walking speed when worker 3 moves between points at work site 2 based on the position and orientation of worker 3 acquired in step S3. This makes it possible to determine the correspondence between a specific section at work site 2 and the walking speed when walking in that section.

[0045] After associating the predetermined section at the work site 2 with the walking speed when the worker 3 walked the predetermined section, the process proceeds to step S9. In step S9, the obtained analysis result regarding the walking of the worker 3 is output. That is, the result is displayed on an output unit 16 such as a display included in the information processing device 8. At this time, for example, the section between two points is displayed superimposed on a map of the work site 2, and the predetermined worker 3, the section, and the walking speed of the predetermined worker 3 in the section are displayed in correspondence with each other.

[0046] As described above, the mobile object analysis system 1 according to the embodiment of the present invention acquires first-person video data captured by a wearable camera 6 worn by a worker 3 and third-person video data capturing a bird's-eye view of the entire worksite 2. Based on the worker 3, objects, and equipment contained in the video data, a matching is performed to identify which worker 3 in the third-person video data is wearing a specific wearable camera 6. Data related to the walking of the worker 3 is also acquired from each piece of video data. That is, the first-person video data is analyzed using AI to determine the walking time of the worker 3. The third-person video data is also analyzed using AI to determine the walking distance (travel distance) of the worker 3. The walking speed of the worker 3 in a specific section of the worksite 2 can then be calculated based on the walking data acquired from each piece of video data.

[0047] The first-person video data can be compared with the third-person video data to accurately determine whether or not the worker 3 is walking. Furthermore, the third-person video data can be compared with the first-person video data to accurately determine the walking distance of the worker 3. The walking speed of the worker 3 in a specified section of the work site 2 is then determined based on this data, so highly accurate walking data can be acquired and output. Therefore, it is possible to accurately identify areas for improvement at the work site 2, such as so-called overburden, waste, and unevenness, which can be used to improve work efficiency at the work site 2.

[0048] Although the embodiments of the present invention have been described above, the present invention is not limited to the above examples and may be modified as appropriate within the scope of achieving the object of the present invention. For example, in the flowchart shown in FIG. 3, the processes of steps S2 to S6 may be configured to be performed in a different order. In particular, since the analysis of first-person video data and the analysis of third-person video data are different data, these processes may be performed simultaneously. Furthermore, the analysis of the data related to the gait of the worker 3 described above may be configured to be performed via a server. For example, the system may be configured to acquire learning data or a learning model for distinguishing between the walking and work of the worker 3 from an external server. Alternatively, the video data captured by the wearable camera 6 and the fixed camera 7 may be output to an external information processing device 8 via a server, and the gait of the worker 3 may be analyzed by the external information processing device 8. [Explanation of symbols]

[0049] 1. Mobile object analysis system 3. Worker (mobile) 6 Wearable camera (second camera) 7 Fixed camera (first camera) 8. Information processing equipment 9. Video acquisition unit 10 Object detection unit 11 Gait estimation unit (motion estimation unit) 12 Momentum estimation section (motion amount estimation section) 13 Travel amount analysis section 14 Feature determination section 15 Motion analysis section 16 Output section

Claims

1. A moving object analysis system that detects a moving object included in video data and analyzes data related to the movement of the moving object, a first camera that acquires first video data obtained by capturing an overhead view of the moving object; a second camera attached to the moving body and configured to acquire second video data viewed from the moving body; a movement amount estimating unit that estimates a movement amount of the moving object based on the first video data; a motion estimation unit that estimates a motion content of the moving object based on the second video data; a motion analysis unit that analyzes data relating to the motion of the moving object based on the amount of motion of the moving object when the motion content is estimated to be during movement. A mobile object analysis system characterized by:

2. The mobile object analysis system according to claim 1, a movement amount analysis unit that calculates a movement amount of the moving object quantitatively based on the movement amount of the moving object, the motion estimation unit is configured to estimate a travel time of the moving object based on a learning model obtained by machine learning about the moving object; The motion analysis unit is configured to calculate a moving speed of the moving object based on a moving amount and a moving time of the moving object. A mobile object analysis system characterized by:

3. 3. The mobile object analysis system according to claim 1, The present invention further includes a feature determination unit that determines, from the first video data, the moving object to which the second camera is attached, based on at least one parameter selected from the position, orientation, and movement of the moving object. A mobile object analysis system characterized by:

4. 3. The mobile object analysis system according to claim 1, The motion analysis unit is configured to analyze a point where the data regarding the movement of the moving object was acquired based on the first video data and the second video data. A mobile object analysis system characterized by:

Citation Information

Patent Citations

  • Monitoring system

    JP2021196909A

  • Action recognizing device, method, and program

    WO2023100330A1