Image processing apparatus

The image processing device corrects video data to align positions and sizes of moving objects, addressing misalignment issues and enhancing comparison accuracy in video overlays.

JP2025157934APending Publication Date: 2025-10-16TOYOTA JIDOSHA KK

Patent Information

Application Number
JP2024060296
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-03
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing image processing systems struggle to accurately overlay video data when the positions or sizes of moving objects differ, leading to misalignment and difficulty in comparing movements.

Method used

An image processing device that acquires multiple video streams, detects moving objects, calculates reference values for correction, and adjusts video data using correction values to align positions and sizes, enabling accurate superimposition of moving objects.

Benefits of technology

Enables accurate superimposition and comparison of moving objects in video data captured from different viewpoints or distances, facilitating efficient analysis of worker movements and identifying improvements in work sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025157934000001_ABST
    Figure 2025157934000001_ABST
Patent Text Reader

Abstract

To provide an image processing apparatus configured to generate video data of mobile bodies so that the mobile bodies different in size and location can be easily compared in superimposed video data.SOLUTION: An image processing apparatus for generating synthesized video data showing superimposed different video data includes: a mobile body detection unit which detects a mobile body; a reference value calculation unit which calculates a reference value to be used when superimposing a predetermined mobile body included in one piece of video data, out of pieces of superimposed video data, on another mobile body included in the other piece of video data; a correction value calculation unit which calculates a correction value for correcting at least one piece of video data based on the reference value; and an image processing unit which generates superimposed video data formed by correcting at least one piece of video data based on the calculated correction value and superimposing the predetermined mobile body on the other mobile body.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device capable of displaying objects included in different moving images in an overlapping manner. [Background technology]

[0002] Patent Document 1 discloses an image processing device for determining whether a person is exercising correctly. In the system of Patent Document 1, a video camera captures a user exercising on a treadmill and displays the captured video on a monitor. In the system of Patent Document 1, the video camera is installed in front of the treadmill and always captures the user from a fixed distance. The system of Patent Document 1 is also configured to extract video data of either the right or left half of the user's body from original video data, and generate superimposed video data by synchronizing the extracted video data with the original video data of the other half of the user's body. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-66696 Summary of the Invention [Problem to be solved by the invention]

[0004] According to Patent Document 1, a user can walk while comparing correct and incorrect walking. However, the system in Patent Document 1 overlays video data captured by a video camera with the same viewpoint, showing a user walking on a treadmill placed at a specific location. Therefore, when filming, it is necessary to maintain a constant distance between the treadmill and the video camera, as well as the user's position in the video. In other words, when the system in Patent Document 1 overlays video data in which the user's position or distance from the video camera differs, the resulting overlaid video data may have shifted in position or size, making it impossible to accurately compare the person's movements. Alternatively, it may be necessary to manually adjust the position and size of the person in the overlaid video data to accurately overlay and display the person in the different video data.

[0005] This invention has been made with an eye on the above-mentioned technical problems, and aims to provide an image processing device that can generate video data that allows easy comparison of moving objects even when video data in which the positions or sizes of the moving objects are different are superimposed. [Means for solving the problem]

[0006] In order to achieve the above-mentioned object, the present invention provides an image processing device that acquires multiple video data of moving objects captured by an imaging device and generates composite video data by overlaying the different video data together, and is characterized by comprising: a video data acquisition unit that acquires the multiple video data; a moving object detection unit that detects the moving objects included in the video data; a reference value calculation unit that calculates a reference value that serves as a reference when overlaying a specified moving object included in one of the video data to be overlaid with another moving object included in another of the video data to be overlaid; a correction value calculation unit that calculates a correction value that corrects at least one of the video data to be overlaid based on the reference value; and an image processing unit that corrects at least one of the video data to be overlaid based on the calculated correction value and generates the composite video data by overlaying the specified moving object with the other moving object.

[0007] In addition, the correction value calculation unit in this invention may be configured to use the reference value of the specified moving body as a reference and calculate the difference between the reference value of the other moving body and the reference value of the specified moving body as the correction value of the other moving body, and the image processing unit may be configured to correct the other video data based on the calculated correction value.

[0008] In addition, the correction value calculation unit in this invention may be configured to calculate the difference between the reference value of the specified moving body and the reference value of the other moving body as the correction value, and the image processing unit may be configured to correct the video data to be superimposed on each other based on the calculated correction value.

[0009] In addition, the reference value calculation unit in this invention may be configured to, when calculating the reference value, find the reference value for the specified moving object included in the specified image data among the image data constituting the specified video data, and for the other moving object included in the other image data among the image data constituting the other video data.

[0010] In addition, the reference value calculation unit in this invention may be configured, when calculating the reference value, to obtain the reference value for the specified moving object included in all image data constituting the specified video data, and for the other moving objects included in all image data constituting the other video data.

[0011] Furthermore, the moving body detection unit in this invention may be configured to detect the moving body using any one of detection processes using AI, including skeleton estimation, object detection, and object area estimation, when the moving body is a person, and the reference value calculation unit may be configured to calculate the reference value based on parameters obtained by any one of the detection processes. [Effects of the Invention]

[0012] An image processing device according to an embodiment of the present invention is configured to display two pieces of video data in an overlapping manner, with moving objects, such as workers, included in each piece of video data being displayed in an overlapping manner. In this case, the image processing device is configured to correct the position and size of each worker included in the video data to be overlapped, and display the workers so that their positions and sizes match. For example, when correcting one piece of video data based on the other piece of video data, the correction is performed using a predetermined part of the worker's body included in each piece of video data as a reference. Specifically, the length or distance of a predetermined part of the body of each worker in the two pieces of video data is used as a reference value. Then, a correction amount is calculated based on the reference value of the worker in the other piece of video data and the reference value of the worker included in the one piece of video data. The position and size of the worker included in the other piece of video data are corrected based on the calculated correction amount. The corrected worker in the other piece of video data and the worker in the one piece of video data are then displayed in an overlapping manner and output.

[0013] Therefore, even if the two video data are shot using cameras with different viewpoints or the distances from the cameras to the moving object are different, and therefore the positions or sizes of the workers in the two video data are different, the workers can be accurately superimposed and displayed. For example, by setting a standard based on the parts to be compared and matching only the standard and displaying the two video data in an overlapping manner, the movements of the workers can be accurately compared. In this way, the workers moving in the two video data can be accurately superimposed and compared, making it possible to check the proficiency of the workers, analyze the cause of any defects that occur at the work site, and easily identify and eliminate areas for improvement such as overstress, waste, and inconsistency at the work site.

[0014] Furthermore, if the correction value calculation unit is configured to use the reference value of a given moving object as a reference and calculate the difference between the reference value of another moving object as a correction value for that other moving object, and the image processing unit is configured to correct the other moving image data based on the calculated correction value, correction of the moving image data can be easily performed. In other words, since only one of the moving image data is corrected, the load caused by calculation processing in the device can be reduced.

[0015] Furthermore, if the correction value calculation unit is configured to calculate a correction value based on the difference between the reference value of a specific moving object and the reference value of another moving object, and the image processing unit is configured to correct the video data to be superimposed on each other based on the calculated correction value, the size of the moving object depicted in the video data can be adjusted to an appropriate size. For example, even if video data in which a worker is depicted small is used as the reference, the worker in the video data can be enlarged during correction. Therefore, not only can the positions and sizes of the workers in the two video data be matched, but also composite video data that is easy to view for a user using the output video data can be generated. In other words, composite video data can be generated that makes it easier to compare the detailed movements of the moving objects.

[0016] Furthermore, if the reference value calculation unit is configured to calculate reference values ​​for a predetermined moving object included in a predetermined image data among the image data constituting the predetermined video data and for another moving object included in another image data among the image data constituting the other video data, only one of the consecutive image frames constituting the video data is corrected. Therefore, it is possible to output composite video data while reducing the load on the device due to calculations, etc., when performing image processing such as detecting a moving object, calculating a reference value, or calculating a correction amount.

[0017] Alternatively, if the reference value calculation unit is configured to calculate reference values ​​for a given moving object included in all image data constituting the given video data and for other moving objects included in all image data constituting the other video data, the entire video data is corrected as needed, thereby making it possible to output composite video data with more accurate matching of the moving objects.

[0018] Furthermore, if the moving object is a person, the moving object detection unit may be configured to detect the moving object using one of the detection processes of skeleton estimation, object detection, and object region estimation, and the reference value calculation unit may be configured to calculate the reference value based on parameters obtained by one of the detection processes. In this case, workers can be superimposed on each other using an appropriate reference value depending on the detection process. For example, if skeleton estimation is performed, the video data can be accurately corrected by using the center of gravity at all joint positions of the worker and the distance from the center of gravity to the tips of the hands and feet as reference values. Therefore, even if the input videos are not images from the same viewpoint, composite video data can be output that makes it easy to compare the detailed movements of different workers. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is an explanatory diagram for explaining a state in which video data to be analyzed by an image processing device according to an embodiment of the present invention is being acquired; [Figure 2]1 is a block diagram illustrating a functional configuration of an image processing device according to an embodiment of the present invention. [Figure 3] FIG. 2 is an explanatory diagram illustrating a first image frame and a second image frame extracted from two pieces of video data, respectively. [Figure 4] 10A and 10B are explanatory diagrams for explaining a process for determining a reference position and a reference value of a worker by two-dimensional skeleton estimation. [Figure 5] 10A and 10B are explanatory diagrams for explaining a process for determining a reference position and a reference value of a worker by object detection. [Figure 6] 10A and 10B are explanatory diagrams for explaining a process for determining a reference position and a reference value of a worker by object region estimation. [Figure 7] 10A and 10B are explanatory diagrams for explaining a process for determining a reference position and a reference value of a worker by three-dimensional skeleton estimation. [Figure 8] 10A and 10B are explanatory diagrams for explaining a process executed when a correction value for an image frame is calculated based on a reference position and a reference value. [Figure 9A] 10A and 10B are explanatory diagrams for explaining a process of determining a reference position and reference values ​​of a worker by three-dimensional skeleton estimation and correcting an image frame. [Figure 9B] FIG. 9B is an explanatory diagram for explaining the process of determining the reference position and reference values ​​of a worker by three-dimensional skeleton estimation and correcting an image frame, and is an explanatory diagram for explaining the process subsequent to the process of FIG. 9A. [Figure 10] 4 is a flowchart illustrating an example of control executed by the moving object tracking system according to the embodiment of the present invention. [Figure 11] 10 is a flowchart illustrating another example of control executed by the moving object tracking system according to the embodiment of the present invention. [Figure 12] 10 is a flowchart illustrating yet another example of control executed by the moving object tracking system according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0020] Next, the present invention will be described based on the embodiments shown in the drawings. Note that the embodiments described below are merely examples of specific embodiments of the present invention, and are not intended to limit the present invention.

[0021] The image processing device in this embodiment of the present invention displays a superimposed display of multiple video data captured of a work site 2 including a worker (mobile object) 1. The image processing device is configured to detect the worker 1 included in the multiple video data captured by a camera 3 capturing an image of the work site 2, and to create composite video data in which the worker 1 is superimposed and displayed while adjusting and matching the position, size, etc. of each worker 1 in two different video data.

[0022] Fig. 1 shows a schematic diagram of an example of a work site 2 that is the subject of an embodiment of the present invention. At the work site 2 shown in Fig. 1, a worker 1 performs a predetermined task, such as assembling parts, on a work object 5 (workpiece) that has been transported by a transport facility 4. Although not shown in the figure, the work site 2 is configured so that as the work object 5 is transported, the position where the worker 1 works also moves in the same direction as the work object 5 is transported.

[0023] Camera 3 is an imaging device configured to capture an image of a predetermined area in work site 2, and captures the image of the predetermined area from diagonally above. As shown in Fig. 1, camera 3 is fixed to, for example, the ceiling of work site 2, and is positioned so as to capture an overhead image of the entire work site 2, which is the predetermined area. For example, as shown in Fig. 1, video data captured by camera 3 shows an image of a work object 5 being transported from the back to the front of the screen, with worker 1 and work object 5 appearing gradually larger over time.

[0024] This camera 3 may be configured similarly to a conventionally known camera, such as an RGB camera, and acquires captured images and videos as data. The camera 3 is configured to be able to output image data and video data to an information processing device 6 that processes the image data and video data. Note that the image data or video data captured by the camera 3 may be configured to be output to the information processing device 6 via an external server such as a management server.

[0025] The information processing device 6 mainly comprises a processor, a communication unit, a storage unit, etc. The information processing device 6 is configured to perform calculations according to a predetermined program using data acquired from the outside and pre-stored data, and to output the results of the calculations as control command signals. For example, the information processing device 6 executes functions that meet predetermined purposes by having the processor load a program stored on a recording medium into a working area of ​​the storage unit and execute the program, and perform various controls through the execution of the program.

[0026] The processor is, for example, a CPU or a DSP. This processor is configured to control the information processing device 6 and perform various information processing operations. The main memory unit includes, for example, a RAM and a ROM. As described above, the main memory unit has a work area for the processor to execute programs. The auxiliary memory unit includes, for example, an EPROM or a hard disk drive. This auxiliary memory unit may also include a portable recording medium, i.e., a removable medium. The auxiliary memory unit also freely stores various programs, various data, and various tables in the recording medium by reading or writing them. The auxiliary memory unit may also store an operating system. The communication unit is a wireless communication circuit connected to an external communication device via wireless communication so as to be able to communicate data. This wireless communication circuit communicates using cellular communication (mobile communication) such as 5G or 4G (LTE).

[0027] Next, the functional configuration of the information processing device 6 in this embodiment of the present invention will be described. As shown in Fig. 2, the information processing device 6 in this embodiment of the present invention includes a video acquisition unit 7, a worker detection unit 8, a worker extraction unit 9, a reference calculation unit 10, a correction value calculation unit 11, an image processing unit 12, and an output unit 13.

[0028] The video acquisition unit 7 acquires image data or video data captured by the camera 3. As shown in FIG. 1 , the video data acquired by the video acquisition unit 7 is video data of the work site 2 photographed from diagonally above by the camera 3. In other words, the video data is a video as a collection of image frames photographed from diagonally above of the work site 2, worker 1, work object 5, etc. Furthermore, the video data may be configured so that the user inputs any video data to be compared, or may be configured so that the video data is automatically acquired from the camera 3.

[0029] The worker detection unit 8 detects the worker 1 from the acquired video data based on learning data such as the feature amounts of the worker 1 that have been learned in advance. The worker detection unit 8 is configured to detect objects from image frames that make up the video data using a deep learning method such as conventionally known YOLO (YOLOv8) or SSD. That is, the worker detection unit 8 detects the worker 1 using methods such as two-dimensional skeleton estimation, object detection, object region estimation, and three-dimensional skeleton estimation.

[0030] The worker extraction unit 9 extracts the worker 1 detected by the worker detection unit 8 from the video data. The worker extraction unit 9 extracts the worker 1 by deleting or masking the background, such as the work site 2 and the work object 5, other than the worker 1 shown in the video data or image data.

[0031] The reference calculation unit 10 calculates a reference value for comparing the extracted workers 1 as described above. First, the reference calculation unit 10 sets a reference portion (parameter) from multiple video data according to the portion to be compared. Then, the reference value is calculated by calculating length, distance, etc. based on the set reference portion. For example, when the workers 1 shown in two video data are displayed superimposed and the movements of the workers 1 are compared, a reference value for superimposing the workers 1 is calculated. An example of the processing performed to calculate the reference value will be described with reference to FIG. 3.

[0032] FIG. 3 shows an example of a case where a first worker 1A included in a first image frame 14A extracted from one video data and a second worker 1B included in a second image frame 14B extracted from the other video data are displayed in an overlapping manner. As shown in FIG. 3, in the first image frame 14A, the first worker 1A is displayed slightly to the lower left of the center (center position) of the screen, and in the second image frame 14B, the second worker 1B is displayed slightly to the upper right of the center (center position) of the screen. When calculating reference values ​​for displaying the first worker 1A and the second worker 1B in an overlapping manner from such two video data, the reference calculation unit 10 determines a reference position and reference value. For example, when matching the position and size of the first worker 1A with those of the second worker 1B, the reference calculation unit 10 calculates a part of the first worker 1A and the length of that part and the same part and length of the second worker 1B. Note that since the parameters obtained differ depending on the detection method used by worker 1, such as the two-dimensional skeleton estimation, object detection, object region estimation, and three-dimensional skeleton estimation described above, the reference position and reference values ​​change depending on the detection method. The following describes the processing executed to determine the reference position (part) and reference value for each detection method. Note that first worker 1A corresponds to a predetermined moving body in this embodiment of the present invention, and second worker 1B corresponds to another moving body in this embodiment of the present invention.

[0033] First, with reference to FIG. 4, a description will be given of a reference position and a reference value when the worker 1 is detected from the image frame 14 by two-dimensional skeletal estimation. Two-dimensional skeletal estimation recognizes body parts of the worker 1 from the image frame and connects them to detect the skeletal structure in two dimensions, thereby estimating the posture of the worker 1. When the worker 1 is detected by two-dimensional skeletal estimation, a reference value is calculated based on the coordinates of each joint position of the detected worker 1 on the image. In two-dimensional skeletal estimation, the skeletons of the first worker 1A and the second worker 1B are estimated. Since each joint position is estimated based on the estimated skeletons, a reference position RP for correcting the image frame is set based on the joint positions. For example, the reference position RP in two-dimensional skeletal estimation may be the center of gravity of one of the joint positions or all of the joint positions. Specifically, as shown in FIG. 4(a), the reference position RP may be the waist of the worker 1 or the toes of either the right or left foot.

[0034] Then, a reference value RV is set based on the part based on the reference position RP set in this way. The part for obtaining the reference value RV is determined based on the length of the bones of the worker 1 obtained by skeletal structure estimation. For example, the length of any bone, the total length of all bones in the work, or the combined length of the neck bone, spine, and leg bones, i.e., the length equivalent to the height of the worker 1, is set. Specifically, as shown in FIG. 4(b), the length of the spine of the worker 1, etc., is set as the reference value RV.

[0035] Next, with reference to FIG. 5, a reference position and a reference value when a worker 1 is detected from an image frame 14 by object detection will be described. In object detection, an object (moving object) detected from an image frame is detected as the worker 1 based on pre-learned feature amounts of the worker 1, etc. When the worker 1 is detected by object detection, a reference value is calculated based on the area of ​​the detected worker 1. In object detection, the worker 1 is detected by an area (person area) enclosed by a square frame, and a reference position RP for correcting the image frame is set from the square area (bounding box) of the detected worker 1. For example, the reference position RP in object detection may be one of the four corners of the square area, or the center of gravity of the square area, as shown in FIG. 5(a).

[0036] Then, a reference value RV is set based on the reference position RP thus set. The part for obtaining the reference value RV is determined based on the dimensions of a square area detected by object detection. For example, as shown in FIG. 5(b), the width of the square area, the length of the diagonal of the square area, or the like may be set as the reference value RV. Furthermore, without being limited to these, the height of the square area or the area of ​​the square area may also be used. Note that when the width or height of the square area is used, the direction in which the image is enlarged or reduced may be limited to the width direction or the height direction.

[0037] Using FIG. 6, a reference position and reference value when worker 1 is detected from image frame 14 by object region estimation will be described. Object region estimation distinguishes worker 1 in units of pixels constituting the image frame and detects the pixel region of worker 1. When worker 1 is detected by object region estimation, a reference value is calculated based on the pixel region of the detected worker 1. In object region estimation, the pixel region of worker 1 is identified in the image frame, and a reference position RP for correcting the image frame is set from the region of worker 1. For example, the reference position RP in object region estimation may be one of the endpoints of the pixel region or the center of gravity of the pixel region, as shown in FIG. 6(a).

[0038] Then, a reference value RV is set based on the reference position RP thus set. The part for calculating the reference value RV is determined based on the dimensions of the pixel region detected by object region estimation. For example, as shown in FIG. 6(b), the reference value RV may be the area of ​​the pixel region, the maximum value of the distance from the center of gravity of the pixel region to an end point, or the like. Furthermore, without being limited to these, the reference value RV may also be the maximum value of the distance between end points of the pixel region. Note that after determining a rectangular region surrounding the pixel region, the reference value RV may be calculated by processing similar to the processing performed in object detection.

[0039] Using FIG. 7, a description will be given of the reference position and reference value when the worker 1 is detected from the image frame 14 by three-dimensional skeleton estimation. Three-dimensional skeleton estimation recognizes the skeleton, shape, and posture of the worker 1 from the image frame, and connects them to detect the skeletal structure in three dimensions and estimate the worker 1 as three-dimensional coordinates. When the worker 1 is detected by three-dimensional skeleton estimation, a reference value is calculated based on the coordinates of each joint position of the detected worker 1 on the image, similar to when the worker 1 is detected by two-dimensional skeleton estimation. In three-dimensional skeleton estimation, each joint position is estimated based on the estimated skeleton and posture of the worker 1, and a reference position RP for correcting the image frame is set based on the joint position. For example, the reference position RP in three-dimensional skeleton estimation may be the center of gravity of one of the joint positions or all of the joint positions, as in the case of two-dimensional skeleton estimation. Specifically, the reference position RP may be the waist of the worker 1 or the toes of either the right or left foot, as shown in FIG. 7.

[0040] Then, a reference value RV is set based on the reference position RP set in this way. The parts for obtaining the reference value RV are determined based on the lengths of the bones of the worker 1 obtained by skeletal structure estimation. For example, the length of any one bone, the total length of all bones in the work, or the combined length of the neck bone, spine, and leg bones, i.e., the length equivalent to the height of the worker 1, is set. Specifically, the length of the spine of the worker 1, etc. is set.

[0041] The results of the three-dimensional skeleton estimation may be projected onto a two-dimensional plane, and processing similar to that of the two-dimensional skeleton estimation may be performed. In this case, the length of one of the bones of the worker 1 obtained by the three-dimensional skeleton estimation may be used. The three-dimensional skeleton estimation can estimate the posture of the worker 1, such as whether the back of the worker 1 is bent, relatively accurately, and therefore the reference value RV can be calculated more accurately.

[0042] The reference calculation unit 10 calculates the reference position RP and the reference value RV for each of the first worker 1A in the first image frame 14A and the second worker 1B in the second image frame 14B using any of the methods described above.

[0043] The correction value calculation unit 11 calculates the amount of correction for one of the first image frame 14A and the second image frame 14B based on the reference position RP and reference value RV obtained from the reference calculation unit 10. In other words, the correction value calculation unit 11 calculates the amount of correction for the position and size of the worker 1 in one image frame so that the reference position RP and reference value RV of the worker 1 in one image frame match the reference position RP and reference value RV of the worker 1 in the other image frame, either the first image frame 14A or the second image frame 14B.

[0044] FIG. 8(a) illustrates a case where a correction amount for the position of the worker 1 in the image frame 14 is calculated based on the calculated reference position RP. In the example shown in FIG. 8(a), the waist position of the second worker 1B in the second image frame 14B estimated by two-dimensional skeleton estimation is set as the reference position RP, and the position of the first worker 1A in the first image frame 14A is corrected based on the reference position RP. The correction value calculation unit 11 acquires two-dimensional coordinates of the reference position RP of the first worker 1A from the reference calculation unit 10 and calculates the difference between the two-dimensional coordinates and the two-dimensional coordinates of the waist position of the first worker 1A in the first image frame 14A. In the example shown in FIG. 8(a), differences occur on both the vertical and horizontal axes in the first image frame 14A. The correction value calculation unit 11 calculates the amount of parallel movement of the first worker 1A along the vertical and horizontal axes based on the calculated difference, and acquires the amount of parallel movement as a position correction value, which is a correction value for the position.

[0045] Next, a case where a correction amount for the size of the worker 1 in the image frame is calculated based on the calculated reference value RV will be described with reference to FIG. 8(b). In the example shown in FIG. 8(b), the diagonal length of the square area of ​​the first worker 1A in the first image frame 14A detected by object detection is set as the reference value RV, and the size of the second worker 1B in the second image frame 14B is corrected based on the reference value RV. The correction value calculation unit 11 obtains the diagonal length of the square area, which is the reference value RV of the first worker 1A, from the reference calculation unit 10 and calculates the ratio between the diagonal length and the diagonal length of the square area of ​​the second worker 1B in the second image frame 14B. The correction value calculation unit 11 obtains the calculated ratio as a size correction value, which is a correction value for the size of the second worker 1B.

[0046] When the reference position RP and reference value RV are input three-dimensionally by three-dimensional skeleton estimation or the like, the correction value calculation unit 11 is configured to calculate the correction amount for the coordinate systems of the first image frame 14A and the second image frame 14B of the camera 3 using the world coordinate system. That is, when the reference position RP and the reference value RV are input three-dimensionally, the coordinate systems of the respective cameras 3 are virtually positioned in the world coordinate system. At this time, the correction amounts for the reference position RP and reference value RV of the first worker 1A and the second worker 1B, i.e., the position correction value and the size correction value, are calculated so that the positions and sizes of the first worker 1A and the second worker 1B coincide with each other.

[0047] The correction value calculation unit 11 calculates a position correction value and a size correction value for each of the first worker 1A in the first image frame 14A and the second worker 1B in the second image frame 14B using the method described above. In other words, the correction value calculation unit 11 is configured to calculate a correction amount that aligns the position and size of one of the first worker 1A and the second worker 1B with those of the other worker 1.

[0048] Image processing unit 12 corrects at least one of first image frame 14A and second image frame 14B based on the position correction value and size correction value of first worker 1A or second worker 1B calculated by correction value calculation unit 11. That is, image processing unit 12 corrects the position and size of either image frame by translating, enlarging, or reducing first worker 1A or second worker 1B based on the position correction value and size correction value, so that first worker 1A and second worker 1B overlap with each other in the same position and size.

[0049] For example, when correcting to align first worker 1A with second worker 1B as a reference, image processor 12 translates first image frame 14A vertically and horizontally based on the position correction value, and corrects the position of first image frame 14A so that the reference position RP of second worker 1B coincides with the reference position RP of first worker 1A. Image processor 12 also enlarges or reduces first image frame 14A based on the size correction value, and processes the image so that the lengths or distances of reference portions (parameters) of second worker 1B and first worker 1A coincide with each other.

[0050] When a position correction value and a size correction value are input simultaneously to the image processing unit 12, the image frame is corrected using an image transformation matrix (affine transformation) that combines translation and linear transformation to simultaneously satisfy both corrections. That is, the coordinates of the image frame are transformed using a matrix to enlarge, reduce, and translate. For example, the image frame is transformed or corrected based on a simultaneous coordinate system expressed by a 3x3 matrix that adds rotation to the enlargement, reduction, and translation of the image frame. Note that the image processing unit 12 may also be configured to perform skew (shear), a process that transforms a rectangle into a parallelogram, as needed.

[0051] When the reference position RP and reference value RV are input three-dimensionally by three-dimensional skeleton estimation or the like, the correction value calculation unit 11 is configured to calculate the correction amount for the camera coordinate system using the world coordinate system. That is, the correction value calculation unit 11 calculates the correction amount by virtually arranging the camera coordinate systems in the first image frame 14A and the second image frame 14B so that the coordinates of the reference position RP match a world coordinate system that defines a coordinate system based on the X-axis, Y-axis, and Z-axis for the entire space. Then, based on the correction amount, the two image frames 14 are corrected to obtain matching positions.

[0052] 9A, based on the position correction values ​​and size correction values ​​obtained as described above, first image frame 14A and second image frame 14B are virtually placed in the world coordinate system so that the coordinates of the reference position RP and reference value RV of first worker 1A and second worker 1B match. When placing them, the respective camera coordinate systems are virtually placed in the world coordinate system as a common coordinate system using external parameters such as the mounting position and orientation of camera 3. Note that if the video data is of the same area photographed from the same position by the same camera 3, the camera coordinate system may also be used.

[0053] In addition, in the world coordinate system, a plane parallel to the image plane arranged so that the reference positions RP and reference values ​​RV of the first worker 1A and the second worker 1B coincide is defined as a predetermined plane P. Then, as shown in FIG. 9B , the first image frame 14A and the second image frame 14B are projected onto the predetermined plane P using internal parameters such as the focal length of the lens, the position of the optical axis, and the lens distortion of the camera 3. A virtual camera having an image plane overlapping the predetermined plane P is set, and images projected onto the imaging area of ​​the virtual camera are acquired. The image frames are then corrected based on the acquired images. At this time, the internal parameters of the virtual camera and the positions of the virtual camera in directions other than the optical axis direction, i.e., the X-axis and Y-axis directions, may be set arbitrarily.

[0054] Note that, in the case of video data captured from the same position by the same camera 3 of the same area, the predetermined plane P may be a plane that coincides with the image plane of any of the cameras 3, or a plane that is parallel to the image planes of any of the cameras 3. The internal parameters of the virtual camera may be adjusted to match the internal parameters of the camera 3. Furthermore, the internal parameters of the virtual camera may be configured to be set so as to include the entire image area after the first image frame 14A and the second image frame 14B are projected onto the predetermined plane P.

[0055] Image processing unit 12 performs the image conversion as described above based on the position correction value and size correction value calculated by correction value calculation unit 11, thereby aligning the position and size of one of first worker 1A and second worker 1B with the other worker 1. Then, first image frame 14A and second image frame 14B are configured to be superimposed on each other in this state.

[0056] The output unit 13 displays the video data superimposed by the image processing unit 12. The output unit 13 may be configured, for example, by a PC monitor or a display of a mobile terminal.

[0057] Next, an example of control executed in the image processing device configured as described above will be described with reference to Fig. 10. The flowchart shown in Fig. 10 illustrates processing for displaying composite video data in which the position and size of the worker 1 displayed in two pieces of input video data are corrected and superimposed. The flowchart shown in Fig. 10 illustrates control for correcting an image frame 14 at a predetermined timing, such as the start of video data, and superimposing and displaying subsequent video data in accordance with the corrected image frame 14.

[0058] 10, first, in step S1, video data to be displayed in an overlapping manner is input. In step S1, for example, the user manually inputs the video data to be displayed in an overlapping manner into the information processing device 6. Alternatively, the video data captured by the camera 3 may be automatically transferred to the information processing device 6 and stored in the information processing device 6. In that case, the input may be performed by the user selecting the video data to be compared or displayed in an overlapping manner from the video data stored in the information processing device 6.

[0059] When two pieces of video data to be superimposed have been input, the process proceeds to step S2. In step S2, a reference position RP and a reference value RV are determined to align the position and size of an object to be superimposed and displayed in each image frame 14 constituting the two pieces of input video data. In step S2, the object captured in each of the two image frames 14, i.e., the worker 1, is identified, and the reference position and reference value, i.e., length and distance, for the identified worker 1 are determined. For example, in step S2, if the waist of one worker 1 in one piece of video data is set as the reference position RP and the distance from the waist to the toes of the right foot is set as the reference value RV, a similar reference position RP and reference value RV are determined for the other worker 1 in the other piece of video data.

[0060] After the reference position RP and reference value RV of each worker 1 included in each video data are determined, the process proceeds to step S3. In step S3, the correction amount for the image frame 14 is calculated based on the reference value RV determined in step S2. That is, in step S3, the difference in the reference position RP and the difference in the reference value RV between the workers 1 in each image frame 14 are determined.

[0061] For example, if the reference position RP of one worker 1 in one image frame 14 is misaligned with the reference position RP of the other worker 1 in the other image frame 14, the coordinate movement amount required to align the reference positions RP of each worker 1 is calculated. Furthermore, if the reference value RV of one worker 1 in one image frame 14 is different from the reference value RV of the other worker 1 in the other image frame 14, the enlargement or reduction rate of the image frame 14 required to align the reference values ​​RV of each worker 1 is calculated. Alternatively, if correction of both the reference position RP and the reference value RV is required, an image transformation matrix capable of correcting both is calculated. In step S3, the correction amount of the image frame 14 is calculated in this manner. Note that if the reference position RP and the reference value RV are input three-dimensionally, the camera coordinate system is positioned in the world coordinate system so that the coordinates of the reference position RP and the reference value RV coincide, and an image is projected onto a predetermined plane P based on the internal parameters of the camera 3. The correction amount is then calculated by acquiring an image projected onto the imaging area of ​​a virtual camera having an image plane overlapping the predetermined plane P.

[0062] After the correction amounts for the position and size of the worker 1 in the image frame 14 have been calculated in this way, the process proceeds to step S4. When the process proceeds to step S4, one of the two image frames 14 is corrected based on the correction amount calculated in step S3. For example, the reference position RP of the other worker 1 in the other video data is corrected based on the calculated position correction amount. Also, the reference value RV of the other worker 1 in the other video data is corrected based on the calculated size correction amount. In step S4, based on the corrected image frame 14, the correction has been applied to all image frames 14 that make up the other video data, and composite video data obtained by overlaying the two video data is output.

[0063] 10 is configured to automatically determine and correct the reference position RP and reference value RV of each worker 1 based on two pieces of input video data. However, the image processing device according to the embodiment of the present invention is not limited to such a configuration, and may be configured so that the reference position RP or reference value RV is manually set by a user or the like, and the video data is corrected in accordance with the setting.

[0064] Control in such a case will be described using the flowchart shown in Fig. 11. In Fig. 11, the same steps as those shown in Fig. 10 are given the same reference numerals, and their description will be omitted or simplified. In the control example shown in Fig. 11, first, the same process as step S1 described above is performed. That is, two pieces of video data to be superimposed are input.

[0065] After the two video data are input, the process proceeds to step S11. In step S11, the user inputs settings for correcting the video data. For example, in step S11, the user sets whether to correct both the position and size of the worker 1, or whether to correct either the position or the size of the worker 1. In addition, when correcting the video data, the user sets which method to use for correcting the video data, such as the above-mentioned two-dimensional skeleton estimation, object detection, object region estimation, and three-dimensional skeleton estimation. Note that, when the process of step S11 is performed, recommended regulation settings may be provided in advance as initial settings. It is preferable that the regulation settings be set so that the user can change them as desired.

[0066] After the user has set the video data correction, the process proceeds to step S2. That is, a reference position RP and a reference value RV are calculated to align the positions and sizes of the objects to be displayed in superimposed fashion in each of the image frames 14 constituting the two input video data. At this time, the reference position RP and the reference value RV are calculated based on the correction settings made by the user as described above.

[0067] After the reference position RP and reference value RV of each worker 1 included in each image frame 14 are determined based on the user's settings, the process proceeds to step S3. In step S3, as described above, the correction amount for the image frame 14 is calculated based on the reference position RP and reference value RV determined in step S2.

[0068] After the correction amount for the position and size of the worker 1 in the image frame 14 has been calculated in this way, the process proceeds to step S4. In step S4, one of the two image frames 14 is corrected based on the calculated correction amount, as described above. Then, based on the corrected image frame 14, the correction is applied to all image frames 14 that make up the other moving image data, and composite moving image data in which the two moving image data are superimposed is output.

[0069] 10 and 11 are configured to correct a specific image frame 14 in moving image data and then display two pieces of moving image data in an overlapping manner based on the correction. However, the image processing device according to the embodiment of the present invention is not limited to such a configuration and may be configured to perform correction on all image frames 14 that make up the moving image data.

[0070] Control in such a case will be described using the flowchart shown in Fig. 12. In Fig. 12, the same steps as those shown in Fig. 10 or 11 are given the same reference numerals, and their description will be omitted or simplified. In the control example shown in Fig. 12, first, the same process as step S1 described above is performed. That is, two pieces of video data to be superimposed are input.

[0071] After the two video data are input, the process proceeds to step S21. In step S21, a predetermined image frame 14 is extracted from the two video data. For example, in step S21, the first image frame within the range of images in which the two video data are to be superimposed is extracted. At this time, if there is an image frame 14 that has already been corrected, the next image frame 14 following the corrected image frame 14 is extracted.

[0072] After the image frames 14 are extracted from each of the two video data in this manner, the process proceeds to step S2. In step S2, a reference position RP and a reference value RV are calculated to align the positions and sizes of the objects to be displayed in superimposed fashion in each of the extracted image frames 14. At this time, if the user has made correction settings as described above, the reference position RP and the reference value RV are calculated based on those settings.

[0073] After the reference position RP and reference value RV of each worker 1 included in each image frame 14 are determined, the process proceeds to step S3. In step S3, as described above, the correction amount for the image frame 14 is calculated based on the reference position RP and reference value RV determined in step S2.

[0074] After the correction amount for the image frame 14 has been calculated in this manner, the process proceeds to step S22. In step S22, it is determined whether the image frame 14 for which the correction amount has been calculated is the last image frame 14 in the video data. That is, in step S22, it is determined whether calculation of the correction amount for one of the two video data, that is, the video data to be corrected, has been completed. If the determination in step S22 is NO because calculation of the correction amount for the video data has not been completed, the process returns to step S21. That is, because there is an image frame 14 constituting one of the video data for which the correction amount has not yet been calculated, the process returns to step S21, and the correction amount for the next image frame 14 consecutive to the corrected image frame 14 is calculated.

[0075] Conversely, if the correction amounts for all image frames 14 constituting the video data have been calculated and the determination in step S22 is YES, the process proceeds to step S4. In step S4, as described above, one of the two video data is corrected based on the calculated correction amounts. Then, the corrected one video data and the other video data are superimposed and displayed and output. When all image frames constituting the video data are corrected in this way, video data can be created in which the two video data are superimposed with the workers 1 more accurately aligned.

[0076] As described above, the image processing device according to the embodiment of the present invention can display two pieces of video data in an overlapping manner. The image processing device is also configured to display each worker 1 included in the two pieces of video data in an overlapping manner. In this case, the image processing device is configured to correct the position and size of each worker 1 included in each piece of video data, so that the positions and sizes of the workers 1 are aligned and displayed. For example, using a predetermined part of the body of each worker 1 included in the two pieces of video data as a reference, the position and size of the predetermined part of the body of the worker 1 included in the other piece of video data is corrected. In this case, the predetermined part of each worker 1 in the two pieces of video data is set as a reference position RP, and the length or distance from the reference position RP to the other part is set as a reference value RV. Then, a correction amount is calculated and corrected so that the reference position RP and reference value RV of the worker 1 included in the other piece of video data are aligned with the reference position RP and reference value RV of the worker 1 included in one piece of video data. The corrected worker 1 from the other piece of video data and the worker 1 from the one piece of video data are then displayed and output in an overlapping manner.

[0077] Therefore, even if the two video data are shot by cameras 3 from different viewpoints or the sizes of the workers 1 captured are different due to differences in distance from the cameras 3, it is possible to accurately superimpose and display each worker 1. In other words, by setting a standard based on the parts to be compared and matching only that standard to display the two video data in an overlapping manner, it is possible to accurately compare the movements of each worker 1. Therefore, because it is possible to accurately superimpose and compare each worker 1 moving in the two video data, it is possible to confirm the proficiency of the worker 1, analyze the cause of any problems that occur at the work site 2, and easily find and eliminate areas for improvement such as overstress, waste, and unevenness at the work site 2.

[0078] Although the embodiments of the present invention have been described above, the present invention is not limited to the above examples and may be modified as appropriate within the scope of achieving the object of the present invention. For example, in the flowchart shown in FIG. 12, the process for superimposing and displaying each worker 1 included in two video data is configured to be performed entirely automatically. This configuration is not limited to this, and the system may be configured to execute a process in which the correction method, conditions, etc. are manually set by the user, as shown in step S11 of FIG. 11. This configuration increases the degree of freedom in correcting the video data, making it possible to obtain video data that more accurately reflects the user's desired superimposition.

[0079] In the above-described embodiment, the first worker 1A and the second worker 1B are described as different workers 1, but the first worker 1A included in one video data and the second worker 1B included in the other video data may be the same worker 1. In such a configuration, it is possible to find areas for improvement for a specific worker 1 and confirm changes in work.

[0080] Furthermore, when correcting the worker 1 captured in the image frames 14 constituting the video data, both of the image frames 14 of the two video data may be corrected. For example, the difference between the reference position RP and the reference value RV of each worker 1 in the two image frames 14 is calculated. Then, based on the difference, a correction amount for each worker 1 in each image frame 14 is calculated, and the position and size of each worker 1 in each image frame 14 are corrected based on the correction amount. The video data may be corrected and superimposed so that the corrected workers 1 overlap each other, and output to a monitor or the like. With this configuration, if the worker 1 shown in the reference video data is small or large, it is possible to prevent or suppress the composite video data that is corrected based on such a small or large worker 1 and output as a result of correction being difficult to see.

[0081] Alternatively, one of the moving image data or an enlarged or reduced image frame 14 constituting one of the moving image data may be used as a reference, and the other moving image data may be corrected from that reference state. That is, a reference position RP and a reference value RV of a predetermined part of the worker 1 in the enlarged image frame 14 in one of the moving image data and the worker 1 in the other moving image data are obtained, and a correction amount is calculated based on the reference position RP and the reference value RV. The other moving image data may then be corrected based on the correction amount, and composite moving image data in which the two moving images are superimposed may be output.

[0082] Furthermore, while the above-described embodiment describes the case where two video data captured from the same viewpoint are superimposed, the image processing device according to the embodiment of the present invention can also be applied to cases where two video data captured from different viewpoints are superimposed and displayed. For example, when superimposing two video data of a worker 1 captured at different angles, the position, size, and angle of the worker 1 are detected using the above-described known detection method. That is, a predetermined part of the worker 1 in an image frame 14 constituting one of the two video data is used as a reference, and a reference position RP, reference value RV, and reference angle of the worker 1 are determined. Note that the reference angle may be, for example, the angle of the face or body orientation relative to a plane. Similarly, a reference position RP, reference value RV, and reference angle of a predetermined part of the worker 1 in an image frame 14 constituting the other of the two video data are determined. Then, the difference between the reference position RP, reference value RV, and reference angle is calculated, and the position, size, and angle of the worker 1 included in the image frame 14 of the other video data are corrected based on the difference. The other video data may then be corrected based on the correction of the image frame 14, and the one video data and the corrected video data may be superimposed and displayed. With this configuration, even if the multiple video data input to the image processing device are video data shot from different angles, the movements of the displayed workers can be accurately superimposed and output. [Explanation of symbols]

[0083] 1. Worker 2. Work site 3 Camera 4. Conveying equipment 5. Work Objects 6. Information processing equipment 7 Video Acquisition Unit 8. Worker detection unit 9 Worker extraction part 10 Reference calculation section 11 Correction value calculation unit 12 Image processing section 13 Output section 14 Image Frames P given plane RP reference position RV reference value

Claims

1. An image processing device that acquires a plurality of pieces of video data obtained by capturing images of a moving object using an imaging device, and generates composite video data by superimposing different pieces of video data together, a video data acquisition unit that acquires a plurality of pieces of video data; a moving object detection unit that detects the moving object included in the video data; a reference value calculation unit that calculates a reference value that serves as a reference when superimposing a predetermined moving object included in one of the moving image data to be superimposed on another moving image data to be superimposed on another of the moving image data to be superimposed on each other; a correction value calculation unit that calculates a correction value for correcting at least one of the moving image data to be superimposed on each other based on the reference value; an image processing unit that corrects at least one of the pieces of moving image data to be superimposed on each other based on the calculated correction value and generates the composite moving image data in which the predetermined moving object and the other moving object are superimposed on each other.

1. An image processing device comprising:

2. 2. The image processing device according to claim 1, the correction value calculation unit uses the reference value of the predetermined moving body as a reference and calculates a difference between the reference value of the other moving body and the predetermined moving body as the correction value of the other moving body; The image processing unit is configured to correct the other video data based on the calculated correction value.

1. An image processing device comprising:

3. 2. The image processing device according to claim 1, the correction value calculation unit calculates a difference between the reference value of the predetermined moving body and the reference value of the other moving body as the correction value; The image processing unit is configured to correct the moving image data to be superimposed on each other based on the calculated correction value.

1. An image processing device comprising:

4. 4. The image processing device according to claim 1, The reference value calculation unit is configured to calculate the reference value for the predetermined moving object included in predetermined image data among image data constituting the predetermined moving image data, and for the other moving object included in other image data among image data constituting the other moving image data.

1. An image processing device comprising:

5. 4. The image processing device according to claim 1, The reference value calculation unit is configured to calculate the reference value for the predetermined moving object included in all image data constituting the predetermined moving image data and the other moving objects included in all image data constituting the other moving image data when calculating the reference value.

1. An image processing device comprising:

Citation Information

Patent Citations

  • Image processing system and image processing method

    JP2013066696A

Cited By

  • Stable gel composition having high oil content, and preparation method therefor and application thereof

    US12582600B2