Image processing device, method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2024-06-17
- Publication Date
- 2026-04-30
AI Technical Summary
Conventional image stitching techniques often fail to generate a composite image effectively when the object is uneven, has changing shadows, or significantly deviates from the expected shape model, leading to fragmentation or distortion, making it difficult to obtain a normal composite image, especially in challenging environments like tunnel structures.
An image processing device and method that performs feature point matching, determines the quality of the composite image based on fragmentation and distortion criteria, and employs additional synthesis processes to correct distortions and arrange fragments based on shooting position information, ensuring a composite image with reduced fragmentation and acceptable distortion is generated.
The solution enables the generation of effective composite images that provide an overview of the object even when traditional stitching fails, ensuring sufficient image quality for understanding the object's outline, particularly in complex environments like tunnel structures.
Smart Images

Figure 2025028045000001 
Figure 2025028045000002
Abstract
Description
Image processing device, method and program
[0001] The present invention relates to an image processing device, method, and program, and more particularly to an image processing device, method, and program for joining a plurality of images together to generate a single composite image.
[0002] Stitching is a known technique for generating a single image that captures a wide range. In stitching, an object is divided into multiple images, and the resulting images are stitched together to generate a single composite image (see, for example, Patent Documents 1 to 4).
[0003] JP 2023-7662 A JP 2005-30961 A JP 2003-111073 A JP 2002-188998 A
[0004] One embodiment of the technique of the present disclosure provides an image processing device, method, and program that can generate an image that is effective for grasping the outline of an object even when a normal composite image cannot be generated.
[0005] (1) An image processing device comprising a processor, which acquires information on an image group and the shooting locations of the images that make up the image group, performs a first synthesis process on the image group to generate a first synthetic image, determines whether the quality of the first synthetic image satisfies a first criterion, and if the quality of the first synthetic image does not satisfy the first criterion, performs a second synthesis process on the image group or the first synthetic image based on the information on the shooting locations to generate a second synthetic image.
[0006] (2) The image processing device described in (1) in which the processor performs, as the first synthesis process, feature point matching between images that constitute the image group and generates a first synthetic image from the results of the feature point matching.
[0007] (3) The image processing device according to (1) or (2), wherein the processor determines the degree of fragmentation of the first composite image and determines whether the quality of the first composite image satisfies a first criterion.
[0008] (4) In the image processing device described in (3), the processor performs a second synthesis process in which at least some of the fragments of the first composite image are placed on fragments different from the fragments based on information about the shooting position and synthesized, thereby generating a second composite image that is less fragmented than the first composite image.
[0009] (5) An image processing device described in (4), in which the processor extracts fragments from the multiple fragments that have distortion beyond a predetermined range before performing the second synthesis process, and further performs processing to correct the distortion of the extracted fragments that have distortion.
[0010] (6) An image processing device described in any one of (1) to (5), in which the processor determines whether the quality of the second composite image satisfies a second criterion, and if the quality of the second composite image does not satisfy the second criterion, generates an image as the second composite image in which the images constituting the image group are arranged based on information about the shooting positions.
[0011] (7) An image processing device described in any one of (1) to (6), in which the processor determines whether the degree of distortion of at least the first composite image is within an acceptable range and determines whether the quality of the first composite image satisfies a first standard.
[0012] (8) The image processing device according to (7), wherein the processor performs the second synthesis process by arranging the images that make up the image group based on information about the shooting positions to generate a second synthesized image.
[0013] (9) An image processing device described in (7) or (8), in which the processor measures the maximum and minimum values of the width in the first direction of the first composite image, calculates the difference between the maximum and minimum values of the width in the first direction, determines whether the difference is less than or equal to a threshold, and determines whether the degree of distortion of the first composite image is within an acceptable range.
[0014] (10) An image processing device described in any one of (7) to (9), in which the processor acquires information on the shooting conditions of the images that constitute the image group, estimates the width in the second direction of the first composite image based on the information on the shooting conditions, measures the width in the second direction of the first composite image, calculates the difference between the estimated value of the width in the second direction and the measured value, determines whether the difference is less than or equal to a threshold, and determines whether the degree of distortion of the first composite image is within an acceptable range.
[0015] (11) An image processing device described in any one of (7) to (10), in which the processor calculates the degree of deviation of synthesis parameters determined from the results of feature point matching between adjacent images, determines whether there are any images with a degree of deviation greater than a threshold, and determines whether the degree of distortion of the first synthesis image is within an acceptable range.
[0016] (12) The image processing device according to any one of (1) to (11), wherein the processor performs feature point matching between adjacent images based on information about the shooting positions.
[0017] (13) An image processing device described in any one of (1) to (12), wherein the image group is composed of images taken while changing positions using a photographing device equipped with multiple cameras, and the information on the photographing positions is composed of information on the placement positions of the cameras in the photographing device and information on the positions where the photographs were taken by the photographing device.
[0018] (14) An image processing device described in any one of (1) to (13), in which a processor acquires an image group for each section of an object divided into multiple sections in the longitudinal direction, and information on the shooting positions of the images that make up the image group, generates a first composite image or a second composite image for each section, and generates a first output image in which the first composite image or the second composite image generated for each section is arranged based on the arrangement of the sections.
[0019] (15) An image processing device described in any one of (1) to (14), in which the processor analyzes the images constituting the image group, the first composite image, or the second composite image to detect damage on the surface of the object, and generates a second output image in which the damage detection results are superimposed on the first composite image or the second composite image.
[0020] (16) An image processing method that acquires information on an image group and the shooting locations of the images that make up the image group, performs a first synthesis process on the image group to generate a first synthetic image, determines whether the quality of the first synthetic image satisfies a first criterion, and if the quality of the first synthetic image does not satisfy the first criterion, performs a second synthesis process on the image group or the first synthetic image based on the information on the shooting locations to generate a second synthetic image.
[0021] (17) An image processing program that causes a computer to perform the following functions: a function of acquiring information on the image group and the shooting locations of the images that make up the image group; a function of performing a first synthesis process on the image group to generate a first synthetic image; a function of determining whether the quality of the first synthetic image satisfies a first criterion; and a function of performing a second synthesis process on the image group or the first synthetic image based on the information on the shooting locations to generate a second synthetic image if the quality of the first synthetic image does not satisfy the first criterion.
[0022] FIG. 1 is a diagram showing a schematic configuration of a photography system; A perspective view showing the configuration of a multi-eye photography device; A front view showing the configuration of a multi-eye photography device; A side view showing the configuration of a multi-eye photography device; A block diagram showing the electrical configuration of a multi-eye photography device; A diagram showing an example of the hardware configuration of a control device; A functional block diagram of the photography control function of the control device; A block diagram of the main functions of the camera control unit; A diagram showing an example of a live view display; A functional block diagram of the image processing function of the control device; A conceptual diagram of composition processing by feature point matching; A diagram showing an example of a fragmented composite image; A diagram showing an example of a distorted composite image; A diagram showing an example of a distorted composite image; A conceptual diagram of composition processing in the second composition processing unit; A diagram showing an example of a composition result; A conceptual diagram of generation of a composite image by the third composition image processing unit; A diagram showing an example of a composite image generated by the third composition processing unit; A flowchart showing the procedure for processing to generate a composite image; A conceptual diagram of a method for identifying images captured in duplicate; A diagram showing an example of a distorted, fragmented primary image; A functional block diagram showing an example of generating a secondary composite image by correcting distortion; A diagram showing an example of generating a secondary composite image by performing distortion correction; A diagram showing an example of displaying the composition result when a composite image is generated by dividing it into multiple sections; A diagram showing an example of displaying an enlarged entire image; A diagram showing an example of displaying the damage detection result; A diagram showing an example of photographing an object with one camera
[0023] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0024] Here, as a stitching synthesis process, a case where synthesis is performed by so-called feature point matching (corresponding point search) will be described as an example.
[0025] In feature point matching, feature points are extracted from each image, similar feature points are matched, and synthesis parameters are generated based on their correspondence to synthesize the images. Therefore, synthesis may fail if there are gaps in the image, if there is insufficient overlap between adjacent images, or if the image is blurred or blurred. While these issues are caused by the image capture, synthesis may also fail due to the subject being photographed. For example, if the surface is highly uneven, and the shadowing or shape of the captured image changes depending on the shooting position, feature point matching may not be possible, resulting in synthesis failure. Synthesis may also fail if the shape of the photographed subject significantly deviates from the expected shape model. Synthesis failure can result in fragmented images or significantly distorted images. "Fragmented images" refers to the image being generated by splitting into multiple parts.
[0026] If the synthesis fails, it is necessary to take another photograph to obtain a normal synthesized image. However, depending on the object, it may be difficult to take another photograph. On the other hand, depending on the application, it may be sufficient to grasp the outline of the object.
[0027] The present embodiment aims to provide an image processing device that can generate an image that is effective for grasping the outline of an object even when a normal composite image cannot be generated.
[0028] [Photography System] Here, an example will be described in which the present invention is applied to a photography system used for inspecting tunnel structures.
[0029] Tunnel structures such as water conduits for hydroelectric power plants and subway tunnels are inspected periodically to ensure their safety. Recently, visual inspections have been replaced by image-based inspections. Image-based inspections involve photographing the walls of tunnel structures with a camera and then visually or by image processing to detect cracks and other damage in the images.
[0030] The area to be photographed is divided into multiple parts (so-called segmented photography). In addition, to generate a composite image (so-called panoramic composite image) from the photographed images, each image is photographed with a portion overlapping with adjacent images.
[0031] [Configuration of Imaging System] FIG. 1 is a diagram showing a schematic configuration of an imaging system.
[0032] As described above, the photography system 1 of this embodiment is configured as a system for photographing the inner wall surface of a tunnel structure TS. The tunnel structure TS, which is the object (subject) to be photographed, has an arc-shaped cross section (semicircular).
[0033] As shown in Figure 1, the photography system 1 of this embodiment includes a multi-eye photography device 10 that uses multiple cameras to photograph the inner wall surface of a tunnel structure TS, and a control device 100 that controls the multi-eye photography device 10 and processes images photographed by the multi-eye photography device 10.
[0034] The multi-eye photography device 10 is mounted on, for example, a cart Tr and photographs while moving within the tunnel structure TS (taking photographs while changing its position). The cart Tr may be equipped with an electric assist function (a function that assists human power with an electric motor) as needed.
[0035] [Multi-eye Photography Apparatus] FIG. 2 is a perspective view showing the configuration of the multi-eye photography apparatus. FIG. 3 is a front view showing the configuration of the multi-eye photography apparatus. FIG. 4 is a side view showing the configuration of the multi-eye photography apparatus. In FIGS. 2 to 4, the x-axis, y-axis, and z-axis are three mutually orthogonal axes. The plane including the x-axis and y-axis is the horizontal plane, and the z-axis is the vertical direction. The x-axis direction is the traveling direction of the dolly Tr, and the + direction of the x-axis (rightward in FIG. 4) is the traveling direction during photography. Therefore, the + direction of the x-axis (leftward in FIG. 4) is the forward direction (forward direction) of the dolly Tr and the multi-eye photography apparatus 10, and the - direction (leftward in FIG. 4) is the backward direction (rearward direction) of the dolly Tr and the multi-eye photography apparatus 10.
[0036] The multi-eye photography device 10 is configured using multiple cameras and multiple lighting devices. The number of cameras and lighting devices can be increased or decreased as needed depending on the subject being photographed. Here, we will explain an example in which the multi-eye photography device 10 is configured using five cameras C1 to C5 and five lighting devices L1 to L5.
[0037] The multi-eye photographing device 10 has a frame 11 for mounting a plurality of cameras C1 to C5 and a plurality of lighting devices L1 to L5.
[0038] The frame 11 is mainly composed of a flat base plate 12, a rectangular column 13 installed on the base plate 12, and a disk-shaped mounting base 14 attached to the column 13. The base plate 12 functions as an installation section on the dolly Tr. The column 13 and mounting base 14 are installed perpendicular to the base plate 12. An axis Ax that passes through the center of the mounting base 14 and is parallel to the x-axis is defined as the axis of the multi-eye photography device 10.
[0039] Cameras C1 to C5 and lighting devices L1 to L5 are attached to mounting base 14 via brackets B1 to B5. Hereinafter, cameras C1 to C5 will be distinguished from one another as necessary by referring to camera C1 as the "first camera C1," camera C2 as the "second camera C2," camera C3 as the "third camera C3," camera C4 as the "fourth camera C4," and camera C5 as the "fifth camera C5."
[0040] Furthermore, the lighting devices L1 to L5 are distinguished from one another by referring to the lighting device L1 as the "first lighting device L1," the lighting device L2 as the "second lighting device L2," the lighting device L3 as the "third lighting device L3," the lighting device L4 as the "fourth lighting device L4," and the lighting device L5 as the "fifth lighting device L5."
[0041] Furthermore, bracket B1 will be referred to as the "first bracket B1," bracket B2 as the "second bracket B2," bracket B3 as the "third bracket B3," bracket B4 as the "fourth bracket B4," and bracket B5 as the "fifth bracket B5" to distinguish between brackets B1 to B5.
[0042] The first camera C1 and the first lighting device L1 are attached to the mounting base 14 via a first bracket B1. The second camera C2 and the second lighting device L2 are attached to the mounting base 14 via a second bracket B2. The third camera C3 and the third lighting device L3 are attached to the mounting base 14 via a third bracket B3. The fourth camera C4 and the fourth lighting device L4 are attached to the mounting base 14 via a fourth bracket B4. The fifth camera C5 and the fifth lighting device L5 are attached to the mounting base 14 via a fifth bracket B5.
[0043] Each of the brackets B1 to B5 is arranged on the same circumference centered on the axis Ax with respect to the mounting base 14. Furthermore, each of the brackets B1 to B5 is attached to the mounting base 14 so as to be movable in the circumferential direction within a predetermined angular range (for example, 30°) around the axis Ax. Each of the brackets B1 to B5 is fixed to the mounting base 14 with a clamp (for example, a toggle clamp) CL. Therefore, the circumferential position of each of the brackets B1 to B5 can be adjusted by loosening the clamp CL.
[0044] Each of the cameras C1 to C5 is attached to a camera mounting portion provided on the brackets B1 to B5. Also, each of the lighting devices L1 to L5 is attached to a lighting mounting portion provided on the brackets B1 to B5. Each of the cameras C1 to C5 is attached to the camera mounting portion using, for example, a screw hole for a tripod. Also, each of the lighting devices L1 to L5 is attached to the lighting mounting portion by fixing the arm portion with a bolt.
[0045] The cameras C1 to C5 and the lighting devices L1 to L5, which are attached to the mounting base 14 via brackets B1 to B5, are arranged on the frame 11 in a predetermined orientation. Specifically, they are arranged within a plane (zy plane) perpendicular to the axis Ax of the multi-eye imaging device 10, facing outward in a radial direction (normal direction) centered on the axis Ax of the multi-eye imaging device 10. More specifically, the cameras C1 to C5 are arranged with their imaging optical axes facing outward in a radial direction (normal direction) centered on the axis Ax of the multi-eye imaging device 10. Furthermore, the cameras C1 to C5 are attached with the bottom surfaces of their camera bodies parallel to the mounting base 14 (parallel to the zy plane) (the bottom edge of the image sensor is attached parallel to the zy plane). As a result, the cameras C1 to C5 are arranged at predetermined intervals circumferentially within the zy plane, centered on the axis Ax of the multi-eye imaging device 10. Illumination devices L1 to L5 are arranged with their illumination directions facing outward in radial directions (normal directions) centered on axis Ax of multi-eye imaging device 10. As a result, cameras C1 to C5 and illumination devices L1 to L5 are arranged radially within the zy plane with axis Ax of multi-eye imaging device 10 as the center.
[0046] As described above, the brackets B1 to B5 are attached so as to be movable within a predetermined angular range in the circumferential direction around the axis Ax of the multi-eye imaging device 10. FIGS. 3 and 4 show the brackets B1 to B5 fixed at their reference positions. By fixing the brackets B1 to B5 at their reference positions, the first camera C1 and the first lighting device L1 are positioned at 330° (-30°) in a front view ( FIG. 3 ). The second camera C2 and the second lighting device L2 are positioned at 30°. The third camera C3 and the third lighting device L3 are positioned at 90°. The fourth camera C4 and the fourth lighting device L4 are positioned at 150°. The fifth camera C5 and the fifth lighting device L5 are positioned at 210°.
[0047] Each bracket B1 to B5 is attached so as to be movable within a range of ±15° in the circumferential direction from a reference position, and therefore the placement positions of each camera C1 to C5 and lighting device L1 to L5 can be adjusted within a range of ±15° in the circumferential direction from the reference position.
[0048] In the multi-eye photography device 10 configured as described above, five cameras C1-C5 and lighting devices L1-L5 are arranged on a circle centered on the axis Ax of the device. The positions of the cameras C1-C5 are adjusted so that the shooting areas of adjacent cameras overlap. In this case, it is preferable to adjust the positions of the cameras C1-C5 so that an overlap rate of at least 10% is ensured. The overlap rate refers to the percentage of overlap between the shooting areas of adjacent cameras (the percentage of overlap between the images captured).
[0049] The cameras C1 to C5 used are digital cameras. There are no particular limitations on the type of digital camera, as long as it has the function of electrically recording images (still images or moving images). As an example, a digital camera with an interchangeable lens is used. In this embodiment, the cameras C1 to C5 each have storage (storage media) and store captured images in the storage. The storage may be a built-in memory or an exchangeable memory card.
[0050] The lighting devices L1 to L5 used are not particularly limited. As an example, a halogen lamp is used. Other lighting devices such as LED (light emitting diode) lamps and xenon lamps can also be used. In this embodiment, lighting devices with an adjustment function for the illumination angle (illumination direction) are used. Each of the lighting devices L1 to L5 rotates (swings back and forth) around an axis perpendicular to the optical axis of the cameras C1 to C5 to adjust the illumination angle (illumination direction). Each of the lighting devices L1 to L5 has an illumination range that can cover the shooting range of the corresponding camera C1 to C5.
[0051] [Electrical Configuration of the Multi-Eye Photography Apparatus] FIG. 5 is a block diagram showing the electrical configuration of the multi-eye photography apparatus.
[0052] As shown in FIG. 5, the multi-eye photographing device 10 has a relay device 20, and is connected to the control device 100 via the relay device 20 so as to be able to communicate with each other.
[0053] The relay device 20 is configured, for example, by a computer equipped with a communication function. The cameras C1 to C5 and the lighting devices L1 to L5 are connected to the relay device 20. The connection between the cameras C1 to C5 and the relay device 20 is not particularly limited. They may be connected via a wired connection to enable communication, or may be connected wirelessly to enable communication.
[0054] The form of communication between the control device 100 and the relay device 20 is not particularly limited. Communication may be wired or wireless. As an example, in this embodiment, the control device 100 and the relay device 20 are connected via a wireless local area network (LAN).
[0055] [Control Device] FIG. 6 is a diagram illustrating an example of the hardware configuration of the control device.
[0056] As shown in Fig. 6, the control device 100 includes a central processing unit (CPU) 111, a read only memory (ROM) 112, a random access memory (RAM) 113, an auxiliary storage device 114, an input device 115, a display device 116, and a communication interface (I / F) 117. Generally, this type of configuration can be realized by a computer. As an example, in this embodiment, the control device 100 is configured as a notebook personal computer. The control device 100 is an example of a processing device.
[0057] The control device 100 functions as a control device when a CPU 111, which is a processor, executes a predetermined program. The program executed by the CPU 111 is stored in the ROM 112 or the auxiliary storage device 114.
[0058] The auxiliary storage device 114 constitutes a storage unit of the control device 100. The auxiliary storage device 114 is constituted by, for example, a hard disk drive (HDD), a solid state drive (SSD), or the like.
[0059] The input device 115 constitutes an operation unit of the control device 100. The input device 115 is constituted by, for example, a keyboard, a mouse, a touch panel, and the like.
[0060] The display device 116 constitutes a display unit of the control device 100. The display device 116 is constituted by, for example, an LCD (liquid crystal display), an OLED (organic light-emitting diode) display, or the like.
[0061] The communication interface 117 constitutes a communication unit of the control device 100. The communication interface 117 is configured to be able to communicate with at least the relay device 20 using a predetermined communication method. As an example, in this embodiment, the communication interface 117 is configured to be able to communicate using a wireless LAN.
[0062] [Functions of the Control Device] The control device 100 has a function to control the multi-eye photography device 10 and a function (image processing function) to process images captured by the multi-eye photography device 10. The function to control the multi-eye photography device 10 includes a function to control photography by the multi-eye photography device 10 (photography control function).
[0063] [Photography Control Function] FIG. 7 is a functional block diagram of the photography control function of the control device.
[0064] 7, the control device 100 has, as its imaging control functions, functions such as a camera control unit 111A and an illumination control unit 111B. The functions of the camera control unit 111A and the illumination control unit 111B are realized by the CPU 111 executing a predetermined program.
[0065] FIG. 8 is a block diagram showing the main functions of the camera control unit.
[0066] As shown in FIG. 8, the camera control unit 111A mainly has the functions of a photographing control unit 111A1, a photographed image acquisition unit 111A2, a photographed image display control unit 111A3, and a photographed image recording control unit 111A4.
[0067] The shooting control unit 111A1 controls the cameras C1 to C5 mounted on the multi-eye photography device 10 and causes each of the cameras C1 to C5 to shoot. Shooting includes both still image shooting and video shooting. Still image shooting also includes so-called interval shooting. Interval shooting is a function that repeatedly shoots still images at regular intervals. The shooting control unit 111A1 causes each of the cameras C1 to C5 to shoot based on operation input (instructions to shoot) from the input device 115. In the case of video shooting and interval shooting, shooting is started in response to an instruction to start shooting, and stopped in response to an instruction to stop shooting.
[0068] The cameras C1 to C5 shoot in synchronization. Therefore, when shooting still images, the cameras C1 to C5 shoot images simultaneously (including a range that can be considered as nearly simultaneous). When shooting moving images, the cameras C1 to C5 start and finish shooting simultaneously.
[0069] The captured image acquisition unit 111A2 acquires images captured by each of the cameras C1 to C5. These images include not only images that have actually been captured (images obtained in response to a capture instruction) but also so-called live view images.
[0070] The captured image display control unit 111A3 controls the display of images (including live view images) captured by the cameras C1 to C5.
[0071] FIG. 9 is a diagram showing an example of a live view display.
[0072] As shown in FIG. 9 , five image display areas IDA1 to IDA5 are set on the display screen of the display device 116, and live view images from each of the cameras C1 to C5 are displayed in the image display areas IDA1 to IDA5. The image display areas IDA1 to IDA5 are arranged in a layout corresponding to the arrangement of the cameras C1 to C5 in the multi-eye imaging device 10. Therefore, in this embodiment, the image display areas IDA1 to IDA5 are arranged in an arc shape. The first image display area IDA1 displays an image from the first camera C1. The second image display area IDA2 displays an image from the second camera C2. The third image display area IDA3 displays an image from the third camera C3. The fourth image display area IDA4 displays an image from the fourth camera C4. The fifth image display area IDA5 displays an image from the fifth camera C5.
[0073] The captured image recording control unit 111A4 controls the recording of images captured by each of the cameras C1 to C5. In this embodiment, the captured image recording control unit 111A4 creates an image database (DB) 120 in the auxiliary storage device 114 and records the images captured by each of the cameras C1 to C5 in the image database 120. The images from each of the cameras C1 to C5 are recorded in units of images. One unit of image capture is the image taken to generate one composite image. Therefore, for example, when generating a composite image of the entire length of a tunnel, images captured along the entire length of the tunnel structure are recorded as a single group (image group) in a manner that allows them to be distinguished from one another. The captured image recording control unit 111A4 associates information about the camera that captured the images and the order in which they were captured, and records the images from each of the cameras C1 to C5. In other words, the images from each of the cameras C1 to C5 are recorded in a manner that allows them to be distinguished from one another by which camera the images were captured. By recording the images in this manner, the capture position of each image can be estimated. That is, the relative positional relationship between the cameras C1 to C5 (the arrangement of the cameras C1 to C5) is known, and the images are taken at approximately constant distance intervals. Therefore, if it is possible to distinguish which camera took which image and in what order, the approximate shooting position can be identified. For the same reason, it is also possible to grasp the relative positional relationship between each image. In other words, it is possible to identify adjacent images. Therefore, in this embodiment, information about the camera that took the image and information about the shooting order constitute information about the image shooting position.
[0074] Note that the "shooting location" information does not require the ability to identify an exact geographical location; it is sufficient if it is information that can at least identify the relative positional relationship between the images that make up the image group, or information that can at least identify adjacent images.
[0075] Furthermore, "information about the camera that took the image" is information when the camera location is known. Therefore, "information about the camera that took the image" is synonymous with information about the location of the camera that took the image (the location on the camera array).
[0076] The form of association is not particularly limited. It is sufficient that the camera that captured each image can be identified and the order in which the images were captured can be identified. As an example, in this embodiment, information about the camera that captured the image and information about the order in which the images were captured are added to the image as additional information (e.g., meta information), and the images captured by each of the cameras C1 to C5 are recorded.
[0077] The illumination control unit 111B controls the illumination devices L1 to L5 mounted on the multi-eye imaging apparatus 10. In other words, it controls the on / off of the illumination light emitted from the illumination devices L1 to L5. The illumination control unit 111B emits the illumination light based on operation input (on command and off command) from the input device 115.
[0078] [Image Processing Function] FIG. 10 is a functional block diagram of the image processing function of the control device.
[0079] The control device 100 has an image processing function of generating a composite image from a group of images captured by the multiple-eye imaging device 10. In this embodiment, the control device 100 is an example of an image processing device.
[0080] 10 , the control device 100 has functions for generating a composite image, such as a processing target image acquisition unit 111C, a first composite processing unit 111D, a first pass / fail determination unit 111E, a second composite processing unit 111F, a second pass / fail determination unit 111G, a third composite processing unit 111H, a composite image recording control unit 111J, and a composite image display control unit 111K. The functions of each unit are realized by the CPU 111 executing a predetermined program (image processing program).
[0081] [Processing Target Image Acquisition Unit] The processing target image acquisition unit 111C acquires a group of images to be processed in the compositing process. That is, it acquires a group of images for generating a composite image. The processing target image acquisition unit 111C acquires the processing target image group from the image database 120.
[0082] As described above, information about the camera that captured the images and the order in which they were captured is added to the images recorded in the image database 120. This information provides information about the capture location. Therefore, by acquiring the image group to be processed, information about the capture location (approximate capture location within the tunnel) of each image that makes up the image group can be obtained at the same time.
[0083] [First Combining Processor] The first combining processor 111D performs a predetermined combining process on the group of images acquired by the processing target image acquiring unit 111C to generate a combined image. In this embodiment, the combining process is performed by so-called feature point matching (corresponding point search). Feature point matching is a process of matching feature points that have a high degree of similarity between images. In general, feature points are detected and feature descriptors are calculated for two images, and feature points that have a high degree of similarity are matched.
[0084] In the synthesis process using feature point matching, parameters required for the synthesis process (synthesis parameters) are determined from the results of the feature point matching, and the synthesis process is performed based on the determined synthesis parameters. More specifically, an image is projected onto a shape model based on the determined synthesis parameters to generate a single synthesized image.
[0085] FIG. 11 is a conceptual diagram of the synthesis process using feature point matching.
[0086] In the synthesis process using feature point matching, the synthesis parameters are determined from the results of feature point matching between images, including the orientation parameters of the shape model, the orientation parameters (rotation matrix, translation vector) of each camera corresponding to each image, and the lens distortion parameters of each camera corresponding to each image.
[0087] The shape model is selected according to the object to be photographed (subject). If the object to be photographed is a plane, a planar model is selected. In the case of a planar model, a rotation matrix and a translation vector are determined as pose parameters. If the object to be photographed is a curved surface, a cylindrical model is selected. In the case of a cylindrical model, a rotation matrix, a translation vector, and the radius of the cylinder are determined as pose parameters.
[0088] In the case of a tunnel structure with an arc-shaped inner wall (curved tunnel), a cylindrical model is selected as the shape model. Therefore, in this case, the rotation matrix, translation vector, and cylinder radius are determined as the orientation parameters of the shape model.
[0089] The first synthesis processing unit 111D performs feature point matching between images of the group of images acquired by the processing target image acquisition unit 111C and determines synthesis parameters from the results of the feature point matching. Furthermore, the first synthesis processing unit 111D projects the images onto a shape model based on the determined synthesis parameters to generate a single synthesized image. Furthermore, when a cylindrical model is selected as the shape model, the first synthesis processing unit 111D develops the image projected onto the shape model into a plane to generate a synthesized image.
[0090] Hereinafter, the composite image generated by the first synthesis processing unit 111D will be referred to as a "primary composite image" as necessary to distinguish it from other composite images.
[0091] In the present embodiment, the compositing process by feature point matching performed by the first compositing processing unit 111D is an example of a first compositing process. Also, in the present embodiment, the composite image (primary composite image) generated by the first compositing processing unit 111D is an example of a first composite image.
[0092] [First Pass / Fail Judgment Unit] The first pass / fail judgment unit 111E judges whether the composite image (primary composite image) generated by the first composition processing unit 111D passes or fails. In this embodiment, the pass / fail judgment is made based on the quality of the primary composite image. Specifically, if the quality of the primary composite image meets a specified quality standard, it is judged as pass. Therefore, if the quality of the primary composite image does not meet the specified quality standard, it is judged as fail.
[0093] In this embodiment, first, the degree of fragmentation of the generated primary composite image is determined to determine whether the quality of the primary composite image satisfies a specified quality standard, and second, the degree of distortion of the generated primary composite image is determined to determine whether the quality of the primary composite image satisfies a specified quality standard.
[0094] Here, "fragmenting a composite image" means that a composite image is generated by dividing it into a plurality of images (fragments).
[0095] FIG. 12 is a diagram showing an example of a fragmented composite image.
[0096] FIG. 12 shows an example in which an image that should have been generated as a single image is separated (fragmented) into three images (fragments) IF1 to IF3.
[0097] 13 and 14 are diagrams showing an example of a composite image in which distortion occurs.
[0098] 13 and 14 show examples of images that should have been synthesized as a nearly rectangular image but are partially distorted. Fig. 13 shows an example of an image that is synthesized with a distorted central portion of the bottom side. Fig. 14 shows an example of an image that is synthesized with a distorted upper right corner.
[0099] Fragmentation and distortion of the composite image occur due to a failed synthesis process. In synthesis using feature point matching, synthesis fails due to missing or incomplete shots, insufficient overlap between adjacent images, blurred or out-of-focus images, etc., resulting in fragmentation and distortion. Furthermore, if there are differences in the appearance of the same location between images in the overlapping area (for example, differences in the way shadows are cast), feature point matching may not be possible, and synthesis may fail. Furthermore, synthesis may fail if the shape of the target object (subject) significantly deviates from the assumed shape model.
[0100] (1) Pass / Fail Judgment Based on the Degree of Fragmentation The first pass / fail judgment unit 111E judges the degree of fragmentation of the generated primary composite image and judges whether the quality of the primary composite image meets the specified quality standard. That is, it judges whether the degree of fragmentation meets the specified standard and judges whether the quality of the primary composite image meets the specified quality standard. If the degree of fragmentation meets the specified standard, the quality of the primary composite image is judged to meet the specified quality standard, and the primary composite image is judged to pass. On the other hand, if the degree of fragmentation does not meet the specified standard, the quality of the primary composite image is judged to not meet the specified quality standard, and the primary composite image is judged to fail.
[0101] As an example, in this embodiment, the presence or absence of fragmentation is used as a criterion for determining the "degree of fragmentation." That is, the presence or absence of fragmentation in the primary composite image is determined to determine whether the quality of the primary composite image satisfies the specified quality standard. Therefore, if a primary composite image is generated through fragmentation, the quality of the primary composite image is determined to not satisfy the specified quality standard, and the image is determined to be unsatisfactory. On the other hand, if a primary composite image is generated as a single composite image without fragmentation, the quality of the primary composite image is determined to satisfy the specified quality standard, and the image is determined to be acceptable.
[0102] In this embodiment, the determination criterion for determining the degree of fragmentation of the generated primary combined image is an example of the first criterion.
[0103] (2) Pass / Fail Judgment Based on the Degree of Distortion The first pass / fail judgment unit 111E judges the degree of distortion of the generated primary composite image and judges whether the quality of the primary composite image satisfies the specified quality standard. If the degree of distortion is within the allowable range, it is judged that the specified quality standard is met and the image is judged as pass. On the other hand, if the degree of distortion exceeds the allowable range, it is judged that the quality does not satisfy the specified quality standard and the image is judged as fail.
[0104] As an example, in this embodiment, the degree of distortion is determined by determining the degree of deviation from the image shape that should have been generated. The "image shape that should have been generated" refers to the shape of the image that would have been generated if the synthesis process had been performed correctly, and is determined by the shooting range. When photographing a structure for inspection, a rectangular cut-out range is generally set as the shooting range (in the case of a tunnel structure, the range that becomes rectangular when unfolded on a plane is set as the shooting range).
[0105] When a tunnel structure is photographed using the multi-eye imaging device 10, the composite image (a properly composited composite image) generated by planar expansion is an approximately rectangular image. Therefore, the degree of distortion can be determined by calculating the degree of deviation from a rectangle. The degree of deviation from a rectangle can be calculated from the difference between the maximum and minimum widths of the image in the horizontal or vertical direction. In other words, the greater the distortion from a rectangle, the greater the difference, and the degree of distortion can be determined from the difference. In this embodiment, the difference between the maximum and minimum widths of the image in the horizontal direction and the maximum and minimum widths of the image in the vertical direction is calculated and compared with a threshold. If at least one of the differences exceeds the threshold, the image is determined to have a large distortion (the distortion exceeds the allowable range). In other words, the degree of distortion is determined to exceed the allowable range. Therefore, in this embodiment, if the difference between the maximum and minimum widths in the horizontal and vertical directions of the image is equal to or less than a threshold, the degree of distortion is determined to be within the allowable range.
[0106] Fig. 13 shows an example where a portion of the image (the center portion) is distorted vertically. In this case, the difference between the maximum and minimum values of the image's vertical width exceeds the threshold. Fig. 14 shows an example where a portion of the image (the upper right corner portion) is distorted horizontally. In this case, the difference between the maximum and minimum values of the horizontal width exceeds the threshold.
[0107] The width of the image is calculated, for example, from the number of pixels. In this case, the difference between the maximum and minimum values of the horizontal width of the image is calculated from the difference between the maximum and minimum values of the horizontal number of pixels. Similarly, the difference between the maximum and minimum values of the vertical width of the image is calculated from the difference between the maximum and minimum values of the vertical number of pixels.
[0108] Generally, when generating a composite image, the composite result is shaped into a rectangular image and output. That is, even if the composite image is distorted, it is shaped into a rectangular image and output by so-called padding. Padding is a process of filling in with meaningless pixels. For example, in Figures 13 and 14, the surrounding black area is a padded area. Figures 13 and 14 show an example of shaping into a rectangular image by padding with black pixels.
[0109] The "image width" here does not refer to the width of the image after padding, but rather to the width (vertical and / or horizontal) of the image itself generated by compositing. Therefore, for example, in the case of a composite image after padding, it refers to the width (vertical and / or horizontal) of the area excluding the padding area. In the examples of Figures 13 and 14, it refers to the width of the image area excluding the surrounding black area.
[0110] In this embodiment, the criterion for determining the degree of distortion of the generated primary composite image is another example of the first criterion, and the measurement direction (horizontal and / or vertical directions) of the image width when calculating the degree of distortion is an example of the first direction.
[0111] [Second Composition Processing Unit] The second composition processing unit 111F performs composition processing on the fragmented primary composite images to generate a composite image with a less degree of fragmentation. Since the processing target is a fragmented primary composite image, the image to be processed by the second composition processing unit 111F is a primary composite image that has been determined to be unacceptable by the first pass / fail determination unit 111E due to fragmentation.
[0112] In this embodiment, the fragmented primary composite image is synthesized using the information on the shooting position of each image. That is, the information on the shooting position (in this embodiment, information on the camera that took the images and information on the shooting order) is used to adjust the arrangement positions of the multiple fragments, and a composite image with a less fragmented degree, or more preferably, a single composite image (one composite image), is generated.
[0113] FIG. 15 is a conceptual diagram of the compositing process in the second compositing processor.
[0114] 15 shows an example in which a primary composite image is generated by separating (fragmenting) it into three images (fragments) IF1 to IF3. Hereinafter, as necessary, "image IF1" will be referred to as the "first image fragment IF1," "image IF2" as the "second image fragment IF2," and "image IF3" as the "third image fragment IF3" to distinguish between the images.
[0115] As described above, in the image group used to generate the primary composite image, each image has information about the camera that took the image and the order in which it was taken, as information about the photographing position. By having this photographing position information, each image constituting the image group can identify its adjacent images. This relationship also applies to composite images.
[0116] First, consider the relationship between the first image fragment IF1 and the second image fragment IF2.
[0117] In the first image fragment IF1, an image Im(1, n) constituting a part of it is assumed to be the nth image captured by the first camera C1, and in the second image fragment IF2, an image Im(1, n+1) constituting a part of it is assumed to be the (n+1)th image captured by the first camera C1.
[0118] Image Im(1,n) of the first image fragment IF1 and image Im(1,n+1) of the second image fragment IF2 are images captured by the same first camera C1. Furthermore, image Im(1,n+1) of the second image fragment IF2 is the image captured chronologically after image Im(1,n) of the first image fragment IF1. From this, it can be seen that image Im(1,n) of the first image fragment IF1 is the image that should be placed immediately to the left of image Im(1,n+1) of the second image fragment IF2.
[0119] Therefore, by positioning the image Im(1, n) of the first fragment image IF1 to the left of the image Im(1, n+1) of the second fragment image IF2 and combining them, the separation between the first fragment image IF1 and the second fragment image IF2 can be eliminated.
[0120] Next, the relationship between the second image fragment IF2 and the third image fragment IF3 will be considered.
[0121] In the second image fragment IF2, an image Im(5, m) constituting a part of it is assumed to be the mth image captured by the fifth camera C5. Also, in the third image fragment IF3, an image Im(5, m+1) constituting a part of it is assumed to be the (m+1)th image captured by the fifth camera C5.
[0122] Image Im(5,m) of the second image fragment IF2 and image Im(5,m+1) of the third image fragment IF3 are images taken by the same fifth camera C5. Furthermore, image Im(5,m+1) of the third image fragment IF3 is the image taken chronologically after image Im(5,m) of the second image fragment IF2. From this, it can be seen that image Im(5,m) of the second image fragment IF2 is the image that should be placed immediately to the left of image Im(5,m+1) of the third image fragment IF3.
[0123] Therefore, by positioning the image Im(5, m) of the second fragment image IF2 to the left of the image Im(5, m+1) of the third fragment image IF3 and combining them, the separation between the second fragment image IF2 and the third fragment image IF3 can be eliminated.
[0124] FIG. 16 is a diagram showing an example of a synthesis result.
[0125] In the example shown in Figure 16, images Im(1,n) and Im(1,n+1) are arranged adjacent to each other, and a first fragment image IF1 and a second fragment image IF2 are composited together. Also, images Im(5,m) and Im(5,m+1) are arranged adjacent to each other, and a second fragment image IF2 and a third fragment image IF3 are composited together. The generated composite image CI2 is of inferior quality compared to a normally generated primary composite image, but an image sufficient for grasping the state of the object (subject) is obtained.
[0126] In the above example, adjacent images are identified using information on the order of shooting (information on the position in the direction of movement of the multi-eye photography device 10), but it is also possible to identify adjacent images using information on the camera that took the images (information on the camera's placement position), or both.
[0127] Furthermore, in the above example, the configuration focuses on only one image, identifies adjacent images, and determines the placement position, but it is also possible to identify the adjacency relationship between multiple images and determine the placement position.
[0128] Furthermore, it is preferable that the composite image is generated not by simply arranging the fragment images in predetermined positions, but by applying processes such as enlarging, reducing, and rotating the fragment images as necessary. These processes are facilitated by identifying the adjacency relationships among multiple images and using the results. Therefore, it is preferable that the composite image is generated by identifying the adjacency relationships among multiple images.
[0129] In this way, the second composition processing unit 111F uses information on the image capture positions to adjust the arrangement positions of the image fragments and compose them, thereby eliminating fragmentation. By eliminating all fragmentation, a single composite image is generated.
[0130] In the present embodiment, the compositing process performed by second compositing processing unit 111F is an example of a second compositing process. Also, in the present embodiment, the composite image generated by second compositing processing unit 111F is an example of a second composite image. Hereinafter, as necessary, the composite image generated by second compositing processing unit 111F will be referred to as a "secondary composite image" to distinguish it from other composite images.
[0131] [Second Pass / Fail Judgment Unit] The second pass / fail judgment unit 111G judges whether the composite image (secondary composite image) generated by the second composition processing unit 111F passes or fails. In this embodiment, similar to the first pass / fail judgment unit 111E, the pass / fail judgment is made based on the quality of the secondary composite image. Therefore, the secondary composite image is judged as pass only when the quality meets the specified quality standard (if the quality of the secondary composite image does not meet the specified quality standard, it is judged as fail).
[0132] Similar to the first pass / fail judgment unit 111E, the second pass / fail judgment unit 111G first judges the degree of fragmentation of the generated secondary composite image to determine whether the quality of the secondary composite image satisfies a specified quality standard, and second judges the degree of distortion of the generated secondary composite image to determine whether the quality of the secondary composite image satisfies a specified quality standard.
[0133] (1) Pass / Fail Judgment Based on the Degree of Fragmentation In this embodiment, whether or not the specified quality standard is satisfied is judged based on whether or not fragmentation has been resolved. In other words, whether or not the specified quality standard is satisfied is judged based on whether or not a single composite image has been generated. If the second compositing processing unit 111F has been able to generate a single composite image, it is judged that the specified quality standard is satisfied and the result is a pass. On the other hand, if the second compositing processing unit 111F has not been able to generate a single composite image, i.e., if compositing has failed, it is judged that the specified quality standard is not satisfied and the result is a fail.
[0134] In this embodiment, the determination criterion for determining the degree of fragmentation of the generated secondary composite image is an example of the second criterion.
[0135] (2) Pass / Fail Judgment Based on the Degree of Distortion The second pass / fail judgment unit 111G judges the degree of distortion of the generated secondary composite image and judges whether the quality of the secondary composite image satisfies the specified quality standard. If the degree of distortion is within the allowable range, it is judged that the quality meets the specified quality standard and is judged as pass. On the other hand, if the degree of distortion exceeds the allowable range, it is judged that the quality does not meet the specified quality standard and is judged as fail.
[0136] As with the first pass / fail judgment unit 111E, in this embodiment, the degree of distortion is determined by calculating the degree of deviation from the image shape that should have been generated. Therefore, the degree of distortion is determined by calculating the difference between the maximum and minimum values of the horizontal width and the maximum and minimum values of the vertical width of the generated secondary composite image. If at least one of the differences exceeds a threshold, the distortion is determined to be large (the distortion exceeds the allowable range) and the image is determined to be unacceptable.
[0137] Note that the "image width" here does not refer to the width of the image after padding, but rather to the width of the image itself generated by synthesis, i.e., the width of the area excluding the padding area.
[0138] [Third Composition Processing Unit] The third composition processing unit 111H generates a composite image of a predetermined format from the image group when the primary composite image is determined to be unacceptable because the degree of distortion does not meet the specified standard, and when the secondary composite image is determined to be unacceptable. The secondary composite image is determined to be unacceptable when the degree of distortion does not meet the specified standard, and when fragmentation cannot be resolved (when a single composite image cannot be generated).
[0139] FIG. 17 is a conceptual diagram of generation of a composite image by the third composite image processing unit.
[0140] The third composition processing unit 111H generates a composite image in which each image constituting the image group is arranged based on information about the shooting position. In this embodiment, each image is arranged in a predetermined position based on information about the camera that took the image and information about the shooting order, and a single composite image is generated. Specifically, the images are arranged in a matrix with the vertical axis (vertical direction) representing the position of the camera that took the image and the horizontal axis (horizontal direction) representing the shooting order of each camera, and a single composite image is generated.
[0141] As shown in FIG. 17 , the images captured by the first camera C1 are chronologically ordered (photographed) as Im(1,1), Im(1,2), Im(1,3), ..., Im(1,n). The images captured by the second camera C2 are chronologically ordered as Im(2,1), Im(2,2), Im(2,3), ..., Im(2,n). The images captured by the third camera C3 are chronologically ordered as Im(3,1), Im(3,2), Im(3,3), ..., Im(3,n). The images captured by the fourth camera C4 are chronologically ordered as Im(4,1), Im(4,2), Im(4,3), ..., Im(4,n). The images captured by the fifth camera C5 are chronologically ordered as Im(5,1), Im(5,2), Im(5,3), ..., Im(5,n).
[0142] The composite image IC3 is generated by arranging the images captured by each of the cameras C1 to C5 in chronological order (shooting order) in each row. Furthermore, each column contains images captured at the same time by each of the cameras C1 to C5. In other words, in the multi-eye imaging device 10 of this embodiment, the cameras C1 to C5 synchronously capture images, so when the images captured by each of the cameras C1 to C5 are arranged in shooting order, each column contains images captured at the same time (including substantially the same time).
[0143] FIG. 18 is a diagram showing an example of a composite image generated by the third composite processing unit.
[0144] Figure 18 shows an example of a composite image CI3 generated from images obtained by taking interval photographs while moving at a substantially constant speed inside a tunnel structure. In particular, Figure 18 shows an example in which images are taken nine times using five cameras C1 to C5. In this case, one composite image CI3 is generated from a group of 5 x 9 = 45 images.
[0145] 18, in the generated composite image CI3, the images captured by the cameras C1 to C5 are arranged in the order of capture in each row, and the images captured by the cameras C1 to C5 in the same order (images captured at the same timing) are arranged in each column.
[0146] In this way, the third synthesis processing unit 111H arranges each image constituting the image group in a predetermined array based on information about the shooting position (in this embodiment, information about the camera that took the images and information about the shooting order), and generates a single synthetic image CI3. Hereinafter, as necessary, the synthetic image CI3 generated by the third synthesis processing unit 111H will be referred to as a "juxtaposed synthetic image" to distinguish it from other synthetic images. The juxtaposed synthetic image CI3 is of inferior quality compared to a normally generated primary synthetic image, but provides an image sufficient for understanding the state of the object (subject). In this embodiment, the "juxtaposed synthetic image" is an example of a second synthetic image.
[0147] In this embodiment, one of a primary composite image, a secondary composite image, or a juxtaposed composite image is generated from the image group.
[0148] [Composite Image Recording Control Unit] The composite image recording control unit 111J controls the recording of generated composite images (primary composite image, secondary composite image, and juxtaposed composite image). The composite image recording control unit 111J associates the generated composite image with the original image group and records it in the image database 120. The composite image recording control unit 111J records the generated composite image in the image database 120 automatically or in response to a recording instruction from the user.
[0149] [Composite Image Display Control Unit] The composite image display control unit 111K controls the display of the generated composite images (primary composite image, secondary composite image, and juxtaposed composite image). The composite image display control unit 111K displays the generated composite image on the display device 116. The composite image display control unit 111K enlarges, reduces, moves, etc. the composite image displayed on the display device 116 in response to an instruction from the user.
[0150] [Operation of the Imaging System] Here, an example will be described in which the inner wall surface of a tunnel structure is photographed in sections by the multi-eye imaging device 10, and a composite image of the inner wall surface of the tunnel structure is generated from the obtained images.
[0151] The photographing is performed by moving the multi-eye photography device 10 using a trolley Tr. As an example, the inner wall surface of a tunnel structure is photographed by moving the trolley Tr at a substantially constant speed and performing interval photographing. In this case, the traveling speed of the trolley Tr and the interval photographing interval for interval photographing are set so that the overlap rate (overlap ratio) of images in the direction of travel of the trolley Tr satisfies a specified condition. The overlap rate condition is determined from the perspective of synthesis processing. As an example, the overlap rate condition is 10% or more. Therefore, the traveling speed of the trolley Tr and the interval photographing interval for interval photographing are set so that the overlap rate is at least 10% or more.
[0152] A group of images obtained by photography is recorded in the image database 120 on a photography unit basis so that they can be distinguished from other groups of images obtained by photography. Each image constituting the group of images is associated with information about the photography location (in this embodiment, information about the camera that took the image, and information about the photography order), and is recorded in the image database 120.
[0153] After the photographing, a composite image is generated in accordance with an instruction from the user. In the photographing system 1 of this embodiment, the process of generating the composite image is performed in the control device 100. More specifically, the CPU 111 of the control device 100 performs the process of generating the composite image from the group of images obtained by photographing.
[0154] [Composite Image Generation Processing] FIG. 19 is a flowchart showing the procedure of composite image generation processing (image processing method).
[0155] First, a group of images to be processed is acquired (step S11). As described above, in the photography system 1 of this embodiment, a group of images obtained by photography is recorded in the image database 120. The CPU 111 acquires the group of images to be processed from the image database 120.
[0156] Next, a process for generating a primary composite image is performed (step S12). In this embodiment, a synthesis process using feature point matching is performed on the acquired image group to generate a primary composite image.
[0157] Next, the generated primary composite image is subjected to a pass / fail determination process (first pass / fail determination process) (step S13). In this embodiment, the pass / fail determination is made by determining whether the quality of the primary composite image satisfies a specified quality standard. Specifically, first, the degree of fragmentation is determined, the image quality is determined, and the pass / fail determination is made. Second, the degree of distortion is determined, the image quality is determined, and the pass / fail determination is made. In this embodiment, if there is no fragmentation and the distortion is within an allowable range, the image is determined as pass.
[0158] Next, based on the result of the pass / fail determination process, it is determined whether the primary composite image passes or fails (step S14). That is, it is determined whether the compositing process (compositing process by feature point matching) has succeeded or failed. If the pass / fail determination result indicates that the primary composite image passes, the compositing process ends. In this case, the primary composite image is used as the image resulting from the compositing process.
[0159] On the other hand, if the primary composite image is judged to be "fail" as a result of the pass / fail judgment, the reason is determined, that is, whether or not it was rejected due to fragmentation is judged (step S15).
[0160] If the image is determined to be unsatisfactory due to fragmentation, a process of generating a secondary composite image from the fragmented primary composite image is performed (step S16). Specifically, the positions of the multiple fragments are adjusted using the information on the shooting positions of each image, and a composite image with a less fragmented degree (preferably a single composite image) is generated (see FIG. 16).
[0161] When the secondary composite image is generated, a pass / fail judgment process (second pass / fail judgment process) is performed on the generated secondary composite image (step S17). In this embodiment, similar to the pass / fail judgment process on the primary composite image (first pass / fail judgment process), the quality of the generated secondary composite image is judged to be whether it satisfies a specified quality standard, and the pass / fail is judged. Then, based on the pass / fail judgment process result, it is judged whether the secondary composite image is pass / fail (step S18). In other words, the success or failure of the compositing process (compositing process by adjusting the arrangement position based on the information on the shooting position) is judged. If the pass / fail judgment result indicates that the secondary composite image is "pass," the compositing process ends. In this case, the secondary composite image is used as the image resulting from the compositing process.
[0162] If the secondary composite image is determined to be unacceptable (if step S18 is "No"), or if the primary composite image is determined to be unacceptable due to distortion (if step S15 is "No"), a process of generating a juxtaposed composite image is performed (step S19). That is, the images constituting the image group are arranged side by side based on the information on their shooting positions to generate a single composite image (see FIG. 18). If a juxtaposed composite image is generated, the juxtaposed composite image is used as the image resulting from the composition process.
[0163] The composite image generation process is completed through the above series of steps. As described above, if the generation of the primary composite image is successful (if the primary composite image is acceptable), the primary composite image is used as the image resulting from the synthesis process. Also, if the generation of the primary composite image fails but the generation of the secondary composite image is successful (if the secondary composite image is acceptable), the secondary composite image is used as the image resulting from the synthesis process. If the generation of any composite image fails, the juxtaposed composite image is used as the image resulting from the synthesis process. Therefore, one composite image is always generated.
[0164] In this way, according to the control device 100 of the present embodiment, even if generation of the primary composite image fails, one composite image can always be generated. The composite images (secondary composite image and juxtaposed composite image) generated when generation of the primary composite image fails are inferior in quality compared to a normally generated primary composite image (a composite image that has been successfully generated), but an image sufficient for grasping the state of the target object (subject) can be obtained (see FIGS. 16 and 18 ).
[0165] The generated composite images (primary composite image, secondary composite image, and juxtaposed composite image) are displayed on the display device 116. The generated composite images are also recorded in the image database 120 automatically or in response to a recording instruction from the user.
[0166] [Modifications] [Generation of Primary Composite Image] In the above embodiment, the primary composite image is generated by a synthesis process using feature point matching, but the method for generating the primary composite image is not limited to this. The primary composite image may also be generated using other methods (so-called panoramic synthesis methods).
[0167] Furthermore, it is preferable to set the synthesis parameters appropriately according to the subject of photography, etc. For example, in the case of planar synthesis, which targets a plane, the projective transformation matrix of each image can be used as the synthesis parameter.
[0168] [Information about Shooting Position] In the above embodiment, the information about the shooting position is configured to acquire information about the camera that took the image (information about the position of the camera on the multi-eye imaging device) and information about the order in which the images were taken, but the information acquired as the information about the shooting position is not limited to this.
[0169] As described above, the "photography position" information may be information that can identify at least the relative positional relationship between the images that make up the image group, or information that can identify at least adjacent images.
[0170] For example, if cameras C1 to C5 have a global positioning system (GPS) function or an indoor messaging system (IMES) function as an indoor GPS, the GPS function or IMES function may be used to acquire information about the shooting location. In this case, the GPS or IMES location information (latitude, longitude, and altitude) is added to the captured image and recorded (for example, as tag information).
[0171] The GPS function or IMES function may be provided in the multi-eye imaging device 10 or the dolly Tr. If each camera is provided with a GPS function or IMES function, information on the camera's location can be omitted, because the location where the image was taken can be identified using location information from the GPS or IMES.
[0172] Furthermore, when information on the order of shooting is used as information on the shooting position, the information on the order of shooting can be configured to be obtained indirectly from other information. For example, when information on the date and time of shooting (a so-called timestamp) is added as additional information (such as tag information) and images are recorded, the order of shooting can be obtained by determining the order of shooting from the information on the date and time of shooting. Furthermore, when images are recorded with sequential file names (when the number included in the file name is incremented by 1 each time an image is recorded), the order of shooting can be obtained by determining the order of shooting from the file name.
[0173] Furthermore, for example, if the cart Tr is equipped with a odometer or the like, distance information (information about the distance from a reference point) obtained from the odometer can be used to obtain information about the photographing position (information about the distance from the reference point or information about the position coordinates). In this case, distance information at the time of photographing is obtained from the odometer and recorded in association with the image. The reference point is, for example, the position where photographing began (the position where the cart began moving).
[0174] [Feature Point Matching] As described above, in the synthesis process using feature point matching, feature point matching is performed as a pre-processing step for synthesis (a process necessary for determining synthesis parameters).
[0175] Generally, feature point matching is performed comprehensively among all images in an image group. If feature point matching is performed comprehensively among all images, the number of feature points that become matching candidates increases, and the probability of mismatching similar feature points increases.
[0176] Basically, feature point matching only needs to be performed between overlapping images (images with overlapping areas). By limiting the targets for feature point matching to these images, the probability of mismatching can be reduced.
[0177] In the photography system 1 of the above embodiment, each image constituting an image group has information about the photography position, so it is possible to identify images that have been photographed in duplicate. Therefore, when performing feature point matching, the information about the photography position is used to identify images that have been photographed in duplicate. Then, feature point matching is performed only between the identified images. This makes it possible to suppress erroneous matching. Furthermore, by reducing the number of images to be processed, the processing load of calculations can also be reduced.
[0178] FIG. 20 is a conceptual diagram of a method for identifying overlapping images.
[0179] 20 is a diagram showing a group of images captured by a camera array arranged based on their shooting positions. The vertical arrangement corresponds to the placement positions of the cameras, and the horizontal arrangement corresponds to the chronological order of the images.
[0180] In FIG. 20, assume that the image of interest is image Im(i,j). Image Im(i,j) is the jth image captured by the i-th camera. Assuming that each image was captured correctly, at least the following four images are images that overlap with image Im(i,j). The first image is image Im(i,j-1) captured by the i-th camera immediately before image Im(i,j). The second image is image Im(i,j+1) captured by the i-th camera immediately after image Im(i,j). The third image is image Im(i-1,j) captured by the i-1th camera adjacent to the i-th camera at the same time as image Im(i,j). The fourth image is image Im(i,j+1) captured by the i+1th camera adjacent to the i-th camera at the same time as image Im(i,j). These four images are adjacent to image Im(i,j).
[0181] Each image constituting an image group has information about its shooting position, and therefore, by using the information about the shooting position, adjacent images can be identified.
[0182] In this way, information on the shooting positions is used to identify overlapping images and limit the processing targets when performing feature point matching, thereby reducing erroneous matching.
[0183] The range of overlapping images varies depending on the camera arrangement and shooting conditions (movement speed, shooting interval, etc.). Therefore, it is preferable to determine the range of overlapping images (range of adjacent images) depending on the camera arrangement, shooting conditions, etc. For example, if there is a prescribed or larger overlapping area (a compositeable overlapping area) between the image taken two images before the image of interest in chronological order (image Im(i, j-2) in FIG. 20) and the image taken two images after (image Im(i, j+2) in FIG. 20), it is preferable to subject these images to feature point matching. Furthermore, for example, if there is a prescribed or larger overlapping area between the images taken immediately before and after the image of interest in chronological order (image Im(i-1, j-1), image Im(i-1, j+1), image Im(i+1, j-1), and image Im(i+1, j+1) in FIG. 20) by adjacent cameras, it is preferable to subject these images to feature point matching.
[0184] [Acceptance / failure judgment based on distortion] In the above embodiment, the method for judging the degree of distortion is configured to measure the maximum and minimum values of the vertical and / or horizontal widths of the generated composite images (primary composite image and secondary composite image), and judge whether the degree of distortion is within the allowable range depending on whether the difference between the maximum and minimum values is equal to or less than a threshold. The method for judging the degree of distortion is not limited to this.
[0185] (1) Method using information on shooting conditions The shape of the generated composite image can be estimated from the shooting conditions of the source images. For example, in the shooting system 1 of the above embodiment, if the target object (subject) is photographed while the multi-eye camera 10 is moved on the dolly Tr, the generated composite image will be rectangular. Furthermore, the size of the image (vertical width and horizontal width) can also be estimated from the shooting conditions (shooting resolution, shooting interval, etc.).
[0186] Therefore, the vertical width and / or horizontal width of the composite image to be generated is estimated using information on the shooting conditions, and the degree of distortion is determined by comparing with the estimated values.
[0187] Specifically, the vertical and / or horizontal widths of the generated composite image are measured and compared with the estimated values of the vertical and / or horizontal widths of the composite image estimated from the shooting conditions. The comparison is performed, for example, by calculating the difference between the measured value and the estimated value. If the calculated difference is equal to or less than a threshold, it is determined that the distortion is within the allowable range. In other words, if the difference is equal to or less than the threshold, it means that the difference from the estimated shape or width is small, and therefore it can be determined that the distortion is small.
[0188] Preferably, the vertical and horizontal widths of the generated composite image are measured and compared with estimated values of the vertical and horizontal widths of the composite image estimated from the shooting conditions. Then, if the difference between the measured and estimated widths in both the vertical and horizontal directions is equal to or less than a threshold, the composite image is determined to be acceptable. The widths can be measured, for example, by the number of pixels.
[0189] In this example, information on the shooting conditions is required. The information on the shooting conditions may be input by the user via the input device 115, or may be automatically acquired from the camera.
[0190] In this example, the width in the vertical direction and / or the width in the horizontal direction is an example of the width in the second direction.
[0191] (2) Method Using Synthesis Parameters In synthesis processing based on synthesis parameters, the degree of distortion can also be determined from the synthesis parameters. Generally, if there is an image in which the synthesis parameters (camera attitude parameters, camera lens distortion parameters) significantly deviate from those of the surrounding images, distortion will occur in the generated synthetic image. Therefore, by focusing on the synthesis parameters, it is possible to determine whether or not there is distortion in the synthetic image (whether or not there is distortion beyond the allowable range). Specifically, after determining synthesis parameters from the result of feature point matching, it is determined whether or not there is an image in which the synthesis parameters significantly deviate from those of the surrounding images. If there is no image in which the synthesis parameters significantly deviate from those of the surrounding images, it is determined that the distortion is within the allowable range. On the other hand, if there is an image in which the synthesis parameters significantly deviate from those of the surrounding images, it is determined that the distortion exceeds the allowable range. Whether or not an image has significantly deviated from the surrounding images in terms of synthesis parameters (deviation) is determined by calculating the degree of deviation (deviation) of the synthesis parameters from those of the surrounding images (e.g., adjacent images) and comparing the calculated deviation with a threshold value. If the deviation is equal to or greater than the threshold, it is determined that the image has synthesis parameters that deviate significantly from those of the surrounding images.
[0192] (3) Others The degree of distortion can be determined in a composite manner by combining the above-mentioned determination methods. For example, the degree of distortion can be determined in a composite manner by combining a method of determining by measuring the width of the generated composite image and a method of determining by using a synthesis parameter. In this case, for example, if all the determinations result in a pass, the image is determined to be pass.
[0193] [Generation of Secondary Composite Image] The primary composite image may be fragmented in a distorted state. FIG. 21 is a diagram illustrating an example of a primary image fragmented in a distorted state. FIG. 21 illustrates an example of fragmentation into three images IF1, IF2, and IF3. Image IF1 is the first image fragment IF1, image IF2 is the second image fragment IF2, and image IF3 is the third image fragment IF3. Of the three image fragments IF1 to IF3, the first image fragment IF1 is generated with significant distortion. Simply arranging a primary composite image containing such a highly distorted image based on information about the shooting position to generate a secondary image will not produce a normal secondary composite image (a secondary composite image that can be determined to pass in a pass / fail judgment). Therefore, when generating a secondary composite image from a primary composite image containing a highly distorted image (fragment), it is preferable to correct the distortion before generating the secondary composite image.
[0194] FIG. 22 is a functional block diagram showing an example of generating a secondary composite image by correcting distortion.
[0195] As shown in Fig. 22, a pre-processing unit 111L is provided before the second synthesis processing unit 111F. The pre-processing unit 111L performs pre-processing on the primary synthesis image to be processed (fragmented primary synthesis image) before generating a secondary synthesis image. The pre-processing unit 111L has the functions of a distorted image extraction unit 111L1 and a distortion correction unit 111L2. Each function is realized by the CPU 111 executing a predetermined program.
[0196] The distorted image extraction unit 111L1 detects fragments having distortion exceeding an acceptable range from among the multiple fragments of the fragmented primary composite image. As an example, the distorted image extraction unit 111L1 determines the degree of distortion of each fragment based on the synthesis parameters. Specifically, the distorted image extraction unit 111L1 detects whether or not there are any images constituting each fragment whose synthesis parameters deviate significantly from those of the surrounding images. If an image whose synthesis parameters deviate significantly from those of the surrounding images is detected, the distortion is determined to be large and to be outside the acceptable range. In this example, a fragment having distortion exceeding the acceptable range is an example of a fragment having distortion exceeding a predetermined range.
[0197] When the distorted image extraction unit 111L1 extracts a fragment with distortion exceeding an allowable range, the distortion correction unit 111L2 performs image processing on the fragment to correct the distortion. As an example, distortion is corrected using the following procedure. First, an image whose synthesis parameters deviate significantly from those of the surrounding images is identified. Next, the synthesis parameters of the identified image are corrected to the same synthesis parameters as those of the surrounding images with less distortion. This reduces the distortion occurring in the image of the fragment. Other known image processing techniques related to distortion correction can also be used for distortion correction.
[0198] When distortion correction processing has been performed, the second synthesis processing unit 111F performs synthesis processing using the images of the fragments after correction, and generates a secondary synthesized image.
[0199] FIG. 23 is a diagram showing an example of generating a secondary composite image by performing distortion correction.
[0200] Fig. 23 shows an example in which a secondary composite image CI2 is generated by performing distortion correction on the first image fragment IF1 of the three image fragments IF1 to IF3 shown in Fig. 21. In Fig. 23, image IF1+ is the first image fragment after distortion correction.
[0201] In this way, by performing distortion correction on distorted fragmentary images and generating a secondary composite image, it is possible to generate a high-quality secondary composite image.
[0202] [Generation of a juxtaposed composite image (1)] In the above embodiment, the images constituting the image group are arranged side by side without overlapping to generate a juxtaposed composite image (see FIGS. 17 and 18 ). However, it is also possible to arrange the images so that they overlap according to a predetermined arrangement rule to generate a single juxtaposed composite image. For example, it is also possible to arrange adjacent images vertically and horizontally so that they overlap with each other at a predetermined overlap rate to generate a single juxtaposed composite image. In this case, the overlap rate is set, for example, as follows:
[0203] (1) Set to a predetermined value. In this case, for example, set to the same value as the overlap rate set when shooting. That is, set to the same value as the overlap rate between cameras and the overlap rate in the movement direction. For example, if shooting is performed with both settings set to a 10% overlap rate, the images are arranged so that there is a 10% overlap rate on the top, bottom, left, and right of the image, and a single juxtaposed composite image is generated.
[0204] (2) Allowing the user to arbitrarily set the overlap rate In this case, the overlap rate designated by the user is accepted and a juxtaposed composite image is generated.
[0205] (3) The overlap rate is determined based on information about the photographing position of each image, for example, based on the distance from a reference point or position coordinates.
[0206] (4) The overlap rate is determined so that the size of the juxtaposed composite image to be generated is a target size. For example, the overlap rate is determined so that the number of pixels in the vertical and horizontal directions of the juxtaposed composite image to be generated are the target numbers of pixels. In this case, the target size may be set arbitrarily by the user. Also, as will be described later, when a composite image is generated by dividing the subject into multiple sections, the target size may be set so that the size is the same as the size of the composite image (primary composite image) of the section that was successfully combined. For example, the target size is set so that the size is the same as the size of the primary composite image of the closest section that was successfully combined. Also, for example, the target size is set so that the size is the same as the size of the primary composite image of the adjacent section that was successfully combined.
[0207] [Generation of a juxtaposed composite image (2)] In the above embodiment, a juxtaposed composite image is generated when the distortion of the primary composite image exceeds the allowable range or when the secondary composite image is unacceptable. However, a juxtaposed composite image may be generated when the primary composite image is unacceptable. In this case, if generation of the primary composite image fails (if the primary composite image is unacceptable), generation of the secondary composite image is not performed, and a juxtaposed composite image is directly generated.
[0208] [Generation of Multiple Divided Composite Images] In the case of long infrastructure structures such as tunnel structures and bridges, it is not realistic to represent the entire structure in a single composite image. Therefore, it is preferable to divide the structure into multiple sections and generate a composite image for each section. In particular, in the case of long infrastructure structures such as tunnel structures and bridges, it is preferable to divide the structure into multiple sections along its length and generate a composite image for each section. In this case, for example, the infrastructure structure is divided into multiple sections at regular intervals along its length, and a composite image is generated for each section.
[0209] When an object is divided into multiple sections along its length and a composite image is generated for each section, the photographing operation itself can be completed in one go. That is, images of all sections can be photographed in one photographing operation. In this case, for example, a group of images photographed for each section is extracted from a group of images photographed for all sections, and a composite image (a primary composite image, a secondary composite image, or a juxtaposed composite image) for each section is generated using the extracted group of images. The group of images for each section is extracted, for example, using information on the photographing position.
[0210] In this way, when an object is divided into multiple sections and a composite image is generated for each section, an image is generated in which the composite images generated for each section (primary composite image, secondary composite image, or juxtaposed composite image) are arranged based on the arrangement of the sections to form an overall composite image. For example, when a tunnel structure is divided into multiple sections along its length and a composite image is generated for each section, an image is generated in which the composite images generated for each section (primary composite image, secondary composite image, or juxtaposed composite image) are arranged in a horizontal row to form an overall composite image. In other words, it is a composite image of the entire tunnel structure.
[0211] FIG. 24 is a diagram showing an example of a display of a composite result when a composite image is generated by dividing into a plurality of sections.
[0212] Figure 24 shows an example in which a tunnel structure is divided into a plurality of sections (10 sections) at regular intervals along the longitudinal direction, and composite images SI1 to SI10 are generated for each section. In this case, as shown in Figure 24, an image SI is generated in which the composite images SI1 to SI10 generated for each section are arranged in a horizontal row. The generated image SI is displayed on the display device 116 as an image of the entire tunnel structure (overall image).
[0213] The composite images SI1 to SI10 are composed of a primary composite image, a secondary composite image, or a juxtaposed composite image. Therefore, even if a primary composite image cannot be generated for all sections, at least a secondary composite image or a juxtaposed composite image is displayed. Therefore, even if a normal composite image (primary composite image) cannot be generated due to a missing image in some sections, at least a secondary composite image or a juxtaposed composite image is displayed. This makes it possible to confirm the entire infrastructure in an image, even for long infrastructure structures such as tunnel structures and bridges.
[0214] In this example, the image (whole image) SI generated by arranging the composite images SI1 to SI10 in a horizontal row is an example of a first output image.
[0215] It is preferable that the entire image SI can be enlarged or reduced in size in response to an instruction from the user.
[0216] FIG. 25 is a diagram showing an example of enlarging and displaying the entire image.
[0217] The enlargement or reduction instruction is given by, for example, specifying the center point of the enlargement on the entire image SI and inputting the magnification, or by specifying the center point of the enlargement on the entire image SI and operating the mouse wheel.
[0218] When the image is enlarged, it is preferable to display a reduced version of the entire image SI as a map image MI on the screen, as shown in FIG. 25, so that the enlarged area of the image can be distinguished within the map image MI.
[0219] [Damage Detection] Damage (cracks, peeling, corrosion, etc.) that appears on the surface of the object may be automatically detected from the image, and the results may be displayed together with the composite image.
[0220] Damage detection may be performed on the images before compositing (the individual images that make up the original image group), or on the image after compositing (composite image). Composite images include primary composite images, secondary composite images, and juxtaposed composite images.
[0221] Techniques for detecting damage from images are well known, and detailed explanations thereof will be omitted. Generally, damage to the surface of an object is detected by analyzing photographed images of the object. If the damage is chalking, chalk lines (for example, lines drawn with chalk along cracks) are also included in the detection targets.
[0222] Damage detection through image analysis includes damage detection using so-called artificial intelligence. For example, damage detection can be performed using a trained model that has been trained to detect damage from images.
[0223] FIG. 26 is a diagram showing an example of a display of the damage detection results.
[0224] An image IX is generated in which the damage detection result is superimposed on the composite image and output to the display device 116. Fig. 26 shows an example in which a crack is detected as the damage. In this case, an image IX is generated in which a line tracing the detected crack is superimposed on the composite image as the damage detection result.
[0225] The display of the damage detection results may be optionally turned on or off. For example, the damage detection results may be switched between being superimposed on the composite image (display on) and not being displayed (display off) in response to a user instruction.
[0226] Furthermore, when multiple types of damage are detected, the display may be switched for each type of damage.
[0227] Furthermore, the display may be changed depending on the size of the damage so that the size of the damage can be distinguished on the screen. For example, cracks may be displayed in different colors and / or different line types depending on the width of the crack. This allows the size (width) of each crack to be grasped at a glance on the screen.
[0228] In this example, image IX is an example of a second output image.
[0229] [Photographing] In the above embodiment, an example has been described in which an object is photographed using a camera array imaging device equipped with multiple cameras and a composite image is generated, but the method of photographing the object is not limited to this. For example, an object may be photographed in parts using a single camera to obtain a group of images to be composited.
[0230] FIG. 27 is a diagram showing an example of photographing an object with one camera.
[0231] Fig. 27 shows an example of dividing and photographing one area (rectangular area) Ob of a plane with one camera. The symbol Sa in the figure indicates the camera's photographing area Sa. As shown in Fig. 27, photographing is performed while changing the position so that the photographing areas Sa overlap at a predetermined overlap rate (for example, 10% or more) in the vertical direction (y direction), which is the direction of the short side of the rectangle, and in the horizontal direction (x direction), which is the direction of the long side of the rectangle. Fig. 27 shows an example of photographing each vertical column in order.
[0232] The information on the photographing position is obtained, for example, using GPS, etc. Information on the photographing order in the vertical direction (y direction) and horizontal direction (x direction) can also be obtained and used as the photographing position information.
[0233] The photography does not necessarily have to be performed by a person, but may be performed by a mobile robot (self-propelled photography robot). For example, photography may be performed by a drone (unmanned aerial vehicle) equipped with a camera. When photography is performed by a drone, it may be configured to move according to a preset route and take pictures automatically.
[0234] [Photography System] In the above embodiment, the control device 100 is equipped with an image processing function such as generating a composite image, but the image processing function may be realized by a device separate from the control device 100. For example, the image processing function may be provided in a computer (server) on a network. In this case, a group of images to be processed is transmitted to the computer on the network, and image processing such as composition processing is performed by the computer on the network.
[0235] [Hardware Configuration] The hardware that realizes the image processing device of the present invention can be configured with various processors. These processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes programs and functions as various processing units; a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacture; and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration specifically designed to perform specific processing. A processing unit constituting the inspection support device may be configured with one of the above-mentioned various processors, or may be configured with two or more processors of the same or different types. For example, a processing unit may be configured with multiple FPGAs or a combination of a CPU and an FPGA. Alternatively, multiple processing units may be configured with a single processor. A first example of configuring multiple processing units with a single processor is a configuration in which a single processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server computers, and this processor functions as multiple processing units. Second, there is a form in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a System on Chip (SoC). In this way, various processing units are configured as a hardware structure using one or more of the above-mentioned various processors. Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit (circuitry) that combines circuit elements such as semiconductor elements.
[0236] 1...Photographing system 10...Multi-eye photography device 11...Frame 12...Base plate 13...Column 14...Mounting base 20...Relay device 100...Control device 111...CPU 111A...Camera control unit 111A1...Photographing control unit 111A2...Photographed image acquisition unit 111A3...Photographed image display control unit 111A4...Photographed image recording control unit 111B...Lighting control unit 111C...Processing target image acquisition unit 111D...First synthesis processing unit 111E...First pass / fail judgment unit 111F...Second synthesis processing unit 111G...Second pass / fail judgment unit 111H...Third synthesis processing unit 111J...Synthesized image recording control unit 111K...Synthesized image display control unit 111L...Preprocessing unit 111L1...Distorted image extraction unit 111L2...Distortion correction unit 112...ROM 113...RAM 114...Auxiliary storage device 115...Input device 116...Display device 117...Communication interface 120...Image database Ax...Axis B1...Bracket (first bracket) B2...Bracket (second bracket) B3...Bracket (third bracket) B4...Bracket (fourth bracket) B5...Bracket (fifth bracket) C1...Camera (first camera) C2...Camera (second camera) C3...Camera (third camera) C4...Camera (fourth camera) C5...Camera (fifth camera) CI2...Composite image (secondary composite image) CI3...Composite image (juxtaposed composite image) CL...Clamp IC3...Composite image IDA1...Image display area (first image display area) IDA2...Image display area (second image display area) IDA3...Image display area (third image display area) IDA4...Image display area (fourth image display area) IDA5...Image display area (fifth image display area) IF1...Image (first fragment image) IF2...Image (second fragment image) IF3...Image (third fragment image) IX...Image (image in which damage detection results are superimposed on a composite image) L1...Lighting device (first lighting device) L2...Lighting device (second lighting device) L3...Lighting device (third lighting device) L4...Lighting device (fourth lighting device) L5...Lighting device (fifth lighting device) MI...Map image SI...Image (whole image) SI1...Composite image SI2...Composite image SI3...Composite image SI4...Composite image SI5...Composite image SI6...Composite image SI7...Composite image SI8...Composite image SI9...Composite image SI10...Composite image Sa...Photographing area TS...Tunnel structureTr... Cart S11 to S19... Composite image generation process procedure
Claims
1. An image processing device comprising a processor that acquires information on an image group and the shooting positions of the images that constitute the image group, performs a first synthesis process on the image group to generate a first composite image, determines whether the quality of the first composite image satisfies a first criterion, and if the quality of the first composite image does not satisfy the first criterion, performs a second synthesis process on the image group or the first composite image based on the information on the shooting positions to generate a second composite image.
2. The image processing device according to claim 1, wherein the processor performs, as the first synthesis process, a process of performing feature point matching between the images constituting the image group and generating the first synthetic image from the result of the feature point matching.
3. The image processing device according to claim 2, wherein the processor determines a degree of fragmentation of the first composite image to determine whether or not the quality of the first composite image satisfies the first criterion.
4. The image processing device described in claim 3, wherein the processor performs, as the second synthesis process, a process of arranging at least some of the multiple fragments of the first composite image on a fragment different from the fragments based on the information on the shooting position and synthesizing the fragments, and generates, as the second composite image, an image that is less fragmented than the first composite image.
5. The image processing device according to claim 4, wherein the processor further performs a process of extracting fragments having distortion exceeding a predetermined range from among the plurality of fragments before performing the second synthesis process, and correcting the distortion of the extracted fragments having distortion.
6. An image processing device as described in any one of claims 1 to 5, wherein the processor determines whether the quality of the second composite image satisfies a second criterion, and if the quality of the second composite image does not satisfy the second criterion, generates as the second composite image an image in which the images constituting the image group are arranged based on information on the shooting positions.
7. The image processing device according to claim 2, wherein the processor determines whether or not the quality of the first composite image satisfies the first criterion by determining whether or not the degree of distortion of at least the first composite image is within an acceptable range.
8. The image processing device according to claim 7, wherein the processor performs, as the second synthesis process, a process of arranging the images constituting the image group based on the information on the shooting positions to generate the second synthetic image.
9. An image processing device as described in claim 7 or 8, wherein the processor measures the maximum and minimum values of the width in the first direction of the first composite image, calculates the difference between the maximum and minimum values of the width in the first direction, and determines whether the difference is equal to or smaller than a threshold value to determine whether the degree of distortion of the first composite image is within an acceptable range.
10. An image processing device as described in claim 7 or 8, wherein the processor acquires information on the shooting conditions of the images that constitute the image group, estimates the width in the second direction of the first composite image based on the information on the shooting conditions, measures the width in the second direction of the first composite image, calculates the difference between the estimated value and the measured value of the width in the second direction, and determines whether the difference is equal to or smaller than a threshold value to determine whether the degree of distortion of the first composite image is within an acceptable range.
11. The image processing device according to claim 7 or 8, wherein the processor calculates the degree of deviation of synthesis parameters determined from the results of the feature point matching between adjacent images, determines whether there is an image for which the degree of deviation is equal to or greater than a threshold, and determines whether the degree of distortion of the first synthetic image is within an acceptable range.
12. The image processing device according to claim 2, 3, 4, 5, 7 or 8, wherein the processor performs the feature point matching between adjacent images based on the information on the shooting positions.
13. An image processing device as described in claim 1, 2, 3, 4, 5, 7 or 8, wherein the group of images is composed of images taken while changing positions using an imaging device equipped with multiple cameras, and the information on the imaging positions is composed of information on the placement positions of the cameras in the imaging device and information on the positions where the images were taken by the imaging device.
14. An image processing device as described in claim 1, 2, 3, 4, 5, 7 or 8, wherein the processor acquires, for an object divided into a plurality of sections in the longitudinal direction, information on the group of images and the shooting positions of the images constituting the group of images for each section, generates the first composite image or the second composite image for each section, and generates a first output image in which the first composite image or the second composite image generated for each section is arranged based on the arrangement of the sections.
15. An image processing device as described in claim 1, 2, 3, 4, 5, 7 or 8, wherein the processor analyzes the images constituting the image group, the first composite image or the second composite image to detect damage to the surface of the object, and generates a second output image in which the damage detection result is superimposed on the first composite image or the second composite image.
16. An image processing method comprising: acquiring information on an image group and the shooting positions of images constituting the image group; performing a first synthesis process on the image group to generate a first composite image; determining whether the quality of the first composite image satisfies a first criterion; and, if the quality of the first composite image does not satisfy the first criterion, performing a second synthesis process based on the shooting position information on the image group or the first composite image to generate a second composite image.
17. An image processing program that causes a computer to realize the following functions: a function for acquiring information on an image group and the shooting positions of images that constitute the image group; a function for performing a first synthesis process on the image group to generate a first composite image; a function for determining whether the quality of the first composite image satisfies a first criterion; and a function for performing a second synthesis process on the image group or the first composite image based on the information on the shooting positions to generate a second composite image if the quality of the first composite image does not satisfy the first criterion.
18. A non-transitory computer-readable recording medium having the program according to claim 17 recorded thereon.