Information processing system, information processing method, and program

The information processing system enhances imaging by using segmentation masks to automatically adjust focus and exposure, addressing manual adjustment issues and improving accuracy for objects like power lines and steel towers.

JP2026011633AActive Publication Date: 2026-01-23SENSYN ROBOTICS INC +2
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024112399
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2026-01-23
Estimated Expiration
2044-07-12

AI Technical Summary

Technical Problem

Existing imaging systems struggle with manual adjustments of camera focus and exposure, leading to time-consuming and error-prone operations, especially when photographing objects like power lines and steel towers, and fail to accurately focus on objects in backlit environments.

Method used

An information processing system that uses object detection through segmentation to generate a segmentation mask, which aligns the imaging device with the object in the field of view, and applies expansion and weighting processes to improve autofocus and autoexposure accuracy.

Benefits of technology

Automatically sets imaging conditions to accurately focus on objects, reducing adjustment errors and improving focus and exposure accuracy, even in challenging environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026011633000001_ABST
    Figure 2026011633000001_ABST
Patent Text Reader

Abstract

To provide an information processing system, an information processing method, and a program capable of automatically setting a photographing condition for a specific object.SOLUTION: An information processing system comprising: an imaging data acquisition unit configured to acquire a captured image including a target object; a target object detection unit configured to specify the target object in the captured image by segmentation; a mask generation unit configured to generate a segmentation mask for segmenting the target object in the captured image on the basis of a detection result of the segmentation; and an imaging control unit configured to determine an imaging condition by aiming an imaging device at the target object in an imaging field of view using the segmentation mask.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing system, an information processing method, and a program related to setting of imaging conditions. [Background technology]

[0002] As shown in Patent Document 1, a system is known that photographs a predetermined object using an imaging device for inspection. Furthermore, Patent Document 2 discloses a method for setting camera parameters that can be used in an inspection system such as that shown in Patent Document 1. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-089160 [Patent Document 2] Special Publication No. 2018-534599 Summary of the Invention [Problem to be solved by the invention]

[0004] In the above-mentioned Patent Document 2, changes in the spatial relationship between the imaging device and the object to be photographed are determined and camera parameters are automatically corrected. However, even when using the automatic parameter setting method disclosed in Patent Document 2, there are cases where the object to be photographed is not sufficiently focused, and the operator performing the photographing must manually adjust the camera focus, exposure, etc. Such manual parameter adjustment takes time and has the problem of adjustment errors and personal errors.

[0005] An object of the exemplary embodiments of the present disclosure is to provide an information processing system, an information processing method, and a program that can automatically set shooting conditions for a specific object. [Means for solving the problem]

[0006] An information processing system according to one embodiment of the present disclosure includes: an imaging data acquisition unit that acquires a captured image including the object; an object detection unit that identifies the object in the captured image by segmentation; a mask generation unit that generates a segmentation mask that divides the object in the captured image based on the segmentation detection result; and an imaging condition setting unit that uses the segmentation mask to aim the imaging device at the object within the imaging field of view and determines imaging conditions.

[0007] The information processing system has the above-described features, so that the photographing conditions for the object can be automatically set. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of an information processing system according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram showing a hardware configuration of the user terminal shown in FIG. [Figure 3] FIG. 3 is a block diagram illustrating a hardware configuration of the management server illustrated in FIG. [Figure 4] FIG. 4 is a block diagram showing the hardware configuration of the aircraft shown in FIG. [Figure 5] FIG. 5 is a block diagram illustrating the functionality of a user terminal. [Figure 6] FIG. 6 is a conceptual diagram showing an example of operation of the information processing system shown in FIG. [Figure 7] FIG. 7 is a flowchart showing an example of information processing related to automatic adjustment of imaging conditions. [Figure 8] FIG. 8 is a flowchart showing an embodiment of steps SQ103 to SQ104 shown in FIG. [Figure 9] FIG. 9 is a flowchart showing another embodiment of steps SQ103 to SQ104 shown in FIG. [Figure 10] FIG. 10 is a diagram showing an example of a captured image. [Figure 11] FIG. 11 is a conceptual diagram for explaining segmentation of an object. [Figure 12] FIG. 12 is a diagram for explaining a superimposed image to which a segmentation mask is applied. [Figure 13] FIG. 13 is a diagram for explaining the process of removing small segments. [Figure 14] FIG. 14 is a diagram for explaining the process of expanding the target region of the segmentation mask. [Figure 15] FIG. 15 is a diagram for explaining weighting of the target region of the segmentation mask. DETAILED DESCRIPTION OF THE INVENTION

[0009] The information processing method, information processing system, and program of the present disclosure have, for example, the following configuration. [Item 1] an imaging data acquisition unit that acquires a captured image including the object; an object detection unit that identifies the object in the captured image by segmentation; a mask generation unit that generates a segmentation mask that divides the object in the captured image based on the segmentation detection result; an imaging condition setting unit that uses the segmentation mask to aim the imaging device at the object within the imaging field of view and determines imaging conditions. [Item 2] the segmentation mask has a target region corresponding to a detection region identified as the target by the segmentation, and a covering region covering a region other than the target in the captured image, Item 1. The information processing system according to item 1, wherein the mask generation unit sets the target region by expanding a dimension in a predetermined direction of the detection region identified by the segmentation so that at least a portion of the contour of the object is included within the target region. [Item 3] the segmentation mask has a target region corresponding to the detection region identified as the target by the segmentation; Item 2. The information processing system according to item 1, wherein the mask generation unit sets weights for a part or all of the target region and generates a weighted segmentation mask. [Item 4] Item 3. The information processing system according to item 2, wherein the mask generation unit sets weights for part or all of the target area set by expanding the dimensions of the detection area, and generates a weighted segmentation mask. [Item 5] The information processing system according to any one of items 2 to 4, wherein the shooting condition setting unit generates a superimposed image by superimposing the segmentation mask on the captured image, and determines the shooting conditions related to exposure of the shooting device based on the luminance of each pixel included in the target area of ​​the superimposed image. [Item 6] The information processing system described in item 2, wherein the shooting condition setting unit generates a superimposed image by superimposing the segmentation mask on the captured image, and determines the shooting conditions related to the focus of the shooting device based on the contour of the object included in the target area of ​​the superimposed image. [Item 7] 5. The information processing system according to item 3 or 4, wherein the shooting condition setting unit generates a superimposed image by superimposing the segmentation mask on the captured image, and determines the shooting condition related to the focus of the shooting device based on a weight assigned to the target region of the superimposed image. [Item 8] Obtaining a photographed image including the object; Identifying the object in the captured image by segmentation; generating a segmentation mask that segments the object in the captured image based on the segmentation detection result; and using the segmentation mask to aim the imaging device at the object within the imaging field of view and determine imaging conditions. [Item 9] A process of acquiring a photographed image including the object; a process of identifying the object in the captured image by segmentation; generating a segmentation mask that segments the object in the captured image based on the segmentation detection result; and a process of using the segmentation mask to aim the imaging device at the object within the imaging field of view and determine imaging conditions.

[0010] <Details of implementation form> An information processing system according to an embodiment of the present disclosure will be described with reference to the drawings. In the accompanying drawings, identical or similar elements are designated by identical or similar reference symbols and names, and duplicate descriptions of identical or similar elements may be omitted in the description of the embodiment. Note that the contents shown in the drawings are merely examples for explaining the present embodiment and are merely schematic examples for ease of explanation of the present embodiment. The contents of the drawings may be modified or changed within the scope of no technical problem.

[0011] <System Overview> The information processing system according to this embodiment focuses an image capture device on a predetermined object and automatically adjusts the image capture conditions. The object captured by this information processing system is not particularly limited. The method for automatically adjusting image capture conditions executed by the information processing system is particularly effective when capturing an object whose texture (pattern) is not easily visible in the captured image (e.g., overhead power lines, steel towers, pipes, etc.). This method can be utilized when capturing images of such objects using a camera mounted on an unmanned mobile vehicle, such as a drone, an unmanned aerial vehicle (UAV), or an unmanned ground vehicle (UGV). This embodiment describes in detail the information processing involved in the automatic adjustment of image capture conditions, using an example of capturing images of power transmission facilities (particularly power lines) using an unmanned aerial vehicle (FIGS. 1 and 6). However, the systems shown in FIGS. 1 and 6 are merely examples, and the image capture device and image capture method disclosed herein are not necessarily limited. For example, a person such as a worker may photograph the object by operating an imaging device (e.g., a smartphone, tablet terminal, digital camera, or other camera), or the object may be photographed by a camera mounted on a manned moving object.

[0012] As shown in Figure 6, in the information processing system exemplified below, an unmanned aerial vehicle 4 is flown along a plurality of power lines extending side by side and attached to a support (such as a steel tower), and the power lines, cross arms supporting the power lines, span spacers that maintain a constant distance between the power lines, and other power transmission equipment such as steel towers are photographed by a camera 42 mounted on the unmanned aerial vehicle 4. The unmanned aerial vehicle 4 may be flown autonomously or remotely controlled based on instructions from a user terminal 2 owned by a user.

[0013] Inspection systems using such unmanned aerial vehicles (UAVs) 4 have attempted to utilize an autofocus function (automatic focus adjustment function) and / or an autoexposure function (automatic exposure adjustment function). However, with these conventional functions, even when attempting to focus the camera on the power line or other object being inspected, the camera would instead focus on the background, making it difficult to focus on the linear power line. The same is true when inspecting a steel tower with a complex arrangement of steel frames. Furthermore, conventional technologies have had the problem of the autofocus and / or autoexposure functions not working properly in backlit shooting environments. For this reason, in the past, users would adjust the camera's focus, exposure, etc. by viewing the image captured by the camera displayed on the user terminal 2 and clicking on the object (e.g., power line, steel tower) to specify it on the image. Such manual parameter adjustments have been problematic in that they require time for adjustment and are prone to adjustment errors and personal errors.

[0014] In the information processing system of this embodiment, objects (such as power lines and steel towers) are extracted from a captured image by segmentation, and a segmentation mask is generated to distinguish the extracted objects. The segmentation mask, for example, covers the background except for the objects, thereby revealing the extracted object area. The information processing system applies such a segmentation mask to the captured image, and determines the shooting conditions by aligning the aim of the image capture device (camera 42 of the unmanned aerial vehicle 4) with the object in the field of view. This prevents the image capture device from focusing on the background, even in backlit shooting environments, such as when the object is a linear power line or a steel tower with steel frames. In particular, by applying a detection area (segment) expansion process and / or a weighting process (described later) to generate a segmentation mask, the success rate of AF (autofocus) and / or AE (autoexposure) in the information processing system can be further improved.

[0015] <System configuration> As shown in FIG. 1, the information processing system of this embodiment may include a management server 1, one or more user terminals 2, and one or more unmanned aerial vehicles 4. The management server 1, the user terminal 2, and the unmanned aerial vehicles 4 are communicatively connected to each other via a network NW. Note that the illustrated configuration is an example and is not limited to this. For example, the unmanned aerial vehicle 4 may not be connected to the network NW. In this case, the unmanned aerial vehicle 4 may be operated by a transmitter (so-called radio) operated by a user, or image data acquired by a camera of the unmanned aerial vehicle 4 may be stored in an auxiliary storage device (e.g., a memory card such as an SD card and / or a USB memory) connected to the unmanned aerial vehicle 4 and subsequently read and stored by the user from the auxiliary storage device in the user terminal 2 and / or the management server 1. Alternatively, the unmanned aerial vehicle 4 may be connected to the network NW solely for the purpose of either operation or storage of image data.

[0016] <User device 2> 2 is a diagram showing the hardware configuration of the user terminal 2. Note that the configuration shown in the figure is an example, and the user terminal 2 may have a different configuration.

[0017] The user terminal 2 includes at least a processor 20, a memory 21, a storage 22, a transceiver 23, an input / output unit 24, etc., which are electrically connected to each other via a bus 25. The user terminal 2 may be, for example, a general-purpose computer such as a workstation or a personal computer.

[0018] The processor 20 is a computing device that controls the overall operation of the user terminal 2, controls the transmission and reception of data between each element, and performs information processing necessary for application execution and authentication processing. For example, the processor 20 is a CPU (Central Processing Unit) and / or GPU (Graphics Processing Unit), and executes programs stored in the storage 22 and deployed in the memory 21 to perform various information processing.

[0019] The memory 21 includes a main memory configured with a volatile storage device such as a DRAM (Dynamic Random Access Memory), and an auxiliary memory configured with a non-volatile storage device such as a flash memory, an HDD (Hard Disc Drive), etc. The memory 21 is used as a work area for the processor 10, and also stores a BIOS (Basic Input / Output System) that is executed when the management server 1 is started, various setting information, etc.

[0020] The storage 22 stores various programs such as application programs. A database storing data used for each process may be constructed in the storage 22. For example, a storage unit 210 (described later) may be provided in a part of the storage area of ​​the storage 22.

[0021] The transmitter / receiver 23 is a communication interface that enables the user terminal 2 to communicate with the management server 1, the unmanned aerial vehicle 4, etc. via a communication network. The transmitter / receiver 23 may further include a short-range communication interface such as Bluetooth (registered trademark) and BLE (Bluetooth Low Energy) and / or a USB (Universal Serial Bus) terminal.

[0022] The input / output unit 24 is an information input device such as a keyboard and a mouse, and an output device such as a display.

[0023] A bus 25 is commonly connected to the above elements and transmits, for example, address signals, data signals and various control signals.

[0024] <Management Server 1> The management server 1 shown in Figure 3 also includes a processor 10, memory 11, storage 12, a transmission / reception unit 13, an input / output unit 14, etc., which are electrically connected to each other via a bus 15. The management server 1 may be a general-purpose computer such as a workstation or personal computer, or may be logically realized by cloud computing. The functions of each element of the management server 1 can be configured in the same way as the above-mentioned user terminal 2, and detailed description of each element of the management server 1 will be omitted.

[0025] <Unmanned Aerial Vehicle 4> 4 is a block diagram showing the hardware configuration of the unmanned aerial vehicle 4. The flight controller 41 may have one or more processors, such as a programmable processor (for example, a central processing unit (CPU)).

[0026] The flight controller 41 may also have or have access to memory 411. The memory 411 stores logic, code, and / or program instructions that the flight controller can execute to perform one or more steps. The flight controller 41 may also include sensors 412, such as inertial sensors (acceleration sensors, gyro sensors), GPS sensors, and proximity sensors (e.g., lidar).

[0027] The memory 411 may include, for example, a separable medium such as an SD card and random access memory (RAM) or an external storage device. Data acquired from the camera / sensors 42 may be directly transmitted to and stored in the memory 411. For example, still image and video data captured by a camera or the like may be recorded in an internal memory or an external memory, but is not limited to this. The data may also be recorded in at least one of the management server 1 and the user terminal 2 from the camera / sensor 42 or the internal memory via the network NW. The camera 42 is installed on the aircraft 4 via a gimbal 43.

[0028] The flight controller 41 includes a control module (not shown) configured to control the state of the air vehicle 4. For example, the control module has six degrees of freedom (translational motion x, y, and z, and rotational motion θ x , θ y and θ z ) and controls the propulsion mechanism (motor 45, etc.) of the aircraft via an ESC 44 (Electric Speed ​​Controller) to adjust the spatial arrangement, speed, and / or acceleration of the aircraft. The motor 45, powered by a battery 48, rotates a propeller 46, generating lift for the aircraft. The control module of the flight controller 41 may have the function of controlling the driving of the camera 42 and the shooting conditions (e.g., sharpness, focus (focal length), exposure value, shutter speed, ISO sensitivity, lens aperture, etc.). The control module can also control one or more of the states of the onboard components and sensors.

[0029] The flight controller 41 can communicate with a transceiver 47 configured to transmit and / or receive data from one or more external devices (e.g., a transceiver 49, a terminal, a display device, or other remote control). The transceiver 49 may use any suitable communication means, such as wired or wireless communication.

[0030] For example, the transceiver 47 may utilize one or more of a local area network (LAN), a wide area network (WAN), infrared, radio, WiFi, a point-to-point (P2P) network, a telecommunications network, cloud communications, and the like.

[0031] The transceiver unit 47 can transmit and / or receive one or more of the following: data acquired by the cameras / sensors 42, processing results generated by the flight controller 41, predetermined control data, user commands from a terminal or remote controller, etc.

[0032] The cameras / sensors 42 may include inertial sensors (acceleration sensors, gyro sensors), GPS sensors, proximity sensors (e.g., lidar), or vision / image sensors (e.g., cameras).

[0033] <User device 2 functions> FIG. 5 is a block diagram illustrating functions implemented in the user terminal 2. In this embodiment, the user terminal 2 may include an imaging data acquisition unit 201, an object detection unit 202, a mask generation unit 203, and an imaging condition setting unit 204. The storage unit 210 of the user terminal 2 may include various databases such as an information / image storage unit 211. Note that the various functional units illustrated in FIG. 5 are illustrated as functional units in the processor 20 of the user terminal 2, but some or all of the various functional units may be implemented in either the processor 20 or the flight controller 41, depending on the capabilities of the processor 20 of the user terminal 2 or the flight controller 41 of the unmanned aerial vehicle 4.

[0034] The imaging data acquisition unit 201 acquires captured images (images of objects such as power lines captured by the camera 42) captured by the camera 42 mounted on the unmanned air vehicle 4, for example, by wireless communication via a communication interface. Image 50 shown in FIG. 10 is a simulated example of a captured image, and the image 50 shows a power line captured by the camera 42 as an object 60. The imaging data acquisition unit 201 is preferably configured to acquire captured images in real time via wireless communication from the unmanned air vehicle 4 or within the unmanned air vehicle 4. Note that the captured images acquired by the imaging data acquisition unit 201 may be moving images or still images. If the captured images are moving images, the moving images may be divided into still images for each frame, and the still images may be used in each functional unit (202-204) described below. Furthermore, still images may be extracted at predetermined intervals from the still images divided into frames, and the extracted still images may be used in each functional unit. The images captured by the unmanned aerial vehicle 4 or the like may be color images, grayscale images, or black and white images.

[0035] The object detection unit 202 receives the captured image acquired by the imaging data acquisition unit 201 as input information and performs processing to identify the object 60 in the captured image by segmentation. Segmentation is a technique for dividing the captured image into regions according to the type of object depicted in the captured image by determining to which segment (e.g., category such as power line, steel tower, background, etc.) each pixel constituting the captured image belongs and labeling each pixel. The object detection unit 202 may be constructed using an AI model that has undergone machine learning (e.g., deep learning) to extract a predetermined object (e.g., power line, steel tower, etc.) from the captured image. The architecture of the segmentation performed by the object detection unit 202 is not particularly limited, and for example, LR-ASPP (Lite Reduced Atrous Spatial Pyramid Pooling) may be used.

[0036] For example, when extracting a power line as an object 60 from a captured image (image 50), the object detection unit 202 performs segmentation to divide the captured image into at least an area showing the power line (object segment) and an area showing other objects (such as the background). This segmentation not only identifies the segment corresponding to the power line, but may also extract segments representing other specific objects, such as a span spacer, a cross arm, or a steel tower. The area identified as an object through segmentation is referred to as a detection area. FIG. 11 is a diagram illustrating the results of segmentation on image 50 of FIG. 10. The area indicated by the dashed dotted line in FIG. 11 represents a detection area 61 identified as a power line. In segmentation, objects are generally extracted so that the boundary line (contour) of the detection area is located on and / or inside the contour of the object. In FIG. 11, for ease of illustration of the detection area 61, the boundary line of the detection area 61 is shown positioned on the inner edge of the outline of the power line.

[0037] The mask generation unit 203 generates a segmentation mask for segmenting the object in the captured image based on the detection result (extraction result) of the segmentation by the object detection unit 202. Here, "segmenting the object in the captured image" means making the detection area 61 identified as the object in the captured image identifiable. For example, the segmentation mask may be a mask that makes the detection area 61 visible and conceals (conceals) areas other than the object, or a mask in which each segment divided by segmentation, such as the detection area 61, is colored (color-coded) in a different color.

[0038] A mask 70 shown in (a) of Fig. 12 is an example of a segmentation mask generated based on the segmentation detection result shown in Fig. 11, and the mask 70 has a target region 71 corresponding to the detection region 61 identified as the target object 60 by segmentation, and a covering region 72 covering the region other than the target object in the captured image. (b) of Fig. 12 is an image (hereinafter referred to as a superimposed image 52) in which the mask 70 in (a) is superimposed on the captured image 50 in Fig. 10, and in the superimposed image 52, the target object 60 is made visible in the target region 71 corresponding to the detection region 61, and the background other than the target object 60 is obscured in the covering region 72.

[0039] The mask generation unit 203 may perform a small segment removal process as a preprocessing step when generating a segmentation mask such as that shown in FIG. 12. In segmentation, as shown in FIG. 13(a), pixels of a similar color to the object or an object of a similar shape to the object (e.g., a linear object captured with a similar thickness as a power line) may be erroneously detected, resulting in the extraction of a small segment 63 that is different from the object 60. Furthermore, if the power lines captured in the captured image are closely spaced (e.g., power lines 60a and 60b shown in FIG. 13(a)), a region 64 between the power lines may be extracted as a region that may represent a linear structure. In the preprocessing step when generating the mask, a process may be performed to remove the erroneously detected small segments 63 and 64, as shown in FIG. 13(a). The method of removing the small segments performed by the mask generation unit 203 as a preprocessing step is not necessarily limited. For example, the mask generation unit 203 may perform an opening process.

[0040] In the opening process, each segment extracted by segmentation (excluding the covered region 72 determined to be the background) is eroded once, and then dilated again by the same amount as the erosion amount, thereby removing small segments. The erosion may be performed multiple times in succession on a pixel-by-pixel basis, and the erosion eliminates small segments 63 and 64 that have been erroneously detected (e.g., FIG. 13(b)). In the dilation process, each eroded segment is dilated on a pixel-by-pixel basis the same number of times as the erosion was performed, returning the detection region 61 to its original size (e.g., FIG. 13(c)). This opening process makes it possible to remove unnecessary small segments 63 and 64 without significantly changing the size of the detection region 61, thereby further improving the accuracy of focusing on the object 60, which will be described later.

[0041] When setting the target region 71 in the segmentation mask 70, the mask generation unit 203 may perform a process (hereinafter referred to as the expansion process) of expanding the detection region 61 identified as the target 60 by segmentation so that at least a portion of the contour (particularly the outer edge) of the target 60 is included inside the target region 71. As described above, when the target 60 is extracted by high-precision segmentation, there is a high probability that the outer edge of the extracted detection region 61 will overlap with the contour line of the target 60 or be located inside the contour line. In the expansion process, when generating the segmentation mask 70, the dimensions of the detection region 61 identified by segmentation are expanded (for example, as shown in FIG. 14(a)), and the expanded region is set as the target region 71. Through this expansion process, the outer edge portion of the contour where the target 60 transitions to the background in the captured image is included inside the target region 71, which is the manifestation region of the mask 70 (as shown in FIG. 14(b)). In this way, by including the outer edge of the object 60 within the object area 71, it becomes easier to aim the camera 42 at the object 60 using the outline of the object 60 as a landmark, thereby further improving the accuracy of focusing.

[0042] The detection area 61 is expanded in pixel units, and the degree of expansion (expansion amount; for example, approximately 3 to 10 pixels) is not particularly limited. The expansion amount may be set in advance through tests before operation, or may be arbitrarily set by the user. The direction of expansion is also not particularly limited. The detection area 61 may be expanded in a predetermined direction (for example, a direction perpendicular to the longitudinal direction of the power line) so that at least a portion of the outline of the object 60 is included within the target area 71, or the detection area 61 may be expanded in all directions (i.e., the entire circumference of the detection area 61). Note that if the aforementioned opening process is performed as pre-processing before setting the target area 71 by the expansion process, the mask generation unit 203 may perform the expansion process of the detection area 61 immediately after the expansion step in the opening process.

[0043] The mask generation unit 203 may also set weights for part or all of the target region 71 (hereinafter referred to as weighting processing) to generate a weighted segmentation mask. Here, setting weights for the target region 71 means, for example, multiplying the pixel values ​​(brightness) of part or all of the pixels corresponding to the target region 71 by a predetermined coefficient. In weighting processing, two or more regions with different weights are formed within the target region 71, intentionally creating a texture (pattern) within the target region 71. Objects such as power lines and steel towers captured in a captured image may have almost no texture within the object, and the sparse texture may be a factor that prevents focusing on the object. By weighting the target region 71 of the segmentation mask 70, an intentional texture is formed on the object 60 in the captured image, allowing the camera 42 to aim using this intentionally formed texture as a landmark, thereby further improving focusing accuracy.

[0044] The pattern of weighted areas set in the weighting process (such as the arrangement, size, and combination of multiple areas to which different weights are set) and the weight values ​​are not particularly limited, as long as they are set so as to form some kind of texture within the target area 71. For example, in FIG. 15, (b) shows a mask 70b to which weights have been set for the target area 71 of the mask 70a shown in (a). In the segmentation mask 70b illustrated in (b) of FIG. 15, a higher weight is set for a central area 71a in the width direction of the power line (a direction perpendicular to the longitudinal direction) than for other areas, and a lower weight is set for an outer area 71b located outside the central area 71a and including the outline of the power line than for the central area 71a.

[0045] The mask generation unit 203 may perform a weighting process on the target region 71 corresponding to the unexpanded detection region 61 without performing the above-described expansion process, or may perform a weighting process on the target region 71 after it has been expanded by the expansion process. In other words, the mask generation unit 203 may generate the segmentation mask 70 by applying only either the expansion process shown in Fig. 14 or the weighting process shown in Fig. 15, or may generate the segmentation mask 70 by applying both processes.

[0046] The photographing condition setting unit 204 executes a process of determining (automatically adjusting) the photographing conditions by aligning the aim of the camera 42 with the object 60 in the photographing field of view using the segmentation mask 70 generated by the mask generation unit 203. Specifically, the photographing condition setting unit 204 may generate a superimposed image 52 by superimposing the segmentation mask 70 on the photographed image, and may execute AF adjustment and / or AE adjustment by aligning the aim of the camera 42 with the object 60 that is made apparent in the target area 71 of the superimposed image 52.

[0047] In the AF adjustment, the photographing conditions (focus position) related to the focus of the camera 42 are determined, and the specific method of the AF adjustment is not necessarily limited. For example, the AF adjustment may be performed using a hill climbing algorithm. In the hill climbing algorithm, first, the direction (focus direction) in which the sharpness in the superimposed image 52 increases when the focus position is moved is detected. Then, while moving the focus in the increasing direction, the position at which the sharpness is maximum (the position at which the slope of the sharpness increase curve is zero) is searched for, and the photographing conditions related to the focus are determined (automatically adjusted) using the position at which the sharpness is maximum as the focus position.

[0048] Sharpness may be calculated based on the contour and / or texture (particularly, intentional texture represented by weights assigned to the target region 71) of the object 60 represented within the target region 71 of the superimposed image 52. For example, the shooting condition setting unit 204 calculates edges included in the target region 71 of the superimposed image 52 using a predetermined edge enhancement filter (e.g., a Laplasian filter), and calculates the variance of the edges as sharpness. In this case, if the contour of the object 60 is made apparent within the target region 71 by the above-described extension process, the maximum sharpness point can be easily identified, and the in-focus position can be more accurately determined. Furthermore, if a texture based on weights is represented within the target region 71 by the above-described weighting process, the boundary between regions with different weights is also calculated as an edge, and the maximum sharpness point can be easily identified, and the in-focus position can be more accurately determined. That is, when the mask generation unit 203 executes the expansion process and / or the weighting process to generate the segmentation mask 70, the shooting condition setting unit 204 identifies the in-focus position based on the outline of the object 60 included in the target area 71 and / or the weighting assigned to the target area 71. This can further improve the accuracy of the AF adjustment.

[0049] In the AE adjustment, the shooting conditions related to the exposure of the camera 42 (e.g., one or more of the shutter speed, ISO sensitivity, and lens aperture) are determined. The AE adjustment method is not necessarily limited. For example, the shooting condition setting unit 204 may determine the exposure value based on the average luminance within the target region 71 of the superimposed image 52. In this case, the shooting condition setting unit 204 calculates the average luminance within the target region 71 by summing the luminance of each pixel included in the target region 71 and dividing the sum by the number of pixels included in the target region 71. When a weighted segmentation mask 70 is used, the average luminance may be calculated based on the luminance after the weight is applied, or may be calculated based on the luminance before the weight is applied. The shooting condition setting unit 204 searches for an exposure value that minimizes the difference between the calculated average luminance and a target luminance (a target value previously set to achieve optimal brightness, for example, a fixed value of 170), and determines (automatically adjusts) the shooting conditions related to exposure using the exposure value that minimizes the difference.

[0050] Even in the case of AE adjustment, the mask generation unit 203 performs an expansion process and / or a weighting process to generate a segmentation mask 70, thereby reducing adjustment errors caused by areas other than the subject, such as the background in the captured image, and further improving the accuracy of AE adjustment.

[0051] When adjusting AF and / or AE, the shooting condition setting unit 204 may generate instruction information indicating the camera parameters at the time of adjustment and transmit it to the camera control unit 420 of the unmanned aerial vehicle 4 via the transceiver unit 23, or may transmit the shooting conditions after automatic adjustment determined based on the search results (search results for maximum sharpness and / or search results for exposure value) to the camera control unit 420. The camera control unit 420 of the unmanned aerial vehicle 4 controls the operation of the camera 42 based on the instruction information etc. sent from the shooting condition setting unit 204.

[0052] The automatic adjustment flow of the shooting conditions by each functional unit from the imaging data acquisition unit 201 to the shooting condition setting unit 204 is executed when the unmanned aerial vehicle 4 captures an image including the target object 60 at, for example, the first waypoint on its flight path after starting flight. The automatic adjustment flow of the shooting conditions may be executed at predetermined intervals, such as at each waypoint on its flight path, or may be executed by real-time image analysis of captured images sequentially sent from the unmanned aerial vehicle 4. When executing real-time image analysis, the shooting condition setting unit 204 automatically adjusts the focus of the camera 42 by searching for the maximum value of sharpness within the target area 71 from the captured images input in real time by the imaging data acquisition unit 201. The shooting condition setting unit 204 also constantly calculates the average brightness within the target area 71 from the captured images input in real time by the imaging data acquisition unit 201, and automatically adjusts the exposure value of the camera 42 so that the difference between the average brightness and the target brightness is kept at a minimum.

[0053] The information / image storage unit 211 stores captured images sent from the unmanned aerial vehicle 4 and information associated with the captured images. The information / image storage unit 211 may also store information used by the functional units 201-204. Examples of the information used by the functional units 201-204 include the shooting conditions initially set at the start of flight, setting information related to opening processing (e.g., setting the number of times the reduction process / expansion process is performed), setting information related to expansion processing (e.g., setting the expansion amount and expansion direction), setting information related to weighting processing (settings related to the weight value, weighting pattern, etc.), and information indicating the target brightness.

[0054] <An example of how to automatically adjust the shooting conditions> Next, a method for automatically adjusting the imaging conditions by the information processing system according to this embodiment will be described with reference to FIG.

[0055] First, the imaging data acquisition unit 201 of the user terminal 2 acquires a captured image including an object 60 (e.g., a power line, a steel tower, etc.) captured by the camera 42 mounted on the unmanned aerial vehicle 4, as shown in Fig. 7 (step S101). Next, the object detection unit 202 performs segmentation on the captured image using the captured image acquired by the imaging data acquisition unit 201 as input information (step SQ102). The object detection unit 202 identifies the object 60 in the captured image through this segmentation (Fig. 11).

[0056] Next, the mask generation unit 203 generates a segmentation mask 70 that segments the object 60 in the captured image based on the segmentation detection result by the object detection unit 202 (step SQ103). The shooting condition setting unit 204 uses the generated segmentation mask 70 to aim the camera 42 at the object 60 in the field of view and determine (automatically adjust) the shooting conditions (step SQ104). Specifically, the shooting condition setting unit 204 applies the segmentation mask 70 to the captured image to conceal (conceal) areas other than the object in the captured image and generate a superimposed image 52 that makes the object 60 identified by the segmentation visible ( FIG. 12 ). Then, the shooting condition setting unit 204 aims at the object 60 displayed in the object area 71, which is the visible area, and performs AF adjustment and / or AE adjustment.

[0057] The method shown in FIG. 7 allows the camera 42 to be focused with high precision on a predetermined object in a captured image (in this embodiment, for example, power transmission facilities such as power lines and steel towers).

[0058] From the viewpoint of further improving the accuracy of automatic adjustment such as AF and / or AE, it is preferable to perform expansion processing and / or weighting processing to generate the segmentation mask 70. A more preferable automatic adjustment method will be exemplified below with reference to FIGS.

[0059] Figure 8 is a flowchart illustrating a method for automatically adjusting shooting conditions using extended processing. Steps SQ301 to SQ302 in Figure 8 are steps executed by the mask generation unit 203, corresponding to step SQ103 in Figure 7, and steps SQ401 to SQ403 are steps executed by the shooting condition setting unit 204, corresponding to step SQ104 in Figure 7.

[0060] As shown in FIG. 8 , after identifying the object 60 in the captured image by segmentation, the mask generation unit 203 may perform a process of removing small segments from the segmentation image (step SQ301). Removal of small segments is performed, for example, by an opening process that shrinks and re-expands each segment extracted by segmentation ( FIG. 13 ). The mask generation unit 203 expands the dimensions of the detection region 61 identified as the object 60 in the segmentation in a predetermined direction (step SQ302), either immediately after the expansion step in the opening process or after the opening process, to set the object region 71 in the segmentation mask 70. The width (unit pixel) and expansion direction of the detection region 61 expanded in the expansion process are not necessarily limited. It is sufficient to slightly expand the detection region 61 so that at least a portion of the outline of the object 60 is included within the object region 71, which is the manifestation region ( FIG. 14 ).

[0061] Next, the shooting condition setting unit 204 overlays the segmentation mask 70 having the expanded target region 71 on the captured image to generate the superimposed image 52 (step SQ401, FIG. 14(b)). Step SQ402, which branches off from step SQ401, is a process for performing AE adjustment. In step SQ402, the shooting condition setting unit 204 determines the shooting conditions related to the exposure of the camera 42 based on the brightness of each pixel included in the target region 71 of the superimposed image 52. Specifically, the shooting condition setting unit 204 calculates the average brightness within the target region 71 and adjusts the exposure value so that the difference between the average brightness and the target brightness is minimized. Because the contour of the target object 60 is made apparent within the target region 71 by the expansion process in step SQ302, the exposure value can be controlled to a more appropriate value.

[0062] In step SQ403, which performs AF adjustment, the shooting condition setting unit 204 determines shooting conditions related to the focus of the camera 42 based on the outline of the object 60 included in the object area 71 of the superimposed image 52. Specifically, the shooting condition setting unit 204 calculates the edge corresponding to the outline of the object 60 represented in the object area 71 and calculates the variance of the edge as sharpness. The shooting condition setting unit 204 then identifies the focus position where sharpness is maximized using a hill-climbing method or the like, and adjusts the shooting conditions using the identified position as the in-focus position. Because the outline of the object 60 is made apparent inside the object area 71 by the extension process in step SQ302, it is easy to identify the point of maximum sharpness, and the in-focus position can be identified more accurately.

[0063] Fig. 9 is a flowchart illustrating an example of a method for automatically adjusting shooting conditions to which extension processing and weighting processing are applied. In the automatic adjustment method shown in Fig. 9, the mask generation unit 203 extends the detection area 61 through extension processing (step SQ302) to set a target area 71 of the mask, and then sets weights for part or all of the target area 71 to generate a weighted segmentation mask 70. The weights within the target area 71 are set, for example, so that two or more areas with different weights are formed within the target area 71 and an intentional texture (pattern) is created by these areas.

[0064] 9, in the AE adjustment of step SQ402, the exposure value may be adjusted based on the luminance after weighting, or based on the luminance before weighting. In the former case, the exposure value can be controlled to an appropriate value even for an object 60 (such as a power line or a steel tower) whose texture is essentially invisible in the captured image.

[0065] In the AF adjustment of step SQ404, the shooting conditions related to the focus of the camera 42 are determined based on the weights assigned to the target area 71. Specifically, the shooting condition setting unit 204 calculates edges corresponding to the contours of the object 60 represented in the target area 71 and the texture (boundaries between areas with different weights) due to the weighting. The shooting condition setting unit 204 calculates the variance of the edges based on the contours and weights as sharpness, and then adjusts the focus position of the camera 42 so that the sharpness is maximized. By forming texture in the mask target area 71 through weighting processing, the point of maximum sharpness can be easily identified, allowing for more accurate identification of the focus position. Furthermore, even for an object 60 that would normally have almost no texture in the captured image (e.g., a power line or a steel tower), the focus position for the object 60 can be easily identified.

[0066] Note that Figure 9 illustrates an example of an automatic adjustment process in which a weighting process (step SQ303) is performed after an extension process (step SQ302), but the segmentation mask 70 may also be generated by applying only a weighting process without performing an extension process. Furthermore, the automatic adjustment process of the shooting conditions as shown in Figures 7 to 9 may be performed at the first waypoint of the flight path, and may also be performed continuously in real time using captured images transmitted by the unmanned aerial vehicle 4 during flight as input information.

[0067] The above-described embodiments are merely examples for facilitating understanding of the present disclosure, and are not intended to limit the present disclosure. The present disclosure can be modified or improved without departing from the spirit thereof, and it goes without saying that the present disclosure includes equivalents thereof.

[0068] For example, some or all of the functional units of the imaging data acquisition unit 201, object detection unit 202, mask generation unit 203, and shooting condition setting unit 204 described in the above embodiment may be executed by the flight controller 41 of the unmanned aerial vehicle 4, rather than by the processor 20 of the user terminal 2. Furthermore, while the above embodiment illustrates a system in which an object is inspected by an unmanned mobile vehicle 4, the automatic adjustment process of shooting conditions of the present disclosure may be applied when an object is photographed by a camera mounted on a manned mobile vehicle, or when a person sequentially photographs objects by operating a photographing device such as a smartphone, tablet terminal, or digital camera.

[0069] Furthermore, in the above embodiment, an example was given of photography primarily for the purpose of inspecting power transmission facilities such as power lines and steel towers, but the purpose of the photography is not limited to the example given in the embodiment, and may also be security, monitoring of infrastructure, surveying, disaster response, etc., and the automatic adjustment process of the present disclosure may be applied to photography for these purposes. [Explanation of symbols]

[0070] 1 Management Server 2. User terminal 4 Unmanned Aerial Vehicles

Claims

1. an imaging data acquisition unit that acquires a captured image including the object; an object detection unit that identifies the object in the captured image by segmentation; a mask generation unit that generates a segmentation mask that divides the object in the captured image based on the segmentation detection result; an imaging condition setting unit that uses the segmentation mask to aim the imaging device at the object within the imaging field of view and determines imaging conditions.

2. the segmentation mask has a target region corresponding to a detection region identified as the target by the segmentation, and a covering region covering a region other than the target in the captured image, 2. The information processing system according to claim 1, wherein the mask generation unit sets the target region by expanding a dimension in a predetermined direction of the detection region identified by the segmentation so that at least a partial outline of the object is included within the target region.

3. the segmentation mask has a target region corresponding to the detection region identified as the target by the segmentation; The information processing system according to claim 1 , wherein the mask generation unit sets weights for a part or all of the target region and generates the weighted segmentation mask.

4. The information processing system according to claim 2 , wherein the mask generation unit sets a weight for a part or all of the target region set by expanding the dimensions of the detection region, and generates the weighted segmentation mask.

5. The information processing system according to any one of claims 2 to 4, wherein the shooting condition setting unit generates a superimposed image by superimposing the segmentation mask on the captured image, and determines the shooting conditions related to the exposure of the shooting device based on the brightness of each pixel included in the target area of ​​the superimposed image.

6. The information processing system according to claim 2, wherein the shooting condition setting unit generates a superimposed image by superimposing the segmentation mask on the captured image, and determines the shooting condition related to the focus of the shooting device based on the contour of the object included in the target area of ​​the superimposed image.

7. 5. The information processing system according to claim 3, wherein the imaging condition setting unit generates a superimposed image by superimposing the segmentation mask on the captured image, and determines the imaging condition related to the focus of the imaging device based on a weight assigned to the target region of the superimposed image.

8. Obtaining a photographed image including the object; Identifying the object in the captured image by segmentation; generating a segmentation mask that segments the object in the captured image based on the segmentation detection result; and using the segmentation mask to aim the imaging device at the object within the imaging field of view and determine imaging conditions.

9. A process of acquiring a photographed image including the object; a process of identifying the object in the captured image by segmentation; generating a segmentation mask that segments the object in the captured image based on the segmentation detection result; and a process of using the segmentation mask to aim the imaging device at the object within the imaging field of view and determine imaging conditions.

Citation Information

Patent Citations

  • System, method, and apparatus for unsupervised adaptation of perception of autonomous mower

    JP2013153413A

  • Imaging control device and control method of the same

    JP2022045757A

  • Control unit, imaging apparatus, control method, and program

    JP2022170035A

  • Feature-based image autofocus

    US20210409609A1

  • Image-capturing device, image-capturing method, and program

    WO2021193147A1