Information processing device and method
The information processing device enhances antenna detection in aerial images by aligning annotations with edges, generating adjusted images, and using machine learning to improve detection accuracy and reduce processing load.
Patent Information
- Application Number
- JP2024085496
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2025-09-04
- Estimated Expiration
- 2041-12-27
AI Technical Summary
Existing image processing techniques face challenges with high processing load and accuracy issues when detecting and correcting annotations for objects like antenna devices in aerial images, particularly due to blending with outdoor backgrounds.
An information processing device that includes image acquisition, region identification, edge detection, annotation correction, and machine learning to generate a learning model for accurately detecting and correcting annotations of antenna devices in aerial images, using drones for capturing images and adjusting parameters to enhance detection accuracy.
The solution improves annotation efficiency and reduces processing load by aligning annotations with detected edges, generating adjusted images, and using machine learning to enhance the detection of antenna devices in aerial images, enabling accurate state determination.
Smart Images

Figure 0007734233000001 
Figure 0007734233000002 
Figure 0007734233000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to image processing techniques. [Background technology]
[0002] A technology has been proposed that acquires an image obtained by photographing an antenna and maps feature points contained in the acquired image into a three-dimensional spatial coordinate system (see Patent Document 1). Also, a technology has been proposed that estimates the pointing direction of an object by applying a box boundary having a pointing direction to the object detected using a machine learning model (see Non-Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6443700 [Non-patent literature]
[0004] [Non-Patent Document 1] Jingru Yi, et al., “Oriented Object Detection in Aerial Images with Box Boundary-Aware Vectors”, WACV2021, p.2150-2159 Summary of the Invention [Problem to be solved by the invention]
[0005] Various techniques have been proposed for processing a predetermined object in an image, but these techniques have had problems with processing load or processing accuracy.
[0006] In view of the above-mentioned problems, an object of the present disclosure is to provide a novel information processing technology related to a predetermined object in an image. [Means for solving the problem]
[0007] An example of the present disclosure is an information processing device that includes an image acquisition means that acquires an image used as training data for machine learning, the image having one or more annotations to indicate the position of a specified object in the image, an adjusted image generation means that generates an adjusted image in which parameters of the image have been adjusted, and a machine learning means that generates a learning model for detecting the specified object in an image by performing machine learning using training data including the adjusted image.
[0008] Furthermore, one example of the present disclosure is an information processing device including a processing target acquisition means for acquiring an image to be processed, an image acquisition means for acquiring an image to be processed that has one or more annotations attached to it to indicate the position at which a predetermined object is indicated in the image, an object detection means for detecting the predetermined object in the processing target image using a learning model for detecting the predetermined object in an image that is generated by machine learning using training data that includes the image to which the one or more annotations are attached, and an angle calculation means for calculating the angle of the detected object relative to a predetermined reference in the processing target image.
[0009] The present disclosure can be understood as an information processing device, a system, a method executed by a computer, or a program executed by a computer. The present disclosure can also be understood as such a program recorded on a recording medium readable by a computer or other device, machine, etc. Here, a recording medium readable by a computer, etc. refers to a recording medium that stores information such as data and programs by electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer, etc. [Effects of the Invention]
[0010] According to the present disclosure, it is possible to provide a novel information processing technique related to a predetermined object in an image. [Brief explanation of the drawings]
[0011] [Figure 1]1 is a schematic diagram illustrating a configuration of a system according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an outline of a functional configuration of an information processing apparatus according to an embodiment. [Figure 3] FIG. 1 illustrates an example of an annotated image according to an embodiment. [Figure 4] FIG. 10 is a diagram showing a region identified in an image in an embodiment. [Figure 5] FIG. 10 is a diagram illustrating an example of an image in which annotations have been corrected in the embodiment. [Figure 6] 10 is a flowchart showing the flow of annotation correction processing according to an embodiment. [Figure 7] 10 is a flowchart illustrating a flow of a data extension process according to the embodiment. [Figure 8] 1 is a flowchart illustrating a flow of machine learning processing according to an embodiment. [Figure 9] 10 is a flowchart illustrating a flow of a state determination process according to the embodiment. [Figure 10] FIG. 10 is a diagram showing an outline of calculation of an azimuth angle in a top-view image to be processed in an embodiment. [Figure 11] FIG. 10 is a diagram showing an outline of calculation of tilt in a side-view image to be processed in an embodiment. [Figure 12] FIG. 10 is a diagram illustrating an outline of the functional configuration of an information processing device according to a variation. [Figure 13] FIG. 10 is a diagram illustrating an outline of the functional configuration of an information processing device according to a variation. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of a system, an information processing device, a method, and a program according to the present disclosure will be described with reference to the drawings. However, the embodiments described below are merely examples, and the system, the information processing device, the method, and the program according to the present disclosure are not limited to the specific configurations described below. In implementing the present disclosure, a specific configuration according to the embodiment may be appropriately adopted, and various improvements and modifications may be made.
[0013] In this embodiment, an embodiment will be described in which the technology according to the present disclosure is implemented in a system that checks the installation status of an antenna device of a mobile base station using images taken from the air using a drone. However, the technology according to the present disclosure can be widely used as a technology for detecting a predetermined object in an image, and the application of the present disclosure is not limited to the example shown in the embodiment.
[0014] In recent years, as the number of miniaturized mobile communication base stations has increased, the importance of technology for monitoring whether the angles of the antenna devices that make up the base stations are in the correct state has increased. Conventionally, methods have been proposed for mapping the feature points of the antenna devices in a captured image onto space, but with these conventional techniques, there is a risk that the feature points cannot be sufficiently extracted because the antenna devices and the outdoor background roughly blend together during capture.
[0015] In view of this situation, the system, information processing device, method, and program according to this embodiment generate a group of images that have a common annotation indicating the location of the area corresponding to the antenna device and have different parameter adjustments, thereby expanding the learning data of the learning model that detects antenna devices.
[0016] Furthermore, technologies have been proposed to detect line features within an image, as well as technologies that estimate and correct the position to be labeled from surrounding data by clicking appropriate coordinates without the need to input precise coordinates. These technologies can support labeling and annotation by detecting edges and lines within the image to be labeled or annotated, but there is room for improvement in terms of correcting local annotations made manually and providing highly efficient annotation support.
[0017] In view of this situation, the system, information processing device, method, and program of this embodiment correct the position of annotations that have been manually or automatically added to an image based on the results of edge detection of the image, as an aid to annotations that indicate where the antenna device area is, which are performed on drone aerial images of the antenna device.
[0018] <System configuration> 1 is a schematic diagram showing the configuration of a system according to this embodiment. The system according to this embodiment includes an information processing device 1, a drone 8, and a user terminal 9, which are connected to a network and can communicate with each other.
[0019] The information processing device 1 is a computer including a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a storage device 14 such as an EEPROM (Electrically Erasable and Programmable Read Only Memory) or an HDD (Hard Disk Drive), a communication unit 15 such as a NIC (Network Interface Card), etc. However, the specific hardware configuration of the information processing device 1 can be omitted, replaced, or added as appropriate depending on the embodiment. Furthermore, the information processing device 1 is not limited to a device consisting of a single housing. The information processing device 1 may be realized by multiple devices using so-called cloud or distributed computing technology, etc.
[0020] The drone 8 is a small unmanned aerial vehicle whose flight is controlled in response to external input signals and / or programs stored in the device, and includes a propeller, motor, CPU, ROM, RAM, storage device, communication unit, input device, output device, etc. (not shown). However, the specific hardware configuration of the drone 8 may be omitted, replaced, or added as appropriate depending on the embodiment. The drone 8 according to this embodiment also includes an imaging device 81, which captures an image of a predetermined target (in this embodiment, an antenna device) in response to external input signals and / or programs stored in the device when flying around the target. In this embodiment, the captured image is primarily used to confirm the orientation of the antenna, among other things, of the installation status of the antenna device of the mobile base station. For this reason, the drone 8 and the imaging device 81 are controlled to a position and orientation that allows them to capture the antenna device from directly above the antenna device, thereby capturing an image of the antenna device as seen from directly above (a so-called top view). Furthermore, the drone 8 and the imaging device 81 are controlled to a position and orientation that allows them to capture the antenna device from directly beside the antenna device, thereby capturing an image of the antenna device as seen from directly beside the antenna device, thereby capturing an image of the antenna device as seen from directly beside the antenna device (a so-called side view). The imaging device 81 may be a camera equipped with an image sensor, or may be a depth camera equipped with a ToF (Time of Flight) sensor or the like.
[0021] Furthermore, the image data obtained by capturing may include, as metadata, data output from various devices mounted on the drone 8 or the imaging device 81 when the image was captured. Here, examples of the various devices mounted on the drone 8 or the imaging device 81 include a three-axis acceleration sensor, a three-axis angular velocity sensor, a GPS (Global Positioning System) device, and a direction sensor (compass). The data output from the various devices may include, for example, acceleration about each axis, angular velocity about each axis, position information, and direction. EXIF (exchangeable image file format) is known as a method for adding such metadata to image data, but the specific method for adding metadata to image data is not limited.
[0022] The user terminal 9 is a terminal device used by a user. The user terminal 9 is a computer equipped with a CPU, ROM, RAM, a storage device, a communication unit, an input device, an output device, etc. (not shown). However, the specific hardware configuration of the user terminal 9 can be omitted, replaced, or added as appropriate depending on the embodiment. Furthermore, the user terminal 9 is not limited to a device consisting of a single housing. The user terminal 9 may be realized by multiple devices using so-called cloud or distributed computing technology. Through these user terminals 9, users can create training data by annotating images, transfer images captured using a drone 8 to the information processing device 1, and so on. Note that, in this embodiment, annotation refers not only to the act of annotation but also to one or more points (key points), labels, etc., added to an image by annotation.
[0023] 2 is a diagram illustrating an outline of the functional configuration of the information processing device 1 according to this embodiment. The information processing device 1 functions as an information processing device including an image acquisition unit 21, a region identification unit 22, an edge detection unit 23, an estimation unit 24, an annotation correction unit 25, an adjusted image generation unit 26, a machine learning unit 27, a processing target acquisition unit 28, a target detection unit 29, and an angle calculation unit 30, by a program recorded in the storage device 14 being read into the RAM 13 and executed by the CPU 11, which controls each piece of hardware included in the information processing device 1. Note that in this embodiment and other embodiments described below, each function included in the information processing device 1 is executed by the CPU 11, which is a general-purpose processor, but some or all of these functions may be executed by one or more dedicated processors.
[0024] The image acquisition unit 21 acquires an image to be used as training data for machine learning, which has one or more annotations attached to it to indicate the position where a specified object (in this embodiment, an antenna device) is shown in the image.
[0025] FIG. 3 is a diagram showing an example of an image annotated and used as training data according to this embodiment. In this embodiment, the training data is used to generate and / or update a learning model for detecting antenna devices for mobile phone networks installed on outdoor structures such as utility poles and steel towers from images obtained by aerial photography using a drone 8 in flight. For this reason, annotations indicating the location of the antenna devices are pre-added to the images. In the example shown in FIG. 3, in an image obtained by looking down (in a substantially vertical direction) on an antenna device installed on a base station pole, multiple points are added as annotations to the outlines of the three box-shaped components that make up the antenna device (in other words, the boundaries between the antenna device and the background). (In FIG. 3, the positions of the points are shown as circles for visibility, but the positions where the annotations are made are the centers of the circles.) Note that, in this embodiment, an example has been described in which annotations are added as points indicating positions in an image. However, the annotations may be any annotation that indicates the area in the image where a specific object is captured, and the form of the annotation is not limited. The annotation may be, for example, a line, a curve, a shape, a fill, or the like added to the image.
[0026] The region identification unit 22 identifies a region in the image where one or more annotations satisfy a predetermined criterion. The predetermined criterion can be at least one of the density of annotations in the image, the position of annotations, the positional relationship between annotations, and the arrangement of annotations. For example, the region identification unit 22 may identify a region in the image where the amount of annotations relative to the area satisfies a predetermined criterion. Alternatively, for example, the region identification unit 22 may identify a region where the positions of multiple annotations have a predetermined relationship.
[0027] Fig. 4 is a diagram showing an area that satisfies a predetermined criterion and is identified in an image in this embodiment. According to the example shown in Fig. 4, an area where one or more annotations satisfy a predetermined criterion is identified, rather than the entire image. This limits the area to be subjected to edge detection, which will be described later, and it can be seen that the processing load for edge detection can be reduced compared to when edge detection is performed on the entire image. Here, the method for identifying an area is not limited, but a specific method for identifying an area will be described below.
[0028] First, an example of a method for identifying an area in an image where the amount of annotations relative to the area satisfies a predetermined standard will be described. For example, for each combination of some or all of the annotations in the image (in the example shown in FIG. 4, a combination of four adjacent annotations), the area identification unit 22 calculates the center of gravity of the area formed by connecting these annotations and the area where the density of these annotations is not below a predetermined density (the area may be expressed, for example, as the number of pixels). By setting an area having the area including the center of gravity, for example, the area identification unit 22 can identify an area where the amount of annotations relative to the area satisfies the predetermined standard. Furthermore, for each combination of some or all of the annotations in the image, the area identification unit 22 can set a circumscribing rectangle that includes these annotations. The area identification unit 22 can then expand this rectangular area in all directions until the annotation density calculated using the area of the rectangle and the number of annotations reaches a predetermined threshold, thereby identifying an area where the amount of annotations relative to the area satisfies the predetermined standard. However, the above-described method is an example of a technique for identifying an area, and an area may be any area in which the amount of annotation relative to its area meets a predetermined standard, and other specific techniques may be adopted to identify an area.
[0029] Next, an example of a method for identifying an area where the positions of multiple annotations have a predetermined relationship will be described. For example, the area identification unit 22 identifies the positional relationship of the annotations included in each combination of some or all of the annotations in an image. For example, each of the three box-shaped members constituting the antenna device, which is a predetermined target according to this embodiment, has a substantially polygonal shape (a quadrangle in the example shown in FIG. 4 ) in which the sides have a predetermined length relationship (ratio) in a plan view. Therefore, for each combination of annotations consisting of the same number of annotations as the number of vertices of a polygon (four in the example shown in FIG. 4 ), the area identification unit 22 can identify the predetermined area by determining whether the annotations included in the combination have a predetermined positional relationship as the vertices of a substantially polygon in which the sides have a predetermined length relationship (ratio). Alternatively, for each combination of annotations, the predetermined area may be identified by determining whether the straight lines formed by the multiple annotations are substantially parallel or substantially perpendicular to each other. However, the above-described method is merely an example of a method for identifying an area, and the area may be any area in which the positions of the annotations related to the area are in a predetermined relationship, and other specific methods may be adopted for identifying the area. Also, in this embodiment, an example of identifying a rectangular area has been described, but the shape of the area is not limited, and may be, for example, a circle.
[0030] The edge detection unit 23 performs edge detection with priority in the identified region or in a range set based on the identified region. That is, the edge detection unit 23 may use the region identified by the region identification unit 22 as is, or may set a different range based on the identified region (for example, by setting a margin) and use this range. Edge detection may be performed by selecting an appropriate edge detection method from among conventionally used edge detection methods and edge detection methods to be devised in the future, and therefore a description thereof will be omitted. Conventionally known edge detection methods include, for example, the gradient method, the Sobel method, the Laplacian method, and the Canny method, but the edge detection methods and filters that can be employed are not limited.
[0031] The estimation unit 24 estimates the position where the annotation was intended based on the detected edges. The estimation unit 24 estimates the position where the annotation was intended by referring to edges detected in the vicinity of the position of the annotation. More specifically, for example, the estimation unit 24 may estimate the position closest to the annotation among the edges detected in the region as the position where the annotation was intended. Furthermore, for example, the estimation unit 24 may estimate the position having a predetermined characteristic among the edges detected in the region as the position where the annotation was intended. Here, examples of the position having the predetermined characteristic include a position where edge lines intersect, a position where edge lines form an angle, or a position where edge lines have a predetermined shape.
[0032] The annotation correction unit 25 corrects the position of the annotation to align with the detected edge by moving the position of the annotation to the position estimated by the estimation unit 24. As described above, the position estimated by the estimation unit 24 is, for example, the position closest to the annotation among the edges detected in the region, the position where the edge lines intersect, the position where the edge lines form an angle, or the position where the edge lines have a predetermined shape. In this way, the position of the annotation can be corrected to the outline of a predetermined object in the image (in other words, the boundary with the background) that is thought to have been originally intended by the annotator.
[0033] Fig. 5 is a diagram showing an example of an image in which annotations have been corrected in this embodiment. According to the example shown in Fig. 5, it can be seen that the position of the annotation, which was added at a position shifted from the edge in Fig. 3, has been corrected, and the annotation is now correctly added to the outline (in other words, the boundary with the background) of a predetermined object (in this embodiment, the antenna device).
[0034] The adjusted image generation unit 26 generates an adjusted image in which image parameters have been adjusted. Here, the adjusted image generation unit 26 generates an adjusted image in which image parameters have been adjusted to make it difficult to detect a predetermined target. An example of an adjustment method for making it difficult to detect a predetermined target is to make the parameters of each pixel similar or identical between the pixel in which the predetermined target (in this embodiment, an antenna device) is captured and the pixel in which the background of the predetermined target (e.g., the ground, buildings, plants, structures on the ground, etc.) is captured (in other words, to make the color of the predetermined target a camouflage color against the background color). Here, the adjusted image generation unit 26 may generate an adjusted image in which parameters related to at least one of the image parameters, such as brightness, exposure, white balance, hue, saturation, brightness, sharpness, noise, and contrast, have been adjusted.
[0035] The adjusted image generation unit 26 may also generate multiple different adjusted images based on a single image. That is, the adjusted image generation unit 26 may generate a first adjusted image in which image parameters have been adjusted, and a second adjusted image in which image parameters have been adjusted differently from those of the first adjusted image. In this case, the multiple adjusted images generated may include adjusted images in which the same type of parameter has been adjusted to different degrees, and / or adjusted images in which different types of parameters have been adjusted. Each of the multiple adjusted images may have the same multiple annotations. The annotations may be annotations corrected through edge detection of the single image, or may be annotations corrected through edge detection of any of the adjusted images. The edge detection unit 23 may perform edge detection on the adjusted image generated by the adjusted image generation unit 26. The estimation unit 24 may estimate the position closest to the annotation among the edges detected in the adjusted image as the intended position of the annotation. Here, if the edge lines in the adjusted image generated by the adjusted image generation unit 26 exhibit characteristics such as a predetermined positional relationship, the position of the edges detected in the adjusted image that is closest to the annotation may be estimated as the position where the annotation was intended.
[0036] The machine learning unit 27 generates a learning model for detecting a predetermined object in an image by performing machine learning using teacher data including an image corrected by the annotation correction unit 25 and / or an adjusted image. For example, in this embodiment, as also exemplified in the angle calculation unit 30 described later, generation of a learning model for detecting a predetermined object in an image is exemplified using supervised machine learning using the PyTorch library (see Non-Patent Document 1). However, since an appropriate machine learning algorithm may be selected and used from conventionally used machine learning algorithms and machine learning algorithms to be devised in the future, a description thereof will be omitted.
[0037] Here, the images used as training data by the machine learning unit 27 may be images with one or more annotations indicating the position of a predetermined object in the image, and the type of image used as training data is not limited. The machine learning unit 27 can use, as training data, an image acquired directly by the image acquisition unit 21, an image corrected by the annotation correction unit 25, an adjusted image generated by the adjusted image generation unit 26, an adjusted image generated by the adjusted image generation unit 26 based on an image corrected by the annotation correction unit 25, and the like. Furthermore, as described above, training data including multiple different adjusted images generated based on a single image, i.e., a first adjusted image and a second adjusted image, may be used. The images used as training data and the adjusted images may each be accompanied by the same multiple annotations.
[0038] The processing target acquisition unit 28 acquires a processing target image. In this embodiment, the processing target image is an image captured from the air using an imaging device 81 mounted on a drone 8 in flight. However, the processing target image may be an image in which a predetermined object is to be detected, and may be an RGB image or a depth image, and the type of the processing target image is not limited.
[0039] The object detection unit 29 detects a predetermined object in the processing target image using a learning model for detecting a predetermined object in an image, generated by the machine learning unit 27. In this embodiment, the object detection unit 29 detects an outdoor antenna device as the predetermined object in the processing target image. However, the object detection unit 29 can detect various objects from an image depending on the image used as training data and the annotation target, and the type of predetermined object detected using the technology disclosed herein is not limited. Furthermore, the detected predetermined object is typically identified using a method similar to the annotation added to the training data. That is, if the annotation is a point indicating the outline of the predetermined object, the object detection unit 29 identifies the predetermined object in the processing target image by adding a point to the outline of the predetermined object. However, the method for identifying the predetermined object is not limited, and the predetermined object may be identified using a method different from the annotation.
[0040] The angle calculation unit 30 calculates the angle of the detected object with respect to a predetermined reference in the processing target image. More specifically, in this embodiment, the angle calculation unit 30 calculates the angle of the detected object with respect to one of a predetermined direction, the vertical direction, and the horizontal direction in the processing target image. Here, the method by which the angle calculation unit 30 calculates the angle is not limited, but for example, a method may be adopted in which the direction of the object is detected using a technique such as detection using a machine learning model (see Non-Patent Document 1) or detection by comparison with a predefined object shape, and the angle formed between the direction of the detected object and a reference direction in the processing target image is calculated.
[0041] <Processing flow> Next, a flow of processing executed by the information processing device 1 according to this embodiment will be described. Note that the specific content and processing order of the processing described below are an example for implementing the present disclosure. The specific content and processing order may be selected as appropriate depending on the embodiment of the present disclosure.
[0042] To perform the annotation correction process, data augmentation process, and machine learning process described below, a user prepares training data including annotated images in advance. In this embodiment, the technology disclosed herein is used in a system intended to detect antenna devices installed outdoors as predetermined targets. Therefore, multiple images, including images showing antenna devices, are obtained. Note that the multiple images may also include images that do not show antenna devices. Training data is then created by annotating the multiple images obtained with annotations indicating the outlines of the antenna devices. In this case, the task of annotating the images may be performed manually by an annotator or automatically. Details of the process of annotating images will not be described here, as conventional annotation support technology may be used.
[0043] 6 is a flowchart showing the flow of annotation correction processing according to this embodiment. The processing shown in this flowchart is executed when training data including annotated images is prepared and an instruction for annotation correction is input by the user.
[0044] In step S101, training data including an image with annotations is acquired. The image acquisition unit 21 acquires, as training data, an image with one or more annotations attached to indicate the position of a predetermined object (in this embodiment, an antenna device) in the image. Thereafter, the process proceeds to step S102.
[0045] In steps S102 and S103, regions where one or more annotations satisfy a predetermined criterion are identified, and edges are detected in the identified regions, etc. The region identification unit 22 identifies regions in the image in the training data obtained in step S101 where one or more annotations satisfy a predetermined criterion (step S102). Then, the edge detection unit 23 performs edge detection in the region identified in step S102 or in a range set based on the region (step S103). Thereafter, the process proceeds to step S104.
[0046] In steps S104 and S105, the annotation is corrected so as to align with the detected edge. The estimation unit 24 estimates the intended position of the annotation based on the edge detected in step S103 (step S104). Then, the annotation correction unit 25 corrects the annotation so as to align with the detected edge by moving the position of the annotation to the position estimated in step S104 (step S105). Thereafter, the processing shown in this flowchart ends.
[0047] The annotation correction process described above improves the efficiency of the correction process for annotations attached to images used as training data for machine learning, making it possible to correct annotations with a smaller processing load than conventional methods.
[0048] 7 is a flowchart showing the flow of data augmentation processing according to this embodiment. The processing shown in this flowchart is executed when training data including annotated images is prepared and a data augmentation instruction is input by the user.
[0049] In step S201, training data including an image with annotations is acquired. The image acquisition unit 21 acquires, as training data, an image with one or more annotations attached to indicate the position of a predetermined object (in this embodiment, an antenna device) in the image. Note that the annotated image acquired here is preferably an image to which annotation correction has been applied by the annotation correction process described with reference to FIG. 6, but an image to which annotation correction has not been applied may also be acquired. Thereafter, the process proceeds to step S202.
[0050] In steps S202 and S203, one or more adjusted images are generated. Adjusted image generation unit 26 generates an adjusted image in which the parameters of the image acquired in step S201 have been adjusted (step S202). Once an adjusted image is generated, it is determined whether or not the generation of adjusted images for all preset patterns for the image acquired in step S201 has been completed (step S203). If not completed (NO in step S203), the process returns to step S202. That is, adjusted image generation unit 26 repeats the process of step S202 while changing the content of parameter adjustment based on one image acquired in step S201, and generates multiple adjusted images that are different from each other. If the generation of adjusted images for all preset patterns has been completed (YES in step S203), the process shown in this flowchart ends.
[0051] The data augmentation process described above makes it possible to reduce the effort required to improve the performance of a learning model generated by machine learning using annotated images.
[0052] 8 is a flowchart showing the flow of machine learning processing according to this embodiment. The processing shown in this flowchart is executed when training data including annotated images is prepared and a user inputs a machine learning instruction.
[0053] In step S301, training data including an image with annotations is acquired. The image acquisition unit 21 acquires, as training data, an image with one or more annotations attached to indicate the position of a predetermined object (in this embodiment, an antenna device) in the image. Note that the annotated image acquired here is preferably an image to which annotation correction has been applied by the annotation correction process described with reference to FIG. 6 and / or an adjusted image generated by the data augmentation process described with reference to FIG. 7, but an image to which neither annotation correction nor parameter adjustment has been applied may also be acquired. Thereafter, the process proceeds to step S302.
[0054] In step S302, a learning model is generated or updated. The machine learning unit 27 generates a learning model for detecting a predetermined object (in this embodiment, an antenna device) in the image by performing machine learning using training data including the image acquired in step S301, or updates an existing learning model. Thereafter, the processing shown in this flowchart ends.
[0055] 9 is a flowchart showing the flow of the state determination process according to this embodiment. The process shown in this flowchart is executed when image data of an image to be processed is prepared and a state determination instruction is input by the user.
[0056] The user captures an image of the antenna device of the base station using the imaging device 81 of the drone 8 in flight, and inputs the obtained image data of the image to be processed into the information processing device 1. At this time, the user may capture an image so that multiple antenna devices are included in one image to be processed. When multiple antenna devices are included in one image to be processed, the state determination process is performed for each area of the antenna device included in the image to be processed. Although the imaging method and the method of inputting the image data into the information processing device 1 are not limited, in this embodiment, an antenna device installed on a structure is captured using the drone 8 equipped with the imaging device 81, and the image data transferred from the imaging device 81 to the user terminal 9 via communication or a recording medium is further transferred to the information processing device 1 via a network, thereby inputting the image data of the image to be processed into the information processing device 1.
[0057] In steps S401 and S402, a predetermined object in the processing target image is detected using a learning model. The processing target acquisition unit 28 acquires the processing target image (in this embodiment, an image captured from the air using the imaging device 81 mounted on the drone 8 in flight) (step S401). Then, the object detection unit 29 detects the predetermined object (in this embodiment, an antenna device) in the processing target image acquired in step S401 using the learning model generated by the machine learning process described with reference to FIG. 8 (step S402). Thereafter, the process proceeds to step S403.
[0058] In steps S403 and S404, the tilt of the detected object is calculated. The angle calculation unit 30 calculates the angle of the object detected in step S402 with respect to a predetermined reference in the processing target image.
[0059] FIG. 10 is a diagram illustrating an outline of calculation of an azimuth angle in a top-view image to be processed in this embodiment. FIG. 10 illustrates an outline of a case in which angle calculation unit 30 calculates the angle that the orientation of an antenna device (a predetermined target) detected from the image to be processed forms with respect to a predetermined reference north direction (which may be true north or magnetic north). First, angle calculation unit 30 determines a reference direction (here, north direction) in the image to be processed (step S403). In this embodiment, the image to be processed is assumed to have been corrected in advance so that the directly upward direction of the image is north, and the directly upward direction of the image is determined as the reference direction. However, the reference direction may be determined by other methods. For example, if the image to be processed has not been corrected so that the directly upward direction of the image is north, the north direction in the image may be identified by a method of referring to metadata attached to the image to be processed (such as acceleration on each axis, angular velocity on each axis, position information, and direction) or by a method of comparing the image to be processed with a map image, and the north direction may be determined as the reference direction. Furthermore, a direction other than north may be adopted as the reference direction. For example, the reference direction may be a design-correct installation direction of a predetermined object (in this embodiment, the antenna device), a vertical direction, a horizontal direction, or the like.
[0060] The angle calculation unit 30 then determines the orientation of the detected antenna device (predetermined target) (step S404). The method by which the angle calculation unit 30 determines the orientation of the predetermined target is not limited. For example, the angle calculation unit 30 may use a machine learning model to estimate the orientation of the detected target by applying a box boundary having an orientation direction to the target (see Non-Patent Document 1), or may use a method of determining the front direction of the detected antenna device by reading a combination of a predefined antenna device shape and the front direction of the antenna device for that shape and applying the combination to the outline of the detected antenna device. The angle calculation unit 30 then calculates the angle between the determined reference direction and the determined front direction of the antenna device. In the example shown in FIG. 10, the angle between the reference direction indicated by the thin line with an arrow and the front direction of the antenna device indicated by the thick line with an arrow is calculated. As described above, the reference direction may be a direction, a correct installation direction, a vertical direction, a horizontal direction, or the like, based on the design of the predetermined target.
[0061] FIG. 11 is a diagram illustrating an outline of tilt calculation for a side-view image of a processing target in this embodiment. FIG. 11 illustrates an outline of a case in which angle calculation unit 30 calculates the angle that the tilt of an antenna device (predetermined target) detected from the processing target image forms with respect to the vertical direction, which is a predetermined reference. First, angle calculation unit 30 determines a reference direction (here, the vertical direction) in the processing target image (step S403). In this embodiment, it is assumed that the center pole in the image is correctly installed in the vertical direction, and the longitudinal direction of the center pole is determined as the reference direction. However, the reference direction may be determined by other methods. For example, the vertical direction in the image may be identified by a method of referencing metadata attached to the processing target image (such as the acceleration of each axis and the angular velocity of each axis), and the vertical direction may be determined as the reference direction. In this manner, angle calculation unit 30 can calculate the azimuth, tilt, etc. of the predetermined target. Then, the process proceeds to step S405.
[0062] In step S405, the state of a predetermined object is determined. In this embodiment, the information processing device 1 determines whether the installation state of the antenna device is correct by determining whether the angle calculated in step S404 is within a predetermined range. Thereafter, the process shown in this flowchart ends, and the determination result is output to the user.
[0063] According to the state determination process described above, the angle of a specified object relative to a reference direction can be obtained, and by referring to the obtained angle, it is possible to determine the state of the specified object (in this embodiment, the installation state of the antenna device).
[0064] <Variations> In the above-described embodiment, an example has been described in which the annotation correction process, data augmentation process, machine learning process, and state determination process are executed in a single information processing device, but these processes may be separated and executed by separate information processing devices. In this case, some of the image acquisition unit 21, region identification unit 22, edge detection unit 23, estimation unit 24, annotation correction unit 25, adjusted image generation unit 26, machine learning unit 27, processing target acquisition unit 28, target detection unit 29, and angle calculation unit 30 included in the information processing device 1 may be omitted.
[0065] 12 is a diagram showing an outline of the functional configuration of an information processing device 1b according to a variation. The information processing device 1b functions as an information processing device including an image acquisition unit 21, a region identification unit 22, an edge detection unit 23, an estimation unit 24, an annotation correction unit 25, a machine learning unit 27, a processing target acquisition unit 28, and a target detection unit 29. Each function of the information processing device 1b is generally similar to that of the embodiment described above, except that the adjusted image generation unit 26 and the angle calculation unit 30 are omitted, and therefore description thereof will be omitted.
[0066] 13 is a diagram showing an outline of the functional configuration of an information processing device 1c according to a variation. The information processing device 1c functions as an information processing device including an adjusted image generation unit 26, a machine learning unit 27, a processing target acquisition unit 28, and a target detection unit 29. The functions of the information processing device 1c are generally similar to those of the embodiment described above, except that the image acquisition unit 21, the region identification unit 22, the edge detection unit 23, the estimation unit 24, the annotation correction unit 25, and the angle calculation unit 30 are omitted, and therefore description thereof will be omitted.
[0067] Furthermore, in the above-described embodiment, an example has been described in which aerial photography is performed using the drone 8, but other devices (aircraft, etc.) may also be used for aerial photography. [Explanation of symbols]
[0068] 1. Information processing equipment
Claims
1. An image acquisition means for acquiring an image having one or more annotations to indicate the position of a predetermined object in the image; an adjusted image generating means for generating an adjusted image in which parameters relating to the entire image are adjusted so that pixels in which the predetermined object is captured and pixels in which the background of the predetermined object is captured are similar in parameters for each pixel, so that detection of the predetermined object becomes difficult; a machine learning means for performing machine learning using training data including the image with the one or more annotations and the adjusted image, thereby generating a learning model for detecting the predetermined object in the image; a processing target acquisition means for acquiring a processing target image; an object detection means for detecting the predetermined object in the processing target image using the learning model; An angle calculation means for calculating an angle of the detected object relative to a predetermined reference in the processing object image; An information processing device comprising:
2. the adjusted image generating means generates a first adjusted image in which parameters of the image are adjusted, and a second adjusted image in which parameters of the image are adjusted to be different from those of the first adjusted image; the machine learning means performs machine learning using training data including the first adjusted image and the second adjusted image. The information processing device according to claim 1 .
3. the adjusted image generating means generates the adjusted image in which at least a parameter related to image brightness among the parameters of the image has been adjusted.
3. The information processing device according to claim 1 or 2.
4. the angle calculation means calculates the angle of the detected object relative to any one of a predetermined direction, a vertical direction, and a horizontal direction in the image to be processed; The information processing device according to claim 1 .
5. The computer an image acquisition step of acquiring an image with one or more annotations indicating the location of a predetermined object in the image; an adjusted image generating step of generating an adjusted image in which parameters relating to the entire image are adjusted so that pixels in which the predetermined object is captured and pixels in which a background of the predetermined object is captured are similar in parameters for each pixel, so that detection of the predetermined object becomes difficult; a machine learning step of generating a learning model for detecting the predetermined object in an image by performing machine learning using training data including the image with the one or more annotations and the adjusted image; a processing target acquisition step of acquiring a processing target image; an object detection step of detecting the predetermined object in the processing target image using the learning model; an angle calculation step of calculating an angle of the detected object relative to a predetermined reference in the processing target image; How to do it.
6. The computer In the adjusted image generating step, a first adjusted image in which a parameter of the image is adjusted and a second adjusted image in which a parameter of the image is adjusted to be different from that of the first adjusted image are generated; In the machine learning step, machine learning is performed using training data including the first adjusted image and the second adjusted image. The method of claim 5.
7. The computer In the adjusted image generating step, the adjusted image is generated in which at least a parameter related to image brightness among the parameters of the image is adjusted.
7. The method according to claim 5 or 6.
8. The computer In the angle calculation step, an angle of the detected object relative to a predetermined direction, a vertical direction, or a horizontal direction in the image to be processed is calculated.
8. The method according to any one of claims 5 to 7.
Citation Information
Patent Citations
Method of fixing construction of tubular lock bolt
JP1989043700A
Machine learning program, machine learning method, and machine learning device
JP2019185483A
Expansion device, expansion method and expansion program
JP2020034998A
Learning program, learning method, learning apparatus, detection program, detection method and detection apparatus
JP2020080022A
Data extension system, data extension method and program
JP2020197833A
Cited By
Information processing apparatus and method
US12718434B2