Sample enhancement method and device
By performing perspective transformation and annotation information calibration on drone aerial images, an enhanced sample set is generated, which solves the problem of detection model performance degradation caused by image distortion during drone inspections and improves the detection accuracy and generalization ability of the model.
Patent Information
- Application Number
- CN202511277323.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-09
AI Technical Summary
During drone inspections, aerial images are distorted due to changes in flight attitude, and existing target detection models are difficult to adapt to, resulting in a decrease in detection accuracy and recognition precision.
By constructing the original image sample set, determining the perspective transformation matrix, performing image perspective transformation, calibrating the annotation information, and generating an enhanced sample set, the sample diversity and spatial information consistency are enhanced.
The target detection model's ability to detect targets with different distortions is improved, the acquisition time is reduced, and the model's generalization performance and detection accuracy are enhanced.
Smart Images

Figure CN120766065A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of model training technology, and in particular to a sample enhancement method and device. Background Art
[0002] Drone inspections involve the use of drones equipped with cameras to conduct automated or semi-automated inspection, monitoring, and data collection of target areas. Due to their flexibility, wide coverage, and efficient information collection, drone inspections are widely used in areas such as power inspections, agricultural monitoring, and construction surveying.
[0003] During actual inspections, a drone's flight attitude is affected by factors such as airflow and control accuracy, making it difficult to maintain complete stability. This can lead to distortion in the captured aerial images. A drone's flight attitude is primarily controlled by its yaw, roll, and pitch angles. Changes in the yaw angle have minimal impact on the scale and morphology of the target in the aerial image and can be corrected through image rotation. Changes in the roll angle can cause morphological distortion in the aerial image, meaning the target's posture in the image is distorted. This distortion can lead to misjudgments or missed detections in target detection models that rely on external features for recognition. Changes in the pitch angle can cause both morphological and scale distortion in the aerial image. This means the target's shape and proportions in the aerial image change simultaneously, significantly deviating from its true form. This can lead to misjudgments, missed detections, or large errors in dimensional measurement.
[0004] The training logic of the target detection model is to establish a database based on target features with standard viewing angles, fixed scales, and regular shapes. Different pitch angles or roll angles will cause different degrees of distortion of the targets in the aerial images, resulting in a serious deviation between the actual collected image data and the "standard features" in the training database, which in turn causes the detection accuracy, recognition precision, and recall rate of the target detection model to drop significantly, and may even make it impossible to complete the detection task. Summary of the Invention
[0005] The embodiments of the present application provide a sample enhancement method and device to solve the problem that existing training samples are difficult to adapt to actual flight postures and cannot overcome the performance degradation of target detection models caused by image distortion.
[0006] In a first aspect, an embodiment of the present application provides a sample enhancement method, comprising: constructing an original image sample set based on annotated aerial images; determining a value range of the transformation parameter according to the scale transformation relationship of the aerial image before and after the perspective transformation, the geometric relationship between the shooting parameters of the aerial image, and the range of the target angle; taking values within the value range of the transformation parameter to determine multiple perspective transformation matrices; performing perspective transformation on the aerial images in the original image sample set using multiple perspective transformation matrices to obtain multiple transformed images; updating the target annotation box in the transformed image using coordinate mapping to calibrate the annotation information; and combining the annotated aerial image and its annotation information with the transformed image and its calibrated annotation information to obtain an enhanced sample set.
[0007] In combination with the first aspect, in a possible implementation method, before constructing the original image sample set based on the labeled aerial images, it also includes: cleaning the aerial images to remove noise thereon; and / or, dehazing and enhancing the aerial images to improve their clarity; and / or, standardizing the shooting parameters of the aerial images; and / or, geometrically correcting and segmenting the aerial images.
[0008] In combination with the first aspect, in a possible implementation method, the method for determining the scale transformation relationship of the aerial image before and after the perspective transformation includes: determining multiple original feature points in the aerial image, and obtaining multiple target feature points after performing a perspective transformation on the aerial image; and determining the scale transformation relationship of the aerial image before and after the perspective transformation based on the change in the distance between different original feature points and the distance between the corresponding target feature points.
[0009] In combination with the first aspect, in a possible implementation, the method for determining the geometric relationship between the shooting parameters of the aerial image includes: the shooting parameters include the field of view angle, field of view area, altitude and attitude angle of the aerial photography device; based on the relationship between the distance of the aerial photography device from the field of view area and the field of view angle, determining the geometric relationship between the pitch angle / roll angle and the field of view.
[0010] In combination with the first aspect, in a possible implementation manner, determining the value range of the transformation parameter includes: determining the value range of the transformation parameter according to the range of the target angle, the geometric relationship, and the scale transformation relationship.
[0011] In combination with the first aspect, in a possible implementation method, taking values within the value range of the transformation parameters to determine multiple perspective transformation matrices includes: determining the coordinates of the original feature points and the target feature points based on the transformation parameters and the scale transformation relationship; determining the perspective transformation matrix corresponding to the transformation parameters based on the coordinates of the original feature points and the target feature points.
[0012] With reference to the first aspect, in a possible implementation manner, the updating the target bounding box in the transformed image by using the coordinate mapping comprises: determining coordinates of a new target bounding box in the transformed image according to corner point coordinates of the original target bounding box and the perspective transformation matrix; determining the new target bounding box that is cropped at an edge based on an interaction relationship between the original target bounding box and the new target bounding box; deleting an unqualified target bounding box according to an area and an aspect ratio of the original target bounding box and the new target bounding box; and updating the target bounding box and the labeling information of the transformed image according to the remaining new target bounding boxes and class information of the new target bounding boxes.
[0013] With reference to the first aspect, in a possible implementation manner, the determining the new target bounding box that is cropped at an edge based on an interaction relationship between the original target bounding box and the new target bounding box comprises: if three corner points of the original target bounding box are cropped in the transformed image, and an original center point of the original target bounding box is not on a cropping line, the original target bounding box is deleted; if three corner points of the original target bounding box are cropped in the transformed image, and the original center point of the original target bounding box is on the cropping line, a new target bounding box is constructed with the original center point as a corner point; if two corner points of the original target bounding box are cropped in the transformed image, the remaining original target bounding box is taken as the new target bounding box; and if one corner point of the original target bounding box is cropped in the transformed image, a new target bounding box is constructed with an intersection point of a diagonal line of the original target bounding box and the cropping line as a corner point.
[0014] With reference to the first aspect, in a possible implementation manner, the deleting the unqualified target bounding box according to an area and an aspect ratio of the original target bounding box and the new target bounding box comprises: if the area of the new target bounding box reaches half of the area of the original target bounding box, the aspect ratio is compared; if a ratio of the aspect ratio of the new target bounding box to the aspect ratio of the original target bounding box is within a preset range, the new target bounding box is retained; if the ratio of the aspect ratio of the new target bounding box to the aspect ratio of the original target bounding box is not within the preset range, the new target bounding box is deleted; and if the area of the new target bounding box does not reach half of the area of the original target bounding box, the new target bounding box is deleted.
[0015] In a second aspect, an embodiment of the present application provides a sample enhancement device, comprising: a construction module for constructing an original image sample set based on an annotated aerial image; a transformation parameter module for determining a value range of the transformation parameter based on the scale transformation relationship of the aerial image before and after the perspective transformation, the geometric relationship between the shooting parameters of the aerial image, and the range of the target angle; a perspective transformation matrix module for taking values within the value range of the transformation parameter to determine multiple perspective transformation matrices; a perspective transformation module for performing perspective transformation on the aerial images in the original image sample set using multiple perspective transformation matrices to obtain multiple transformed images; a calibration module for updating the target annotation box in the transformed image using coordinate mapping to calibrate the annotation information; and a combination module for combining the annotated aerial image and its annotation information with the transformed image and its calibrated annotation information to obtain an enhanced sample set.
[0016] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: By determining the range of values for the transformation parameters, the embodiments of the present application can determine multiple different perspective transformation matrices, thereby performing multi-angle perspective transformations on the original image and enriching the diversity of the samples. By calibrating the annotation information of the transformed image after the perspective transformation, the consistency of the image spatial information and the annotation information can be ensured, thereby improving the reliability of model training. This effectively solves the problem that existing training samples are difficult to adapt to actual flight attitudes and cannot overcome the performance degradation of the target detection model caused by image distortion. It can also increase the diversity of samples, allowing the target detection model to better detect target features with different distortions. It can also reduce the acquisition time required to collect diverse samples, and improve the generalization performance of the target detection model in adapting to the target environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments of the present application or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 A flow chart of a sample enhancement method provided in an embodiment of the present application; Figure 2 This is an example diagram of forward perspective transformation provided in an embodiment of the present application; Figure 3 This is an example diagram of negative perspective transformation provided by an embodiment of the present application; Figure 4 An example diagram of the relationship between the shooting parameters of the shooting device provided in an embodiment of the present application and the ground field of view; Figure 5 An example diagram of an aerial image provided in an embodiment of the present application; Figure 6 The embodiment of this application provides Figure 5 Instance image after perspective transformation; Figure 7 Schematic diagrams of various situations in which the original target annotation frame provided in the embodiment of the present application is cropped; Figure 8 A schematic structural diagram of a sample enhancement device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0020] The following description of some of the technologies involved in the embodiments of this application is provided to facilitate understanding and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for the sake of clarity and conciseness, some descriptions of well-known functions and structures are omitted from the following description.
[0021] Figure 1 This is a flow chart of a sample enhancement method provided in an embodiment of the present application, including steps 101 to 106. Figure 1 This is only an execution order shown in the embodiment of the present application, and does not represent the only execution order of the sample enhancement method. If the final result can be achieved, Figure 1 The steps shown may be performed in parallel or reversed.
[0022] Step 101: Construct a set of original image samples based on the labeled aerial images. In this embodiment of the present application, to enable the trained object detection model to better adapt to the complex flight scenarios, the drone can adaptively adjust its flight altitude, sensor parameters, and shooting angle to optimize data resolution and coverage. Furthermore, the collected aerial images should cover inspection scenarios in different seasons, lighting conditions, weather conditions, and shooting angles.
[0023] In embodiments of the present application, to ensure the reliability of sample data for the target detection model, aerial images may be preprocessed, including cleaning the aerial images to remove noise, dehazing and enhancing the aerial images to improve clarity, standardizing the aerial image capture parameters, and / or geometrically correcting and segmenting the aerial images.
[0024] Specifically, aerial images can be cleaned by using denoising and filtering algorithms to remove noise and sensor interference. Alternatively, images can be dehazed and enhanced to improve clarity. Aerial image capture parameters can also be standardized, unifying data formats and dimensions across different capture devices and acquisition times. Aerial images can also be rotated, cropped, and segmented to identify key areas, such as inspection targets, improving image quality and the accuracy of subsequent feature extraction.
[0025] The mainstream training method for object detection models is supervised training, which requires sample labeling. This is achieved by adding precise object annotation boxes to aerial images and labeling the object type within the boxes to guide the object detection model's learning. To improve the accuracy of the annotation information, traditional manual labeling is typically used for the initial annotation. This involves manually selecting the object and marking its category. The object annotation box must be the object's minimum bounding rectangle, and objects of the same type must be of the same category.
[0026] Finally, the preprocessed and labeled aerial images are stored as the original image sample set.
[0027] Step 102: Determine the range of values for the transformation parameters based on the scale transformation relationship of the aerial image before and after the perspective transformation, the geometric relationship between the shooting parameters of the aerial image, and the range of the target angle. In an embodiment of the present application, the method for determining the scale transformation relationship of the aerial image before and after the perspective transformation includes: determining a plurality of original feature points in the aerial image, performing a perspective transformation on the aerial image to obtain a plurality of target feature points. Determine the scale transformation relationship of the aerial image before and after the perspective transformation based on the change in the distance between different original feature points and the distance between corresponding target feature points.
[0028] In an embodiment of the present application, a method for determining a geometric relationship between shooting parameters of an aerial image includes: the shooting parameters include the field of view angle, field of view area, altitude, and attitude angle of the aerial camera device; and determining a geometric relationship between the pitch angle / roll angle and the field of view based on the relationship between the distance of the aerial camera device from the field of view area and the field of view angle.
[0029] It should be noted that this application does not restrict the method for determining the range of target angles. It can be determined based on the allowable shooting angle range or commonly used angle range of the aerial photography device, or it can be customized by the user based on actual needs. Preferably, the preset range can be expanded based on the shooting angle range of the original image sample set.
[0030] In the embodiment of the present application, determining the value range of the transformation parameter includes: determining the value range of the transformation parameter according to the range of the target angle, the geometric relationship, and the scale transformation relationship.
[0031] Specifically, perspective transformation aims to solve the problem of distortion caused by changes in attitude angles (yaw angle, roll angle and pitch angle, especially roll angle and pitch angle) of drone aerial images. By accurately constructing the perspective transformation matrix, image correction is achieved to improve the reliability of inspection analysis. The core of the transformation parameter generation process is to solve 3 3, the perspective transformation matrix is calculated by establishing a geometric relationship between a plurality of groups of feature point pairs corresponding to a plurality of (exemplarily four) original feature points and a plurality of (exemplarily four) target feature points.
[0032] Furthermore, since changes in the pitch angle primarily affect the deformation of the aerial image in the horizontal direction, while changes in the roll angle primarily affect the deformation of the aerial image in the vertical direction, the change from vertical to horizontal can be achieved by rotating the aerial image 90°, and the deformation after rotation also becomes horizontal. In addition, the aerial photography device (such as an electro-optical pod) is bound to the drone, and the pitch angle reflects the angle between the aerial photography device and the horizontal line, which is generally negative downward. If the pitch angle is less than 30°, the target features seen by the aerial photography device in the forward view are no longer sufficient to express the target attributes. If the pitch angle is equal to 90°, it is a front view of the target, and a front view image is obtained. The pitch angle of the drone generally varies within a range of approximately plus or minus 10°. The method for correcting deformation caused by the roll angle is consistent with the method for correcting deformation caused by the pitch angle. The following embodiments are explained in detail using the pitch angle as an example.
[0033] Specifically, the selection of original feature points and target feature points in the image perspective transformation should follow the three core principles of non-collinearity, clear correspondence, and feature stability to ensure that the perspective transformation matrix can accurately map the aerial image to the target perspective. Generally, the original feature points select four non-collinear significant feature points from the aerial image, giving priority to satisfying the geometric distribution rules and anti-interference properties, and are usually selected from corner points, contour points, and reference object vertices. During the flight of the drone, the aerial image is composed of continuous image frames. There is a strong correlation between the contents of the aerial images, and the scale and perspective of the target in the aerial image continue to change. For example: scale difference: at a long distance, the image size of the target is small and the details are blurred. At a close distance, the details of the target are rich, but the field of view is narrow; posture deflection: under an oblique perspective, rectangular targets such as cars and building walls may appear trapezoidal or irregular quadrilateral distorted, and the vehicle may appear in different postures facing the lens in different directions. In addition, the background image in which the target is located is also an influencing factor of target detection, and the scale and posture will also change. Therefore, the original feature points of the perspective transformation can be selected from the entire image, that is, the upper left, upper right, lower left, and lower right corner points of the aerial image.
[0034] Due to the three-dimensional changes in the target's flight angle, its features are no longer limited to a single change in a single directional axis. Projected onto a plane, they are reflected as a comprehensive diagonal change of the original feature points, which can be positive or negative. Positive changes can stretch the compressed size of a distant target to its normal proportions, such as correcting a small vehicle at a distance to a close-up, proportional frontal image. They can also correct previously converged states (such as road edges) to parallel states. They can also enhance target features. Although the original image of a distant target has a low resolution, the transformed image can be improved through interpolation algorithms to provide a more regular image foundation for subsequent analysis. Negative changes can convert irregular perspective distortions, such as trapezoids and rhombuses, into regular geometric shapes (such as rectangles and squares). For example, buildings photographed at a close distance appear to be a trapezoid that is "wide at the top and narrow at the bottom" due to the bird's-eye view, but after the perspective transformation, they are restored to a horizontal, upright rectangle; the relative proportional relationship between the various parts of the target can be restored. For example, a close-range vehicle has the front part stretched and the rear part compressed due to the side view, but after the perspective transformation, it can be corrected to a front view so that the length and width ratio of the vehicle body is consistent with the actual physical size; the target's detailed features can be enhanced. The close-range target itself contains richer texture information. The aerial image after the perspective transformation is presented in an upright manner, which can avoid detail occlusion or deformation caused by the tilt of the viewing angle, and help improve the target recognition accuracy.
[0035] For example, Figure 2 and Figure 3The following are schematic diagrams of positive and negative perspective transformations, respectively. In the figure, the green box represents the original, unprocessed aerial image. A0, B0, C0, and D0 are the four original feature points (corner points) in the aerial image. w0 and h0 are the width and height of the aerial image, respectively. w and h represent the width and height of the aerial image after perspective transformation, which are the trapezoidal parts with black borders in the figure. Their coordinates are A1, B1, C1, and D1, respectively. The four original feature points A0, B0, C0, and D0 in the aerial image are transformed into four target feature points A1, B1, C1, and D1 after perspective transformation. A, B, C, and D are the image regions cropped from the perspective-transformed aerial image, which are the red box parts in the figure. d represents the horizontal transformation parameter of the original feature points after perspective transformation, and y represents the vertical transformation parameter of the original feature points after perspective transformation.
[0036] Furthermore, the coordinates of the original feature points and the target feature points are expressed with the point where D0 or D1 is located as the origin. Below, the coordinates of the four original feature points are expressed as: D0 (0, 0), C0 (w0, 0), B0 (w0, h0), A0 (0, h0) with D0 as the origin.
[0037] Figure 2 The coordinates of the four target feature points in the forward perspective transformation can be expressed as: D1 (0, 0), C1 (w0, 0), B1 (w0+d, h), A1 (-d, h).
[0038] Based on the change in the distance between the different original feature points and the distance between the corresponding target feature points, the scale transformation relationship of the aerial image before and after perspective transformation is determined. That is, based on the change between the width w0 / height h0 of the aerial image and the corresponding width w / height h after perspective transformation, the scale transformation relationship of the aerial image before and after perspective transformation is determined as follows: , , .
[0039] Figure 3 The coordinates of the four target feature points in the negative perspective transformation can be expressed as: D1(0,0), C1(w0,0), B1(w0-d,h0-y), A1(d,h0-y).
[0040] The scale transformation relationship of the aerial image before and after the perspective transformation in the negative perspective transformation is as follows: , , .
[0041] Furthermore, in order to adapt to the perspective transformation of aerial images caused by the pitch angle change during the actual inspection process of the drone, the value range of the transformation parameters is determined. Figure 4 As shown in the figure, point P is the current position of the aerial photography device, PL represents the flight direction of the UAV, point Q is the ground projection corresponding to point P, point O is the intersection of the imaging centerline of the aerial photography device and the ground, points A, B, C, and D are the intersection points of the camera's horizontal and vertical fields of view with the ground under the current attitude flight perspective, that is, the corner points of the ground's field of view area, points E and F are the intersection points of the UAV's flight direction and the field of view area, PF represents the distance from the current position of the aerial photography device to the farthest vertical field of view area, PE represents the distance from the current position of the aerial photography device to the nearest vertical field of view area, PQ represents the height of the aerial photography device from the ground, ∠LPO represents the pitch angle of the aerial photography device, ∠FPE represents the vertical field of view angle of the aerial photography device, and ∠BPC represents the horizontal field of view angle of the aerial photography device.
[0042] If the pitch angle of the current position of the aerial photography device is θ, the following relationship can be derived: , , , , Then we can determine the geometric relationship between the pitch angle and the field of view: .
[0043] For example, if the pitch angle θ=90°, the vertical field angle ∠FPE of the aerial photography device=60°, the horizontal field angle ∠BPC of the aerial photography device=60°, and the field width BC of the front view image=1080px (pixels), according to the proportional relationship, it can be calculated that when the pitch angle θ=30°, the field width BC of the front view image=5341px; when the pitch angle θ=45°, the field width BC of the front view image=3612px; when the pitch angle θ=60°, the field width BC of the front view image=1869px.
[0044] During actual flight, the pitch angle θ is usually around 45°, so the default horizontal resolution of aerial images at this time is 1080px. Based on the proportional relationship, when the pitch angle θ = 30°, the horizontal resolution of the aerial image is 1596px; when the pitch angle θ = 60°, the horizontal resolution of the aerial image is 558px; and when the pitch angle θ = 90°, the horizontal resolution of the aerial image is 322px.
[0045] It's important to note that when the drone's flight altitude changes, the projection of the inherent geometric shape of ground objects in the aerial imagery remains relatively stable. The significant changes are in the spatial resolution and scene coverage caused by the imaging system's field of view. This is because within the typical aerial photography altitude range, which is far above the size of ground objects, altitude changes primarily cause scaling between the imaging system and the scene, rather than perspective distortion. The outlines and relative proportions of ground objects in aerial images are largely determined by their actual geometry and orientation. Increasing or decreasing altitude primarily causes the objects to appear proportionally larger or smaller in the aerial imagery, while the projected geometric properties of their shape features (such as the rectangular outline of a building, the linear structure of a road, and the aspect ratio of a vehicle) remain relatively consistent. Furthermore, the above calculation process only calculates a range of values for the transformation parameter (d). The subsequent determination of the perspective transformation matrix uses values within this range. Therefore, the effect of the aerial photography altitude on the transformation parameters during different flight processes can be ignored.
[0046] From the above scale transformation relationship, it can be seen that the transformation parameter is half the number of pixels of the resolution of the aerial image in the horizontal direction, that is: .
[0047] From this calculation, we know that for a negative perspective change of 30°, the transformation parameter d = -258; for a positive perspective change of 60°, the transformation parameter d = 261; and for a positive perspective change of 90°, the transformation parameter d = 379. Therefore, the actual pitch angle range during flight is between [-30° and -90°], so the range of the above transformation parameter d is [-258, 379].
[0048] Step 103: Determine multiple perspective transformation matrices by selecting values within the range of the transformation parameters. In this embodiment, the coordinates of the original feature points and the target feature points are determined based on the relationship between the transformation parameters and the scale transformation. Based on the coordinates of the original feature points and the target feature points, the perspective transformation matrix corresponding to the transformation parameters is determined.
[0049] Specifically, the coordinates of the original feature points and the target feature points are determined according to the relationship between the transformation parameter d and the scale transformation, that is, Figure 2 or Figure 3 The coordinates of the four sets of feature point pairs A0, B0, C0, D0, A1, B1, C1, and D1.
[0050] For example, assuming that the coordinates of an original point on the aerial image are (x, y, z), the original point is 3 is used for perspective transformation to obtain the first coordinate (X, Y, Z), and the first coordinate is normalized according to the value of Z to obtain the homogeneous coordinate (X', Y', 1) after normalization of the first coordinate.
[0051] The process of the first coordinate normalization is as follows: .
[0052] When homogeneous coordinates , then the point is the two-dimensional plane coordinate of the original point after the perspective transformation, and the following relationship can be obtained: , When settlement is performed, let , and the following is obtained: , In the formula, , , , , , , , , represent 3 9 elements in the perspective transformation matrix of 3.
[0053] In the above expanded equation, there are a total of 8 unknowns, which can be solved by using the coordinates of the original feature points and the target feature points obtained above to solve the values of the elements , , , , , , , in the perspective transformation matrix, thereby obtaining the perspective transformation matrix.
[0054] It should be noted that when the values in the value range of the transformation parameters are taken, they can be randomly taken or user-specified, uniformly taken, etc. Preferably, the values are taken relatively dispersedly in the value range.
[0055] Exemplarily, the transformation parameter d is valued according to a uniform distribution, a plurality of groups of coordinates of the original feature points and the target feature points are determined by the above method, and the values of the elements in different perspective transformation matrices are solved, thereby obtaining a plurality of perspective transformation matrices.
[0056] Step 104: Use multiple perspective transformation matrices to perform perspective transformation on the aerial images in the original image sample set to obtain multiple transformed images. In the embodiment of the present application, the aerial images are perspective transformed by the cv2.warpPerspective function (a function used to perform perspective transformation on an image). The cv2.warpPerspective function traverses each pixel in the aerial image, calculates the coordinates after perspective transformation based on the perspective transformation matrix, and fills the pixel values through the interpolation algorithm to generate a new image with distortion eliminated. The new image is then cropped to obtain a transformed image. As shown in the figure, Figure 5 is the original aerial image, Figure 6 for Figure 5 The transformed image after perspective transformation.
[0057] It should be noted that the number of aerial images corresponding to perspective transformation under each perspective transformation matrix should be similar or equal, that is, the number of transformed images obtained after perspective change of each perspective transformation matrix should be similar or equal.
[0058] Step 105: Update the target annotation frame in the transformed image using coordinate mapping to calibrate the annotation information. In an embodiment of the present application, updating the target annotation frame in the transformed image using coordinate mapping includes: determining the coordinates of the new target annotation frame in the transformed image based on the corner point coordinates of the original target annotation frame and the perspective transformation matrix. Based on the interaction between the original target annotation frame and the new target annotation frame, determine the new target annotation frame with the edges cropped. Delete unqualified target annotation frames based on the area and aspect ratio of the original target annotation frame and the new target annotation frame. Update the target annotation frame and annotation information of the transformed image based on the remaining new target annotation frames and their category information.
[0059] Specifically, updating the annotation information of the transformed image after perspective transformation is a key step to ensure that the annotation information is aligned with the transformed image space. The core is to convert the original target annotation box in the aerial image into a new target annotation box in the transformed image through a coordinate mapping algorithm, and simultaneously update the annotation information stored in the XML file. The aerial images taken during the flight of the drone are continuous, and the position of the target in each aerial image is not fixed. After perspective transformation and image cropping, the target at the edge position is most affected, which may cause Figure 7 The six cases shown in . Figure 7 In the figure, the dotted line represents the cropping line, the black and white parts together are the original target annotation box, the black part on one side of the cropping line represents the original target annotation box that is cropped, and the white part represents the constructed new target annotation box. A point in the figure represents the original center point, that is, the intersection of the diagonal lines of the original target annotation box.
[0060] For the above six situations, a new target annotation box is first constructed based on the interactive relationship between the original target annotation box and the new target annotation box, and then the unqualified target annotation box is deleted according to the area and aspect ratio of the original target annotation box and the new target annotation box.
[0061] In an embodiment of the present application, based on the interactive relationship between the original target annotation frame and the new target annotation frame, a new target annotation frame with the edge cropped is determined, including: if in the transformed image, the three corner points of the original target annotation frame are cropped and the original center point of the original target annotation frame is not on the cropping line, then the original target annotation frame is deleted. If in the transformed image, the three corner points of the original target annotation frame are cropped and the original center point of the original target annotation frame is on the cropping line, then the new target annotation frame is constructed with the original center point as the corner point. If in the transformed image, two corner points of the original target annotation frame are cropped, then the remaining original target annotation frame is used as the new target annotation frame. If in the transformed image, one corner point of the original target annotation frame is cropped, then the new target annotation frame is constructed with the intersection of the diagonal line of the original target annotation frame and the cropping line as the corner point.
[0062] For example, when constructing a new target annotation box, the four sides of the new target annotation box should be parallel to or coincide with the four sides of the original target annotation box.
[0063] In an embodiment of the present application, deleting an unqualified target marking frame based on the area and aspect ratio of the original target marking frame and the new target marking frame includes: if the area of the new target marking frame reaches half the area of the original target marking frame, comparing the aspect ratio. If the ratio of the aspect ratio of the new target marking frame to the aspect ratio of the original target marking frame is within a preset range, retaining the new target marking frame. If the ratio of the aspect ratio of the new target marking frame to the aspect ratio of the original target marking frame is not within the preset range, deleting the new target marking frame. If the area of the new target marking frame does not reach half the area of the original target marking frame, deleting the new target marking frame.
[0064] Exemplarily, the preset range is set based on the range of aspect ratios of original target annotation boxes of aerial images in the original image sample set, or may be set based on user requirements for different applications.
[0065] Exemplarily, the preset range is [0.8, 1.2], that is, the aspect ratio of the new target annotation box is the product of [0.8, 1.2] and the aspect ratio of the original target annotation box.
[0066] The four corner points of the original target annotation box are mapped using the perspective transformation matrix to obtain the four corner points of the new target annotation box. In addition, the new target annotation box can be further optimized using the minimum enclosing rectangle algorithm to ensure that it can completely enclose the target after perspective transformation.
[0067] Because of the randomness of the transformation parameters and the difference in the position of the target in the image, the perspective transformation will cause some targets to exceed the image boundary, which needs to be deleted because such targets can no longer fully express the characteristics of the target. Perspective transformation will cause some targets to appear the phenomenon of aspect ratio distortion, which needs to be deleted because such targets can no longer accurately express the characteristics of the target. Perspective transformation will cause some target features to appear blurred, which needs to be deleted because such target features will reduce the robustness of the characteristics.
[0068] Further, the new target bounding box can also be judged for abnormality, that is, the gradient, size, and aspect ratio of the new target bounding box are determined, and it is judged whether they are within the effective range. If the gradient, area, and aspect ratio of the new target bounding box are all within the effective range, the new target bounding box is retained. Otherwise, the new target bounding box and its labeling information are deleted. The labeling information of the remaining new target bounding box is counted, and is balanced according to the distribution of the types of targets.
[0069] The above abnormality judgment can be realized by a multi-scale detection head network. Specifically, the multi-scale detection head network is used to extract the characteristics of the target, and the size range of the original target bounding box in the aerial image is calculated. If the new target bounding box is not within the range, the new target bounding box and its labeling information are deleted. The gradient and aspect ratio of the target bounding box can also be calculated using this method, or using the labeling information of the original target bounding box. Here, details are not repeated.
[0070] According to the coordinate information of the updated new target bounding box and the target type (the target type of the original target bounding box is used), the labeling information is updated, and is stored as an XML format file, to ensure the spatial consistency of the image data. The labeling information includes target bounding box coordinates, image size information, target set information, target category information, etc.
[0071] Step 106: Combine the labeled aerial image and its labeling information with the transformed image and its calibrated labeling information to obtain an enhanced sample set. In the embodiment of the present application, the transformed image and its corresponding labeling information after perspective transformation and labeling information calibration are combined with the original image sample set and its corresponding labeling information to form an enhanced sample set.
[0072] Specifically, the construction process is divided into two cases: The first case: pitch angle change sample.
[0073] In the flight process of fixed-wing unmanned aerial vehicles, the flight path is generally pre-set. Except for the normal pitch angle changes in the take-off and landing stages, the pitch angle changes in the cruising stage are mainly caused by environmental factors such as atmospheric turbulence or wind shear. There are relatively many targets in the flight process, and the forms of the targets are also relatively many. In addition, since the pitch angle changes can be positive or negative during flight, perspective transformation not only needs randomness, but also needs to maintain the uniform distribution of perspective transformation. Therefore, in order to better match different forms under different pitch angles, the range of image perspective transformation parameters can be appropriately expanded, avoiding excessive positive or negative changes that cause angle imbalance, and avoiding target missed detection caused by insufficient pitch angle adaptation.
[0074] The second case: roll angle change sample.
[0075] In the flight process of fixed-wing unmanned aerial vehicles, the roll angle changes in the cruising stage are mainly caused by turbulence disturbance and turning needs. Since conventional oil and gas pipeline laying is straight-line dominant, in the inspection process, in order to avoid terrain, obstacles and other needs to fly turn, the target under the roll state appears relatively less, so the range of image perspective transformation parameters in this process can be appropriately reduced, avoiding false detection caused by excessive background transformation.
[0076] In addition, the balance of different categories of target data in the target detection model training is the key prerequisite to ensure the generalization ability and classification fairness of the target detection model. The core goal is to eliminate the model bias caused by the difference in the number of samples, and to ensure stable learning of both minority and majority categories of targets. Perspective transformation will increase or decrease the number of samples of different categories, so when constructing the enhanced sample set, the number of samples of different categories needs to be counted, and sample balancing needs to be performed according to the data balancing strategy, or the number or proportion of samples of different categories of targets in the enhanced sample set needs to be adjusted according to the needs, in order to meet the needs of different feature extraction or training.
[0077] The enhanced sample set is used to train the target detection model, such as a convolutional neural network model based on deep learning. The forward propagation calculates the prediction results, and the network parameters are updated through back propagation. During the training process, the performance is evaluated regularly on the validation set, and multiple local optimal models with the highest validation set accuracy are saved. Finally, the global optimal model is taken.
[0078] In order to improve the recognition accuracy and generalization ability of the target detection model, the training parameters need to be adjusted constantly, and the test set needs to be used to evaluate the trained target detection model offline. Considering that the risk needs to be warned in the actual inspection process, the detection rate is taken as an important indicator of the target detection model. The detection rate represents the proportion of targets whose position and category are both detected correctly among all targets.
[0079] To validate the effectiveness of this method, a base sample set consisting of 8,884 training and 1,480 validation images was selected from actual flight videos. The YOLOv5x network model was trained using this base sample set to obtain a base model. Using this method, the base sample set was then perspective-transformed using random transformation parameters to form an enhanced sample set, which was then used to train the YOLOv5x network model to obtain an enhanced model. The object categories in both sample sets included "person," "car / pickup," "truck," "tanker," "engineering vehicle," "bus," "tricycle," and "building." The sample scenes were collected based on oil and gas pipeline inspections and included forests, roads, parking lots, farmland, and factory areas.
[0080] Two test sets were constructed: Test Set 1, consisting of 198 images with scenes roughly comparable to the training samples (basic and enhanced sets), and Test Set 2, consisting of 11 images with scenes completely unrelated to the training samples. The testing process was based on the majority of targets encountered during actual flight, primarily focusing on small "car" targets and large "truck" targets. The training and test results are shown in Table 1 below: Table 1 Test results of the basic model and enhanced model under different test sets
[0081] In the table, Map@0.5 represents the mean average precision. The detection result is considered correct only when the intersection-over-union ratio of the predicted box and the true box is ≥0.5.
[0082] From the above detection results, it can be seen that the enhanced model of this application has a greatly improved detection effect compared with the basic model. It not only has good detection effects on small and large targets, but also has strong generalization capabilities for unknown scenes.
[0083] Although this application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on routine or non-creative work. The order of steps listed in this embodiment is only one way of executing the steps among many, and does not represent the only execution order. When an actual device or client product executes, the method shown in this embodiment or the accompanying drawings may be executed sequentially or in parallel (for example, in a parallel processor or multi-threaded processing environment).
[0084] like Figure 8As shown, the embodiment of the present application further provides a sample enhancement device 800. The device includes: a construction module 801, a transformation parameter module 802, a perspective transformation matrix module 803, a perspective transformation module 804, a calibration module 805 and a combination module 806, as follows.
[0085] The construction module 801 is used to construct an original image sample set based on the labeled aerial images.
[0086] The transformation parameter module 802 is used to determine the value range of the transformation parameter according to the scale transformation relationship of the aerial image before and after the perspective transformation, the geometric relationship between the shooting parameters of the aerial image, and the range of the target angle.
[0087] The perspective transformation matrix module 803 is used to determine multiple perspective transformation matrices by taking values within a range of transformation parameters.
[0088] The perspective transformation module 804 is used to perform perspective transformation on the aerial images in the original image sample set using multiple perspective transformation matrices to obtain multiple transformed images.
[0089] The calibration module 805 is used to update the target annotation box in the transformed image using coordinate mapping to calibrate the annotation information.
[0090] The combining module 806 is used to combine the labeled aerial image and its labeling information with the transformed image and its calibrated labeling information to obtain an enhanced sample set.
[0091] Some modules in the apparatus described herein may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0092] The devices or modules described in the above application embodiments can be implemented by computer chips or physical devices, or by products with certain functions. For ease of description, the above devices are described separately by function in various modules. When implementing the embodiments of this application, the functions of each module can be implemented in the same or multiple software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.
[0093] The methods, devices, or modules described herein can be implemented in the form of computer-readable program code. The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller in pure computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the means for implementing various functions may be considered to be both a software module for implementing the method and a structure within a hardware component.
[0094] An embodiment of the present application further provides a device comprising: a processor; a memory for storing processor-executable instructions; and when the processor executes the executable instructions, the method described in the embodiment of the present application is implemented.
[0095] The embodiments of the present application also provide a non-volatile computer-readable storage medium having a computer program or instruction stored thereon. When the computer program or instruction is executed, the method described in the embodiments of the present application is implemented.
[0096] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist independently, or two or more modules may be integrated into one module.
[0097] The above-mentioned storage media include, but are not limited to, random access memory (RAM), read-only memory (ROM), cache, hard disk drive (HDD), or memory card. Such memory can be used to store computer program instructions.
[0098] Through the description of the above implementation methods, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, or can be embodied through the implementation process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application or certain parts of the embodiments.
[0099] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. All or part of this application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.
[0100] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some or all of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present application.
Claims
1. A sample enhancement method, characterized in that: include: Constructing a set of original image samples based on labeled aerial images; Determine the value range of the transformation parameters based on the scale transformation relationship of the aerial image before and after the perspective transformation, the geometric relationship between the shooting parameters of the aerial image, and the range of the target angle; Taking values within the range of values of the transformation parameters to determine a plurality of perspective transformation matrices; Performing perspective transformation on the aerial images in the original image sample set using the plurality of perspective transformation matrices to obtain a plurality of transformed images; updating the target annotation box in the transformed image using the coordinate mapping to calibrate the annotation information; The labeled aerial image and its labeling information are combined with the transformed image and its calibrated labeling information to obtain an enhanced sample set.
2. The method according to claim 1, characterized in that Before constructing the original image sample set based on the labeled aerial images, the method further includes: Cleaning the aerial image to remove noise; and / or, Dehazing and enhancing aerial images to improve their clarity; and / or, Standardizing the capture parameters of aerial images; and / or, Perform geometric correction and segmentation on aerial images.
3. The method according to claim 1, characterized in that The method for determining the scale transformation relationship of the aerial image before and after the perspective transformation includes: Determine multiple original feature points in the aerial image, and obtain multiple target feature points after performing perspective transformation on the aerial image; According to the change of the distance between different original feature points and the distance between the corresponding target feature points, the scale transformation relationship of the aerial image before and after the perspective transformation is determined.
4. The method according to claim 1, wherein The method for determining the geometric relationship between the shooting parameters of the aerial image includes: The shooting parameters include the field of view angle, field of view area, height and attitude angle of the aerial photography device; The geometric relationship between the pitch angle / roll angle and the field of view is determined based on the relationship between the distance between the aerial photography device and the field of view area and the field of view angle.
5. The method according to claim 1, characterized in that The determining of the value range of the transformation parameter includes: The value range of the transformation parameter is determined according to the range of the target angle, the geometric relationship and the scale transformation relationship.
6. The method according to claim 3, characterized in that The determining of multiple perspective transformation matrices within the range of the transformation parameters includes: Determine the coordinates of the original feature point and the target feature point according to the transformation parameter and the scale transformation relationship; According to the coordinates of the original feature point and the target feature point, a perspective transformation matrix corresponding to the transformation parameters is determined.
7. The method according to claim 1, characterized in that The updating of the target annotation box in the transformed image by using coordinate mapping includes: Determine the coordinates of the new target annotation box in the transformed image based on the corner point coordinates of the original target annotation box and the perspective transformation matrix; Based on the interaction between the original target annotation frame and the new target annotation frame, a new target annotation frame with the edge clipped is determined; Delete unqualified target annotation frames based on the area and aspect ratio between the original target annotation frame and the new target annotation frame; The target annotation box and annotation information of the transformed image are updated according to the remaining new target annotation box and its category information.
8. The method according to claim 7, characterized in that The determining of the new target annotation frame with the edge clipped based on the interaction relationship between the original target annotation frame and the new target annotation frame includes: If the three corner points of the original target annotation box are clipped in the transformed image, and the original center point of the original target annotation box is not on the clipping line, then the original target annotation box is deleted; If the three corner points of the original target annotation box are cropped in the transformed image, and the original center point of the original target annotation box is on the cropping line, then a new target annotation box is constructed with the original center point as the corner point; If two corner points of the original target annotation box are cropped in the transformed image, the remaining original target annotation box is used as the new target annotation box; If a corner point of the original target annotation box is cropped in the transformed image, a new target annotation box is constructed with the intersection of the diagonal line of the original target annotation box and the cropping line as the corner point.
9. The method according to claim 7, characterized in that The step of deleting unqualified target marking frames according to the area and aspect ratio of the original target marking frame and the new target marking frame includes: If the area of the new target annotation box reaches half the area of the original target annotation box, then the aspect ratio is compared; If the aspect ratio of the new target annotation frame is within a preset range compared to the aspect ratio of the original target annotation frame, the new target annotation frame is retained; If the aspect ratio of the new target annotation frame is not within the preset range compared to the aspect ratio of the original target annotation frame, the new target annotation frame is deleted; If the area of the new target labeling box does not reach half of the area of the original target labeling box, the new target labeling box will be deleted.
10. A sample enhancement device for implementing the method according to any one of claims 1 to 9, characterized in that: include: A construction module, used to construct an original image sample set based on the labeled aerial images; A transformation parameter module is used to determine the value range of the transformation parameter based on the scale transformation relationship of the aerial image before and after the perspective transformation, the geometric relationship between the shooting parameters of the aerial image, and the range of the target angle; A perspective transformation matrix module, configured to determine a plurality of perspective transformation matrices by taking values within a range of values of the transformation parameters; a perspective transformation module, configured to perform perspective transformation on the aerial images in the original image sample set using a plurality of perspective transformation matrices to obtain a plurality of transformed images; a calibration module, configured to update the target annotation frame in the transformed image using coordinate mapping to calibrate the annotation information; The combination module is used to combine the labeled aerial image and its annotation information with the transformed image and its calibrated annotation information to obtain an enhanced sample set.
Citation Information
Patent Citations
Target detection method based on data enhancement
CN109063748A
Method and equipment for transversely and longitudinally positioning and ranging rear vehicle
CN112070839A
Splicing method and device for aerial images of multiple unmanned aerial vehicles
CN112288634A
A method for data enhancement of image
CN112396569A
Transformer-based multi-view target detection method and system
CN113673425A