Device and System for Identifying Rural Dwellings under the Vision of Unmanned Aerial Vehicles
Patent Information
- Application Number
- CN202410948045.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2044-07-08
AI Technical Summary
显然,该方案应用到无人机系统中,要求无人机视频画面为固定的同一相对方位角度,才有可比性,而这无疑大大限制了民居尺寸的可量测性
[0055] Compared with the prior art, the present invention has the following advantages: Addressing the unpredictable relative orientation relationship between the attitude and orientation of dwellings when UAVs collect images, the present invention analyzes the characteristic representation of dwelling dimensions in image space within rural dwelling images. It decomposes dwelling dimensions into three dimensions: length, width, and height. It proposes using two opposing sidelines that respectively sandwich a pair of long, wide, and high sides to describe the dimensions of dwellings in these three directions. Using four geometric features—the slope of the two opposing sidelines in any direction, and the perpendicular distance from the image center point to the two sidelines—as constraints, an artificial neural network is established to predict dwelling dimensions. The collected samples are used to identify and train the nonlinear mapping relationship between the size of rural dwellings and the proposed four geometric features. Each sample can identify rural dwellings by classifying and recognizing targets in aerial images. Through the extraction and edge detection of dwellings, three pairs of edges corresponding to the length, width and height of the dwelling contour are obtained by Hough transform. The four geometric features are extracted from two edges of any pair of edges. The results are input into the trained dwelling size prediction model to predict the length, width and height of rural dwellings online. The neural network-based residential building size prediction model established in this invention has good generalization performance. It can accurately predict the three-dimensional dimensions of various types of residential buildings based on images of rural houses collected from uncertain flight paths. It also has a wide range of applications. Since it extracts features from straight lines without relying on precise pixel start and end points, the prediction accuracy is high. At the same time, due to the robustness of the established model, the step size in the cumulative matrix of the Hough line transform does not need to be set too small when extracting the edge lines of residential buildings. Furthermore, the identification of the correspondence between edge points and straight line elements in the matrix is simplified, reducing the amount of computation and speeding up the search, thereby further improving the real-time performance of rural residential building identification. Based on heuristic search, it can automatically find and identify the horizontal edges and vertical side edges of the ground and roof of residential buildings, simplifying the operation confirmation steps.
Smart Images

Figure CN118918464B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of application technology of drones in rural transportation, and specifically relates to a device and system for identifying rural dwellings under the view of a drone. Background Technology
[0002] The application of drones in rural transportation has broad potential and practical value. In rural areas lacking medical resources, drones can rapidly transport emergency medicines, vaccines, blood, and medical equipment, and can also provide telemedicine services to remote areas, quickly delivering medical samples and equipment. Integrating specialized accessories, drones can quickly collect soil, water, and crop samples in farmland, supporting precision agriculture and crop monitoring. In natural disasters or emergencies, drones can rapidly transport and deliver relief supplies, quickly reaching disaster areas to conduct disaster assessments and photography, helping to develop relief plans. In environmental monitoring and protection, drones are used to monitor forests, rivers, wetlands, and other ecological environments, collecting data on vegetation cover, water quality changes, etc.; they can also monitor illegal logging and poaching activities, protecting natural resources and the ecological environment.
[0003] In terms of infrastructure inspection and maintenance, drones are used to inspect infrastructure such as power lines, communication base stations, and bridges to promptly identify and address problems; they are also used to conduct regular inspections of rural houses and public facilities to ensure safety and maintenance needs.
[0004] In e-commerce logistics, drones can provide rapid delivery services for small parcels, letters, and express packages, improving rural logistics efficiency; and transport fresh food and daily necessities for farmers and residents, enhancing convenience. In the future, they are expected to support the logistics and distribution of rural e-commerce platforms, shortening the circulation time of agricultural products and daily necessities; simultaneously, they will provide reverse logistics transportation, offering farmers convenient sales channels for agricultural products and quickly transporting them to urban markets.
[0005] Drone delivery offers advantages such as high efficiency and wide coverage. Drones can quickly reach their destinations, saving time, especially in rural areas with poor transportation; and because they can cover remote or inaccessible rural areas, they can provide comprehensive services.
[0006] In rural transportation, whether for normal landings or emergency evasive landings, drone docking must consider safety, convenience, and efficiency. When selecting suitable docking points in rural residential areas to avoid houses, precise measurements of the house's length, width, and height are essential for navigation planning. On one hand, when setting the drone's minimum flight altitude, the height of rural houses needs to be determined; on the other hand, when circling houses or landing on rooftops, the dimensions and length of the houses must be measured to ensure safe lateral and longitudinal distances.
[0007] Several existing technologies exist for measuring buildings. For example, patent application number 201811029714.X discloses a spatial measurement method for intelligent buildings and surveying. A laser emits a light pulse towards the target, and the distance is obtained from the flight time required for the object to reflect the light. Then, a camera captures a measurement image to calculate the distance between two points in the image. This method requires the camera's optical axis to be aligned with the target plane and can only measure the distance between two points on that single-direction target plane. Patent application number 202310233220.8 discloses a method for identifying, acquiring, and calculating the ratio of length units to pixel values in CAD drawings. It uses AI to automatically acquire pixel values on CAD drawings and the corresponding number of pixels, thereby calculating the ratio of pixel values to the actual space length. This method automatically calculates the length of the actual space represented by pixels in architectural CAD images. Clearly, this method requires a priori knowledge of the pixel-to-actual-space ratio, meaning an actual length needs to be measured or marked beforehand. The invention patent application with application number 202310416333.1 discloses a method for on-site surveying of photovoltaic power stations. It uses any combination of tools, including infrared levels, measuring tapes, tape measures, and laser rangefinders, to measure and cross-check the actual area and dimensions of the roof, making it easier for users to determine whether a photovoltaic power station needs to be installed. Through on-site photography, on-site measurement, and review by the back-end system, it determines whether the roof is suitable for installing a photovoltaic power station based on the surrounding obstacles and the roof conditions. It is evident that this application also requires the use of actual measurement methods to obtain the dimensions of the roof.
[0008] The invention patent application with application number 202110395821.X discloses a method for monitoring building height based on high-resolution optical remote sensing satellite images with corner points. It uses spatial distribution to pair corner points and then uses a geometric model of shadow imaging to calculate the height of the building. It utilizes the geometric relationship between the total length of the building shadow on the remote sensing image, the solar altitude angle, the solar azimuth angle and the building height.
[0009] As can be seen from the above technical solutions, current image-based measurements of house dimensions either rely on prior data obtained from on-site measurements or are directly calculated based on pixel ratios. The latter depends on a precise proportional relationship, requiring the pixel width of the measured line to correspond to a predetermined, a priori proportion. Clearly, applying this approach to drone systems requires the drone video footage to be from a fixed, relative angle for comparability, which significantly limits the measurability of residential dimensions.
[0010] Therefore, there is a need for a device or system for identifying the length, width, and height dimensions of rural dwellings that has strong generalization capabilities and wide applicability. Summary of the Invention
[0011] The purpose of this invention is to provide a device, controller and system for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV), which is used to identify the features of rural dwellings based on images in the view of the UAV, and then calculate their height, as well as the length and width of the roof, etc., without requiring the UAV to maintain a certain precise orientation when sampling images.
[0012] To address the aforementioned technical problems, this invention provides a device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV), comprising a host unit and a user interface unit. The host unit includes an input / output module, a processor, and a storage module, and is configured as follows:
[0013] A rural dwelling image is acquired, and a Region of Interest (ROI) containing the dwelling is determined within the rural dwelling image. The dwelling area is extracted from the ROI as the dwelling target image. Edge detection is performed on the dwelling target image to obtain dwelling edge images. A Hough transform is applied to the dwelling edge images, and a heuristic search is used to obtain dwelling edge images including three pairs of sidelines in the dwelling outline edge. The first pair of sidelines consists of two opposing sidelines at the top and bottom of the dwelling, sandwiched between a pair of side edges. The second and third pairs of sidelines consist of two opposing sidelines sandwiched between a pair of horizontal long sides and a pair of horizontal wide sides, respectively. From the dwelling edge image, the geometric features of the two opposing sidelines relative to the image center are sequentially obtained for any pair of sidelines. These geometric features are input into a trained neural network-based dwelling size prediction model to obtain a predicted dwelling size value, which is then output.
[0014] Preferably, the user interface unit is provided with an operation panel, and the host unit is further configured to: extract the residential area from the ROI region as the foreground object, use the non-foreground object area in the rural residential image as the background object, and set all background objects in the rural residential image to zero to form the residential target image.
[0015] Preferably, the processor includes a general processing module, an acquisition and preprocessing module, a feature extraction module, an iterative learning module, and a residential building size prediction module.
[0016] The acquisition and preprocessing module acquires images of rural dwellings from the image acquisition unit. After filtering, based on deep learning features, a trained target classification model identifies the location of the dwellings in the image as Regions of Interest (ROIs). The portion outside the ROIs in the original rural dwelling image is removed to form a target image of the dwellings. Then, edge detection is performed to form a binarized edge image of the dwellings. After performing a Hough transform on the binarized edge image of the dwellings, the vertical and horizontal line segments with the longest lengths are selected as candidate line segments based on the polar coordinate space parameter features in the Hough transform. Based on the orientation features of the candidate line segments, two opposite lines with a pair of long, wide, and high sides sandwiched in the middle are determined as a pair of side lines for the length, width, and height of the dwellings, respectively, forming three pairs of side lines for the dwellings.
[0017] For each edge of any pair of the three pairs of edge lines, the feature extraction module combines the corresponding binarized residential edge image, describes it using the coordinates of the two endpoints of the line, and calculates the slope of the line and the perpendicular distance from the image center point to the line based on the two endpoints; the slope of the line on both sides of the residential building and the perpendicular distance from the image center point to the line are used as the geometric features of the residential building in the rural residential building image.
[0018] The iterative learning module trains the neural network of the residential building size prediction model in the residential building size prediction module based on the collected training samples.
[0019] The general processing module responds to the message, initiates the identification of dwellings in the acquired images, schedules the identification processing flow, and outputs the size data identified and predicted by the dwelling size prediction module.
[0020] Preferably, the acquisition and preprocessing module includes a target detection unit. Within this unit, a target classification model based on the ASFF-YOLOv5s network is established. After acquiring rural scene images using an overhead camera in the image acquisition unit, the network is trained offline using an image sample set consisting of these rural scene images labeled with rural residential areas to obtain the rural scene target classification model. During online operation, the acquisition and preprocessing module uses the rural scene target classification model to identify and process the rural scene images (including residential buildings) to be tested, obtaining anchor boxes for rural residential areas. The areas within these anchor boxes are designated as the ROIs (Regions of Interest) for rural residential areas. The ASFF-YOLOv5s network adds three layers (24, 25, and 26) between the 23rd layer and the layer below it in the YOLOv5s network. These three layers receive the outputs of layers 17, 20, and 23 as inputs, respectively. After feature fusion at different scales by the ASFF adaptive spatial feature fusion module, these layers sequentially replace the original layers 17, 20, and 23 as inputs to the original output layer 24, i.e., the 27th layer of the new network.
[0021] As a preferred approach, for residential building images captured based on a residential building model, the acquired images are filtered and grayscaled, then binarized to form the residential building target image. For rural residential building images acquired online, a target classification model based on the ASFF-YOLOv5s deep learning network is used to locate the ROI region of the residential building using anchor boxes. Pixels outside the obtained ROI region are then set to zero to form the residential building target image. Edge extraction can be directly performed on the binarized residential building target image. For color or grayscale residential building target images, the color image is first converted to a single-channel image, and then edge detection operators such as Sobel, Canny, and LOG are used to detect edges.
[0022] Preferably, the residential building size prediction model uses an artificial neural network. Its input layer receives four geometric feature inputs from the general processing module: the slopes k1 and k2 of a pair of sides corresponding to the length, width, or height of the residential building, and the perpendicular distances D1 and D2 from the image center point to the pair of sides. The output of the output layer is the length, width, or height of the residential building corresponding to the pair of sides.
[0023] Preferably, the host unit is further configured to: when performing the Hough transform on each edge pixel in the residential edge image, accumulate to obtain the corresponding ρ-θ plane matrix of the image, where ρ represents the distance of the perpendicular segment of the line from the origin, and θ is the angle of the perpendicular line to the horizontal axis; in the accumulation, the foreground pixel (u, v) in the image is connected to a straight line element l i,j (ρ i,j θ i,j The corresponding judgment method is:
[0024] Take a fulcrum H(ρ) on this straight line i,j ·cosθ i,j , ρ i,j ·sinθ i,j Let H(u′, v′) be the variable, and calculate J = |vv′-(uu′)·tanθ. i,j If |J|≤ε, then the pixel is determined to be on the straight line and the sum is accumulated, where ε is a preset small positive number.
[0025] Preferably, the host unit is further configured to: when performing the Hough transform on each edge pixel in the residential edge image, accumulate to obtain the ρ-θ plane matrix corresponding to the image, where ρ represents the distance of the perpendicular line segment from the origin to the straight line, and θ is the angle of the perpendicular line to the horizontal axis.
[0026] Based on the accumulated ρ-θ plane matrix, within a range close to the vertical direction, search for the two or three longest edges as the side edges of the dwellings;
[0027] Within the range outside the side edge direction, the line segments are sorted from largest to smallest according to the values of the planar matrix elements, i.e., their lengths. Starting with the longest main line segment, two secondary line segments with similar directions are sequentially searched. These main line segments and their corresponding two secondary line segments form a candidate line group. For any line segment in the candidate line group, the distance to the upper and lower endpoints of the side edge of the dwelling is calculated. If the minimum value of this distance corresponds to the distance between the line segment and the upper endpoint of a side edge, then the line segment is marked as a candidate line for the roof; otherwise, it is marked as a candidate line for the ground. Multiple candidate line groups are obtained by iteratively selecting lines from different directions.
[0028] Preferably, the host unit is further configured to: in one of the candidate straight line groups, form the first pair of side lines with two opposing side lines consisting of a candidate straight line for a residential roof and a candidate straight line for the ground sandwiched between a pair of side edges; in one of the candidate straight line groups, form the second pair of side lines and the third pair of side lines with two opposing side lines consisting of two candidate straight lines for residential roofs sandwiched between a pair of other straight line segments; or, in the side edges of the residential buildings, form the second pair of side lines and the third pair of side lines with two opposing side lines consisting of two side edges of the residential buildings sandwiched between a pair of other straight line segments.
[0029] Preferably, the host unit is further configured to: calculate and extract geometric features for any pair of the three pairs of edges.
[0030] Assuming the coordinates of the two endpoints of a residential building's boundary line are (x1, y1) and (x2, y2), then its slope is k = y1 - y2 / x1 - x2;
[0031] The perpendicular distance from the center of the image to the edge of the house is
[0032] In the formula, (x0, y0) represents the coordinates of the center of the image, A and B are the coefficients of the general form of the equation of a straight line, and C is a constant term and: A = y2 - y1, B = x1 - x2, C = -x1(y2 - y1) + y1(x2 - x1);
[0033] Based on the calculation formulas for k and d, the slopes k1 and k2 of any pair of edges of the dwellings, as well as the perpendicular distances D1 and D2 from the image center point to the two edges, are calculated as four quantities, which are used as the geometric features of the dwellings in the rural dwelling image.
[0034] Preferably, the residential building size prediction model uses a BP neural network, which has a three-layer structure including one hidden layer.
[0035] The output of the j-th node in the hidden layer is
[0036] The output of the first node in the output layer is
[0037] In this network training, f() is taken as the tansig function, wij and vj are the connection weights from the input layer to the hidden layer and from the hidden layer to the output layer, respectively, aj and b are the thresholds of the hidden layer and the output layer, m=4 is the number of input variables, k is the number of hidden layer nodes, and gradient descent is used for network training.
[0038] Preferably, the host unit is further configured to: before online application, first use a drone to collect images of the constructed residential building model in offline state and form a sample set, and use the sample set to train the residential building size prediction model; the residential building model adopts a cuboid, and the size of the cuboid is a first preset multiple S1 of the actual residential building size, the multiple S1 being set according to the drone's rural cruise flight altitude and camera resolution; the different cuboids used in the residential building model correspond to a size range of S1 times the common size range of rural residential buildings, and the size difference between different cuboids corresponds to S1 times the actual residential building size difference SD.
[0039] Preferably, the first preset multiple S1 is between 1 / 50 and 1 / 20, and the actual residential building size difference SD is between 0.5 and 3 meters; the length, width and height of each cuboid used as the residential building model can preferably form an arithmetic sequence with a tolerance of multiple S1 of SD.
[0040] Preferably, during the collection of the sample set, the residential models are arranged in different areas of the ground view, so that the residential models in the sample set are distributed in each of the nine grids formed by the top, middle, bottom, left, middle, and right sides of the image plane.
[0041] Preferably, the residential building size prediction model established in the residential building size prediction module has an input layer that receives the geometric features of the residential buildings in the rural residential building image from the general processing module, and an output layer that transmits the data to the iterative learning module and the general processing module through the first connection matrix and the second connection matrix, respectively. When training the residential building size prediction model offline, the iterative learning module adjusts the connection weights of the network based on the actual residential building size values input by the general processing module and the residential building size prediction model through the first connection matrix and the output value of the neural network in the model. In the field environment, the first connection matrix is disconnected, and the residential building size prediction model neural network predicts the size of the residential buildings in the current rural residential building image and outputs it to the general processing module through the second connection matrix. After processing and analysis by the general processing module, the data is output through the output section of the input-output module.
[0042] Preferably, the host unit is further configured to: obtain the coordinate and orientation information of the current location of the dwelling through the positioning unit, and after determining the size of the dwelling, output the size of the dwelling and the corresponding coordinate and orientation information through the input-output unit.
[0043] Preferably, after identifying the dimensions of the dwellings, the dimensions are also marked on the corresponding extracted edge segments of the dwellings in the rural dwelling image, such as using arrows to mark the corresponding directional line segments; preferably, the obtained coordinates and location information, as well as the dimensions of the dwellings, are also stored in a list in the storage module, and the list information can be stored in a cloud server.
[0044] When both the length and width in the dimension data exceed the preset safety distance threshold, the coordinates of the center point of the residential roof are provided; otherwise, an alarm is issued indicating insufficient safety distance to the residential roof. During the drone's flight, images are iteratively acquired, and dimension prediction processing is performed based on the acquired images. When a residential building with both length and width exceeding the preset safety distance threshold appears within the field of view, a prompt is issued indicating that a residential building with sufficient safety distance has been found. Preferably, the processed images are updated according to the flight speed and frame rate, so that residential buildings in the field of view are identified one by one. After a site with sufficient safety distance is found, images of residential buildings are acquired from multiple angles and dimension prediction is performed to verify the site's safety distance.
[0045] Preferably, multiple neural network-based residential building size prediction models are created in the host unit, each residential building size prediction model corresponding to a flight altitude and an overhead view angle of the image acquisition unit; when performing residential building size prediction online, the geometric features obtained after current acquisition and processing are input into the residential building size prediction model corresponding to the flight altitude and overhead view angle at the time of acquisition, and the model output value is obtained as the residential building size prediction value.
[0046] Preferably, in the BP neural network of the residential building size prediction model, the input layer receives six input quantities from the general processing module: the slopes k1 and k2 of the straight lines on both sides of the residential building, the perpendicular distances D1 and D2 from the center point of the image to the straight lines on both sides of the residential building, and the flight altitude and the overhead angle during acquisition. The output quantity of the output layer is the size of the residential building.
[0047] In another embodiment of the present invention, a rural dwelling identification system under the view of an unmanned aerial vehicle (UAV) is also provided, comprising: an image acquisition unit for image sensing and acquisition of rural scenes; a user interface unit for parameter input and operation and display of rural dwelling information; a server for data management and information transmission; a handheld terminal for remote operation and interaction with rural dwelling information; and a host unit connected to the user interface unit, the server, and the image acquisition unit; the host unit includes an input / output module, a processor, and a storage module, and is configured as follows:
[0048] A rural dwelling image is acquired, and a Region of Interest (ROI) containing the dwelling is determined within the rural dwelling image. The dwelling area is extracted from the ROI as the dwelling target image. Edge detection is performed on the dwelling target image to obtain dwelling edge images. A Hough transform is applied to the dwelling edge images, and a heuristic search is used to obtain dwelling edge images including three pairs of sidelines in the dwelling outline edge. The first pair of sidelines consists of two opposing sidelines at the top and bottom of the dwelling, sandwiched between a pair of side edges. The second and third pairs of sidelines consist of two opposing sidelines sandwiched between a pair of horizontal long sides and a pair of horizontal wide sides, respectively. From the dwelling edge image, the geometric features of the two opposing sidelines relative to the image center are sequentially obtained for any pair of sidelines. These geometric features are input into a trained neural network-based dwelling size prediction model to obtain a predicted dwelling size value, which is then output.
[0049] Preferably, the UAV-based rural dwelling identification system may not include the server and handheld terminal. Instead, during application, the UAV-based rural dwelling identification system is connected to the server and handheld terminal via a communication network. The UAV-based rural dwelling identification system may also include a positioning unit for coordinate acquisition.
[0050] Preferably, the UAV-based rural dwelling identification system also includes a flight control platform. In response to user operations on the control panel or instructions from the host unit, the flight control platform executes instructions and flies toward the designated residential location.
[0051] The UAV-based rural dwelling identification system continuously detects and predicts the size of multiple frames of rural dwelling images. When the safe distance between the dwelling sites meets the requirements, it compares the sizes of multiple dwelling sites and determines the site with the largest safe distance as the target landing point. The system outputs the coordinate position information of the target landing point to the flight control platform through the output unit, and the flight control platform manipulates the power unit to make the UAV fly towards the target landing point.
[0052] Preferably, when drones are used in group systems such as logistics and transportation, the system can have multiple handheld terminals and drone platforms, and subscription and publishing services can be provided accordingly. The host unit is also configured to: subscribe to the information acquisition topics of the handheld terminals, and in response to the topic message, transmit the detected residential dimensions and corresponding coordinates and location information to the handheld terminals through the server; or transmit the residential dimensions and corresponding coordinates and location information to the server in the form of a size event topic message when the detected residential dimensions reach a set safe distance threshold.
[0053] Preferably, the user interface unit is provided with an operation panel for parameter input, and a display screen for assisting parameter input and displaying the residential building size recognition results in the form of a list; the host unit also displays the recognized residential building size in the form of a curve on the display screen, and marks the coordinates and positions corresponding to the recognized size on the curve.
[0054] Preferably, the positioning unit in the system also acquires heading information, and records and identifies the orientation identified based on the heading, such as due north, and the orientation corresponding to the roof axis of the residential building in the output location information.
[0055] Compared with the prior art, the present invention has the following advantages: Addressing the unpredictable relative orientation relationship between the attitude and orientation of dwellings when UAVs collect images, the present invention analyzes the characteristic representation of dwelling dimensions in image space within rural dwelling images. It decomposes dwelling dimensions into three dimensions: length, width, and height. It proposes using two opposing sidelines that respectively sandwich a pair of long, wide, and high sides to describe the dimensions of dwellings in these three directions. Using four geometric features—the slope of the two opposing sidelines in any direction, and the perpendicular distance from the image center point to the two sidelines—as constraints, an artificial neural network is established to predict dwelling dimensions. The collected samples are used to identify and train the nonlinear mapping relationship between the size of rural dwellings and the proposed four geometric features. Each sample can identify rural dwellings by classifying and recognizing targets in aerial images. Through the extraction and edge detection of dwellings, three pairs of edges corresponding to the length, width and height of the dwelling contour are obtained by Hough transform. The four geometric features are extracted from two edges of any pair of edges. The results are input into the trained dwelling size prediction model to predict the length, width and height of rural dwellings online. The neural network-based residential building size prediction model established in this invention has good generalization performance. It can accurately predict the three-dimensional dimensions of various types of residential buildings based on images of rural houses collected from uncertain flight paths. It also has a wide range of applications. Since it extracts features from straight lines without relying on precise pixel start and end points, the prediction accuracy is high. At the same time, due to the robustness of the established model, the step size in the cumulative matrix of the Hough line transform does not need to be set too small when extracting the edge lines of residential buildings. Furthermore, the identification of the correspondence between edge points and straight line elements in the matrix is simplified, reducing the amount of computation and speeding up the search, thereby further improving the real-time performance of rural residential building identification. Based on heuristic search, it can automatically find and identify the horizontal edges and vertical side edges of the ground and roof of residential buildings, simplifying the operation confirmation steps.
[0056] It should be understood that all combinations of the foregoing concepts and the additional concepts discussed in more detail below (provided that such concepts are not inconsistent with each other) can be contemplated as part of the inventive subject matter disclosed herein. In particular, all combinations of the claimed subject matter appearing in this disclosure can be contemplated as part of the inventive subject matter disclosed herein. Attached Figure Description
[0057] Figure 1 This is a schematic diagram of a rural residential scene.
[0058] Figure 2 A schematic diagram of a device and system for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV).
[0059] Figure 3A This is a structural diagram of the main unit / controller in a device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV). Figure 3B To obtain the structure diagram of the preprocessing module, Figure 3C Here is a structural diagram of the feature extraction module;
[0060] Figure 4A This is a schematic diagram of the neural network model structure for the residential building size prediction module. Figure 4B A schematic diagram of a partial structure of the main unit / controller in a device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV).
[0061] Figure 5A This is a diagram illustrating the posture of the drone during sample collection. Figure 5B A comparative illustration of residential building models of different sizes;
[0062] Figure 6A , Figure 6B These are schematic diagrams of residential building models taken from different angles and orientations during sample collection. Figure 6C This is a schematic diagram of the single-sided outline of a residential building. Figure 6D The image shows a residential building model after Hough line transform and line detection.
[0063] Figure 7A , Figure 7E This is a schematic diagram showing the boundary constraints and geometric features of the proposed residential building model. Figure 7B , Figure 7C , Figure 7D These are schematic diagrams showing the dimensional constraints and geometric features of the length, width, and height of rural dwellings.
[0064] Figure 8 To test the fitting curve between the actual and predicted width values of the residential building model;
[0065] Figure 9 To test the fitting curve between the actual and predicted values of the residential building model length;
[0066] Figure 10To test the fitting curve between the actual and predicted values of the residential building height model;
[0067] Figure 11 A schematic diagram illustrating target annotation in a rural drone field-of-view image.
[0068] Figure 12 A schematic diagram of small, medium and large-sized targets in a rural setting;
[0069] Figure 13A This is a partial diagram of the YOLOv5s network configuration file. Figure 13B This is a schematic diagram illustrating the feature fusion and learning of features at different scales in ASFF-3. Figure 13C , Figure 13D These are partial diagrams of the ASFF-YOLOv5s network configuration file and a schematic diagram of the network structure.
[0070] Figure 14A , Figure 14B A diagram illustrating the comparison of loss curves on the training and validation sets before and after network improvement;
[0071] Figure 15 This is a schematic diagram showing the anchor frame and size markings of residential buildings in a drone's field-of-view image.
[0072] The system includes: a 1000-unit UAV-based rural dwelling identification system, a 900-unit UAV-based rural dwelling identification device, a 100-unit main unit / controller, a 200-unit user interface unit, a 300-unit handheld terminal, a 400-unit server, a 500-unit image acquisition unit, a 600-unit flight control platform, a 700-unit positioning unit, and an 800-unit rural dwelling environment / scene display.
[0073] 110 processors, 120 memory modules, 130 input / output modules.
[0074] 111 General Processing Module, 112 Acquisition and Preprocessing Module, 113 Feature Extraction Module, 114 Iterative Learning Module, 115 Residential Building Size Prediction Module.
[0075] 1121 Image Acquisition Unit, 1122 Target Detection Unit, 1123 Residential Area Extraction Unit,
[0076] 1131 Edge Detection Unit, 1132 Line Extraction (Filtering) Unit, 1133 Slope Feature Extraction Unit, 1134 Distance Feature Extraction Unit, 1135 Side Edge Extraction Unit, 1136 Top and Bottom Edge Extraction Unit, 1137 Roof Parallel Line Extraction Unit.
[0077] 1141 First connecting array, 1142 Second connecting array,
[0078] 1101 Neural Network Model, 131 Input Section, 132 Output Section.
[0079] 210 Display screen, 220 Operation panel. Detailed Implementation
[0080] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings, but the present invention is not limited to these embodiments. The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention.
[0081] Example 1:
[0082] When drones perform missions in rural areas, they need to measure the height of houses to set a minimum flight altitude when flying over villages, and measure the length and width of houses when landing on rooftops or circling around them. Therefore, identifying and measuring the dimensions of rural houses from the drone's field of vision is fundamental to the application of drones in rural areas.
[0083] like Figure 1 As shown, rural scenes typically consist of farmland and vegetable gardens, with houses scattered among them. Drones provide transportation services in rural areas, offering landing sites on rooftops or flat ground around houses. To ensure the safety of drone flights and landings, it is necessary to understand the dimensions of rural houses to make flight and landing decisions.
[0084] Because rural dwellings are often not of uniform specifications and their dimensions vary widely, and because the property rights belong to different entities, it is difficult to collect dimensional data a priori. Therefore, this invention provides a UAV-based rural dwelling identification device that does not rely on prior three-dimensional dimensional measurements and can estimate the size of rural dwellings by acquiring images of them under non-fixed flight directions during flight.
[0085] like Figure 2As shown, the present invention provides a rural dwelling identification device 900 under the view of an unmanned aerial vehicle (UAV), which includes a host unit 100 and a user interface unit 200; wherein, the host unit 100 includes a processor 110, a storage module 120 and an input / output module 130; the host unit 110 is configured to: acquire rural dwelling images, determine ROI regions containing dwellings in the rural dwelling images, extract dwelling regions from the ROI regions as dwelling target images; obtain dwelling edge images by edge detection on the dwelling target images; perform Hough line transform on the dwelling edge images, and obtain three pairs of sides including the outline edges of dwellings through heuristic search. The image shows the outline of a dwelling, wherein the first pair of outlines consists of two opposing outlines at the top and bottom of the dwelling, sandwiched between a pair of side edges; the second and third pairs of outlines consist of two opposing outlines sandwiched between a pair of horizontal long sides and a pair of horizontal wide sides, respectively. From the image, the geometric features of the two opposing outlines relative to the image center are obtained for any pair of outlines. These geometric features are then input into a trained neural network-based dwelling size prediction model to obtain a predicted dwelling size value. This predicted size value is then output from the output unit 132 of the input / output module 130.
[0086] See Figure 1 As shown, in images captured in rural scenes, houses are approximately displayed as cuboids. Generally, when a drone captures images, the horizontal plane of the imaging surface is parallel to the horizontal plane, and the vertical axis in the center of the imaging surface is approximately parallel to the vertical line. However, due to perspective transformation, any two parallel edges or side edges of the actual houses in the image appear to converge at an almost invisible point at infinity. On the one hand, it is obvious that when calculating the distance from the two ends of the edge segments of the houses in the image and then converting it back to the actual width according to a proportional relationship, the proportional relationship is difficult to determine precisely, as it will be directly related to the front-back or left-right position of the edge in the image. On the other hand, referring to... Figure 1 In practical applications, the scenarios shown in the left and right images cannot be forced to be taken directly towards a specific surface of a dwelling for the purpose of scaling calculations. Otherwise, the applicability of the identification will be greatly limited.
[0087] Although the physical dimensions of rural dwellings are definite within the field of view, their edges, both in pixel width and direction, are uncertain in image space when sampled from different flight paths. This definite physical dimension and uncertain image features present a contradictory relationship that is difficult to reconcile. Therefore, it is necessary to decouple and model this contradiction. The flight path of a drone is not fixed, and the orientation of the dwellings varies greatly; therefore, the relative relationship between the drone's flight path and the dwelling's orientation is unpredictable and cannot be predetermined. To improve the universality and applicability of dwelling size identification, it is necessary to study how to represent the relationship between the two under various relative orientations in image space—that is, what features in the image can determine the size of the dwellings?
[0088] Observing the length, width, and height edges (boundaries) of dwellings in various orientations within images, it was found that the key lies in the two straight lines on either side of the length, width, or height edges. However, the dwelling edges themselves do not uniquely determine the size of the dwelling; the angle of the straight line only reflects the relative angular difference between it and the heading. Extensive experiments and tests also revealed that the values of the length, width, or height—that is, the relative distance between the two straight lines on either side of the dwelling edges—affect the imaging effect of the two parallel sides in the image. So, how can the characteristics of these two sides be characterized? Through further research, hypotheses, and reasoning, it was proposed that four geometric feature inputs be used: the slopes k1 and k2 of the two parallel lines at different spacings, and the perpendicular distances D1 and D2 from the image center point to the two lines on either side of the line. These four inputs serve as constraints on the size of one orientation edge of the dwelling within the two lines in the image space. Furthermore, how can this constraint be mapped to the size of the dwelling? Since the relationship between the two is non-linear and coupled, inspired by the fitting application of neural networks in non-linear system modeling, this invention proposes using artificial neural networks to establish a dwelling size prediction model.
[0089] A neural network is a machine learning model that mimics the biological nervous system, and is particularly well-suited for modeling nonlinear mappings. It learns and represents complex functional relationships through a multi-layered structure of artificial neurons. A neural network consists of an input layer, hidden layers, and an output layer; each layer contains several neurons, and each neuron is connected to neurons in the previous layer via weights. The input layer receives the input data, the hidden layers perform nonlinear transformations and feature extraction (and can have multiple layers), and the output layer outputs the final prediction result.
[0090] like Figure 4AAs shown, the residential building size prediction model of the present invention adopts a three-layer BP neural network with one hidden layer. Its input layer receives four geometric feature inputs from the general processing module: the slopes k1 and k2 of a pair of sides corresponding to the length, width, or height of the residential building, and the perpendicular distances D1 and D2 from the image center point to the pair of sides. The output of the output layer is the length, width, or height of the residential building corresponding to the pair of sides.
[0091] To extract the four input quantities or geometric features of rural dwellings in image space, it is necessary to first obtain the two straight lines corresponding to the length, width, or height of the dwelling from the image. For this, see section 2. Figure 3A As shown, the processor 110 of the host unit or controller 100 is equipped with multiple processing modules, namely a general processing module 111, an acquisition and preprocessing module 112, a feature extraction module 113, an iterative learning module 114, and a residential building size prediction module 115. The general processing module 111 responds to messages and operations from the user interface unit, initiates the identification of residential buildings in the acquired images, schedules the identification processing flow, inputs the real-time acquired rural residential building images into the acquisition and preprocessing module 112, and outputs the size data identified and predicted by the residential building size prediction model in the residential building size prediction module 115.
[0092] The acquisition and preprocessing module 112 acquires images of rural dwellings captured by the image acquisition unit 500, removes high-frequency noise through filtering, and then determines the Region of Interest (ROI) including the dwellings based on deep learning features or color and texture features. The portion outside the ROI is removed from the original rural dwelling image to form the dwelling target image. Edge detection is then performed to form a binarized dwelling edge image. After performing a Hough transform on the binarized dwelling edge image, the longest vertical and horizontal line segments are selected as candidate line segments based on the polar coordinate space parameter features in the Hough transform. Based on the directional features of candidate straight line segments, two opposing lines sandwiching a pair of long, wide, and high sides are determined to serve as a pair of lines representing the length, width, and height of the dwelling, respectively, forming three pairs of lines representing the dwelling. For each line in any pair of these three pairs, the feature extraction module combines the corresponding binarized edge image of the dwelling and describes it using the coordinates of its two endpoints. Based on the two endpoints, the slope of the line and the perpendicular distance from the image center to the line are calculated. The slope of the lines on both sides of the dwelling and the perpendicular distance from the image center to the line are used as the geometric features of the dwelling in the rural dwelling image. Preferably, the slope of the line can be obtained from the angle corresponding to the line in the Hough line transform accumulation, i.e., the accumulator array.
[0093] like Figure 3BAs shown, the acquisition and preprocessing module 112 is further divided into an image acquisition unit 1121, a target detection unit 1122, and a residential area extraction unit 1123. The image acquisition unit 1121 acquires images of rural dwellings captured in real-time by the external image acquisition unit 500 via the general processing module 111, or acquires pre-stored residential dwelling images from the storage module 120 in an offline state. The residential area extraction unit 1123 initially determines the ROI regions (Regions of Interest) including dwellings in the images. Preferably, a target classification model based on a YOLOv5s network can be established in the target detection unit 1122. After acquiring rural scene images using the overhead camera in the image acquisition unit, the network is trained offline using an image sample set composed of rural scene images labeled with various target regions such as rural dwellings, to obtain a rural scene target classification model. For online operation, see [link to relevant documentation]. Figure 15 As shown, the residential area extraction unit 1123 uses the rural scene target classification model to identify and process the rural scene image containing residential buildings, i.e., the rural residential image, to obtain the anchor frame of the rural residential area, and takes the area within the anchor frame as the ROI area of the rural residential area.
[0094] Among the YOLOv5 model series, YOLOv5s has the smallest model size and the fewest training parameters. Comparative experiments show that YOLOv5s can balance recognition accuracy and image processing frame rate, making it suitable for deployment on mobile devices. The classic YOLOv5s network structure is divided into three main parts: the first part is the Backbone, which is used to extract features from the input image; the second part is the Neck, which is used to extract features at different scales from the network structure at different depths in the Backbone and fuse these features to maximize the preservation of feature information at different depths, which is more conducive to the learning of network model parameters during training; the third part is the Output, which is the output prediction. This part mainly uses tensors of different scales after feature fusion to predict the input image.
[0095] In rural transportation scenarios, targets with rural characteristics often appear, such as high-voltage power line towers, puddles, streams, small bridges, signal towers, houses, and greenhouses. Under normal weather conditions, a large number of target images in rural scenes were acquired through drone aerial photography. To enrich the dataset, images were taken in both sunny and cloudy weather to account for varying lighting conditions. After the dataset was collected, the targets to be detected in the images were labeled so that the images could be used for supervised learning when trained on the neural network. Figure 11The example illustrates labeling a residential building as a target object. After labeling all targets, click the "Save" button to complete the labeling and save of an image. The labeled file is saved in .txt format, with each label file corresponding to one image file. Each line in the label file represents a target location in the image, and the saved information, from left to right, includes the target type, the x-coordinate of the target's center point, the y-coordinate of the target's center point, the target's width, and the target's height; the data uses normalized numerical values.
[0096] The dataset was randomly divided into training and validation sets according to a certain ratio. The network was trained on the training set and then validated using the validation set. The results showed that the AP (Average Performance) of residential buildings on the validation set was not high, only 0.669. Therefore, after analyzing the characteristics of the samples and the features of the deep network, it was decided to improve the performance of the classification network by focusing on the characteristics of the samples.
[0097] In rural scenes, the size of the targets varies significantly. For example... Figure 12 As shown, high-voltage power line towers, residential buildings, and a small river correspond to small, medium, and large targets in the image, respectively. Based on this feature, this invention proposes introducing an ASFF adaptive spatial feature fusion module into a deep learning network to enhance its ability to predict and recognize features at different scales.
[0098] like Figure 13B As shown, the ASFF module learns the weights of feature maps at different scales as follows: 1) The feature maps of Level 1 and Level 2 are upsampled to the same scale as Level 3. The transformation process is denoted as X. 1→3 X 2→3 and X 3→3 2) Compress the number of channels at the three scales to a uniform length using a 1*1 convolution kernel; 3) Concatenate the compressed channels; 4) Further compress the number of channels to 3 using a 1*1 convolution kernel; 5) Obtain the weight information, α, for the three different scale spaces using the Softmax classification function. 3 β 3 and γ 3 For the i-th layer, the corresponding ASFF fusion feature formula is:
[0099] ASFF i =X 1→i ·α i +X 2→i ·β i +X 3→2 ·γ i
[0100] A partial diagram of the network configuration file after adding ASFF is shown below. Figure 13C As shown. Comparison Figure 13A , Figure 13CAs shown, the ASFF-YOLOv5s network established by the rural scene target classification model of this invention adds three layers 24, 25, and 26 between the 23rd layer and the next layer of the YOLOv5s network. These three layers receive the outputs of the 17th, 20th, and 23rd layers as inputs, and after feature fusion at different scales by the ASFF adaptive spatial feature fusion module, they successively replace the original 17th, 20th, and 23rd layers as inputs to the original output layer 24, which is the 27th layer of the new network. Figure 13D This is a schematic diagram of the ASFF-YOLOv5s network structure.
[0101] After the ASFF-YOLOv5s model is trained, its performance is compared with that of the classic YOLOv5s model. Figure 14A , Figure 14B The diagrams illustrate a comparison of the training set loss curves and the validation set loss curves for both models. It can be seen that the ASFF-YOLOv5s loss value is significantly lower than the YOLOv5s baseline model, and its convergence speed is also significantly faster. At 100 epochs of training, the training_losses for ASFF-YOLOv5s and YOLOv5s are 0.1437 and 0.1140, respectively, and the validation_losses are 0.0951 and 0.0915, respectively.
[0102] More important than the decrease in network model loss is the detection accuracy of ASFF-YOLOv5s on the validation set for each category. A comparison of AP values for each category before and after the network improvement shows that the ASFF-YOLOv5s model improves the AP values for all categories on the validation set compared to the classic YOLOv5s benchmark model. Specifically, it improves by 3.9% for high-voltage power line towers, 9.9% for puddles, 5% for small rivers, 7.8% for small bridges, 8% for residential buildings, 8% for signal towers, and 3.8% for greenhouses, with an overall mAP improvement of 6.5%. This demonstrates that the improved network model significantly improves the detection accuracy for targets of various sizes.
[0103] After the target classification model based on the ASFF-YOLOv5s network locates the ROI region of the residential buildings using anchor boxes, the residential building area is extracted from the ROI region as the foreground object, and the non-foreground object area in the rural residential building image is used as the background object. The background object in the rural residential building image is set to zero to form the residential building target image, which facilitates the subsequent extraction of the straight lines of the residential building edges.
[0104] like Figure 3CAs shown, the feature extraction module 113 also includes an edge detection unit 1131, a straight line extraction and filtering unit 1132, a side edge extraction unit 1135, a top and bottom edge extraction unit 1136, a roof parallel line extraction unit 1137, a slope feature extraction unit 1133, and a distance feature extraction unit 1134.
[0105] For residential building target images, the edge detection unit 1131 performs edge extraction. For binarized residential building target images, edge extraction can be performed directly. For color or grayscale residential building target images, the color image is first converted into a single-channel image, and then the Canny edge detection operator is used to detect the edges.
[0106] The line extraction and filtering unit 1132 extracts the edge lines of the dwellings, including the long side, wide side, and height side, from the binarized image of the dwelling edges. First, how to represent the long side, wide side, and height side of the dwelling? If we directly search for the corresponding line segments, the movement of the endpoints of the line segments will introduce significant errors; more importantly, even if a boundary line segment is obtained, how do we know its actual length? Through modeling and analysis, this invention describes the boundary edge as two opposing edge lines on either side of any boundary edge (length, width, or height), perpendicular to the boundary edge and together sandwiching it. Second, how to extract these opposing edge lines? This invention uses the Hough line transform.
[0107] The Hough line transform maps points in image space to curves in a parameter space, such as polar coordinates. It then finds the intersection points of these curves in the parameter space to detect straight lines. In image space, the equation of a straight line can be expressed as: y = kx + b; using polar coordinates, it can be expressed as: ρ = xcosθ + ysinθ, where ρ is the perpendicular distance from the line to the origin, and θ is the angle between the perpendicular and the positive x-axis. The essence of the Hough line transform is to establish a mapping between image space and parameter space. By fully utilizing the global characteristics of the image and leveraging the point-line duality, it maps pixel coordinates to the straight lines represented by elements in the parameter space. Analysis is then performed in the parameter space, and the detection task is completed through cumulative statistical analysis.
[0108] To detect straight lines, a planar matrix, or accumulator array (ρ, θ), is created in the ρ-θ plane based on the Hough line transform, dividing the parameter space into a discrete grid. Then, for each edge point (x, y) in the image space, the ρ value of ρ = xcosθ + ysinθ under different θ values is calculated, and a vote is taken at the corresponding position in the accumulator array. After voting, peak detection is performed, and the element with the larger value in the accumulator array is the detected straight line parameter (ρ, θ). Simultaneously, a heuristic search is performed to obtain the residential boundary image, including three pairs of edges in the outline of the residential buildings.
[0109] Specifically, refer to Figure 1 As shown, the drone is in level flight when acquiring images; therefore, generally, the edges of a plumb wall on the ground appear approximately parallel to the vertical axis in the image. Based on this heuristic rule, the side edge extraction unit 1135 searches for the two or three longest edges within a range where the straight line is nearly perpendicular, using the accumulator array (ρ, θ), as the side edges of the dwelling; where the perpendicular line is nearly perpendicular, its θ is within the range of 0 degrees. Generally, three or more side edges can be found for a dwelling; only when the drone is directly facing the side wall of the dwelling will only two side edges be found.
[0110] After excluding straight lines within the lateral edge direction, the two longest straight lines with similar directions are selected from the voted line set to form candidate line pairs. Several subsets of lines with similar directions can be clustered based on the similarity of θ values. For each subset, 3 to 5 candidate lines can be selected. Within the non-lateral edge direction, the line segments are sorted from longest to shortest according to the accumulator array element values (i.e., length). Starting with the longest main line segment, two secondary line segments with similar directions are sequentially searched, and these main line segments and their corresponding two secondary line segments form candidate line groups. Preferably, 2 to 3 candidate line groups are obtained; some residential buildings have multi-level stepped shapes, resulting in multiple pairs of parallel edges.
[0111] Combination Figure 1 As shown, it can be observed that the horizontal boundaries of the houses on the left and right sides, including the rooftops and the ground, have nearly parallel straight lines in the field of view. Therefore, all the side edge groups are recorded in array form, and the upper and lower endpoints of each side edge are obtained by comparing the pixel set corresponding to the same (ρ, θ) recorded during the Hough line transform accumulation process, using the minimum and maximum values of the ordinate. For any line segment in the obtained candidate line group, the distance to the upper and lower endpoints of the house side edge is calculated. If the element with the smallest value in the distance set corresponds to the distance between the line segment and the upper endpoint of a side edge, then the line segment is marked as a candidate line for the roof; otherwise, if the element with the smallest value in the distance set corresponds to the distance between the line segment and the lower endpoint of a side edge, then it is marked as a candidate line for the ground.
[0112] Based on the direction corresponding to different candidate straight line groups and the identification of candidate straight lines within each group, the top and bottom edge extraction unit 1136 and the roof parallel line extraction unit 1137 respectively determine the three pairs of edge lines corresponding to the length, width, and height of the residential building outline edge. Each candidate straight line group is sorted according to the length of its main straight line segment, with the two candidate straight line groups with the largest corresponding lengths designated as the first candidate straight line group and the second candidate straight line group, respectively. First, in one of the candidate straight line groups, the top and bottom edge extraction unit 1136 forms the first pair of edge lines using two opposing edge lines: a candidate straight line of the residential building roof sandwiched between a pair of side edges, and a candidate straight line of the ground. Then, in the second candidate straight line group, the roof parallel line extraction unit 1137 forms the second pair of edge lines using two opposing edge lines of the residential building roof sandwiched between a pair of candidate straight lines from the first candidate straight line group; similarly, in the first candidate straight line group, the third pair of edge lines is formed using two opposing edge lines of the residential building roof sandwiched between a pair of candidate straight lines from the second candidate straight line group.
[0113] The first pair of edge lines sandwiches a pair of side edges or high edges of the dwelling. Between the ordinates of the corresponding endpoints of the first pair of edge lines, namely the candidate roof line and the candidate ground line, the ordinates of the upper and lower endpoints of a side edge can be accommodated. That is, the interval covered by the ordinates of the corresponding endpoints of the candidate roof line and the candidate ground line is basically close to or includes the ordinates of the upper and lower endpoints of the side edge.
[0114] In the Hough line transform, the distance resolution (ρ) is typically 1 pixel, while the angle resolution (θ) is typically π / 180. Since the boundary straight lines in rural dwellings exhibit relatively distinct characteristics, the Hough line transform process can be optimized. Preferably, when traversing and accumulating the element values in the ρ-θ plane matrix, the spacing between elements within the value range can be set to a larger value based on prior testing. This increases the spacing between distance and angle values, reduces the number of rows and columns in the matrix, and significantly reduces the computational load and time during the calculation and judgment process.
[0115] Preferably, before peak detection, the elements in the accumulator array (ρ, θ) are merged to be similar. When the values of several neighboring elements in the array are all above a preset pixel threshold, i.e., the minimum length of the straight line, the centroid of the polygon enclosed by the coordinates of the neighboring elements is calculated, and the (ρ, θ) corresponding to the calculated centroid is used as the edge of the residential building.
[0116] For any pair of edges found, the slope feature extraction unit 1133 and the distance feature extraction unit 1134 calculate the slope of the line and the perpendicular distance from the image center point to the line, respectively. (See also...) Figure 7AThe diagram shows the boundary constraints and geometric features of a residential building model, illustrating a pair of sides. The two sides on the left have angles and slopes of the lines they contain, α(k1) and β(k2), respectively, and their perpendicular distances from the image center to these two sides are D1 and D2, respectively. Similarly, the two sides on the right have angles and slopes of the lines they contain, α′(k1′) and β′(k2′), and their perpendicular distances from the image center to these two sides are D1′ and D2′, respectively. In the oblique photographic image plane, the two parallel lines will intersect at a point O at their far ends; point O is called the hidden line disappearance point.
[0117] The algebraic equation of a straight line can commonly be expressed in three forms: the general form, the two-point form, and the slope-intercept form. The general form is: Ax + By + C = 0, where A and B are not simultaneously zero, and the slope of the line is k = -A / B. In a Cartesian coordinate system, the y-intercept of the line is b = -C / B. The general form applies to all linear equations.
[0118] Two-point form is: The two-point form equation is the equation of a straight line passing through two points given their coordinates. The slope of this line is k = (y2 - y1) / (x2 - x1). As can be seen from the two-point form equation, this equation is applicable to describing straight lines that are not perpendicular to the x-axis or y-axis.
[0119] The slope-intercept form is: y = kx + b. The slope-intercept form explicitly expresses the slope k and intercept b of the line and is suitable for representing lines that are not perpendicular to the x-axis.
[0120] In the image, the straight line detected by the Hough transform is described by the coordinates of its two endpoints, denoted as (x1, y1) and (x2, y2). Substituting these coordinates into the two-point equation and transforming it into slope-intercept form, we obtain: Its corresponding slope k = y1-y2 / x1-x2. The only difference is that the slope of the line is calculated in the image coordinate system (the y-axis is downward as the positive direction). Therefore, the slope of the line in the image is equal in magnitude but opposite in sign to that in the Cartesian coordinate system.
[0121] Using the two-point equation, the perpendicular distance from the image center to the edge of the house can be calculated as follows: In the formula, (x0, y0) represents the coordinates of the image center, A and B are the coefficients of the general form of the equation of a straight line, and C is a constant term with the following conditions: A = y2 - y1, B = x1 - x2, C = -x1(y2 - y1) + y1( x2 -x1). Based on the calculation formulas for k and d, the slopes k1 and k2 of the straight lines on both sides of the dwelling, as well as the perpendicular distances D1 and D2 from the center point of the image to the two straight lines, are calculated respectively, which are used as the geometric features of the dwelling in the rural dwelling image.
[0122] Combination Figure 1 Actual residential scenes and Figure 7A The concept, Figure 7B , 7C 7D illustrates the constraints and geometric features of the length, width, and height dimensions of rural dwellings; among them, Figure 7D The two middle edges are the first pair of edges sandwiched between a pair of lateral edges or high edges. Figure 7B , Figure 7C The two side lines in the middle are the second and third pairs of side lines, which are sandwiched between a pair of long sides and a pair of wide sides. See also Figure 7E As shown, for the determined residential building boundaries, a perpendicular line is drawn from the center point of the image to the boundary line, represented by a dashed line. Among these, in... Figure 7E Since the foot of the perpendicular is outside the range of the roof edge line, the roof edge line is extended to the foot of the perpendicular.
[0123] Preferably, a human-computer interaction interface is provided in the user interface unit. For residential buildings with complex outlines, after automatic edge line filtering, the operator confirms and marks three pairs of edges through the human-computer interaction interface. The geometric features of the marked residential building edges are then calculated to predict the building's dimensions. Figure 1 Taking the residential building shown as an example, if there are structures on the roof of the building on the left, the length and width of these structures can be further marked on the human-computer interface. The geometric features of the marked structures can then be calculated to predict their dimensions. Finally, the length and width of the roof plane are determined by subtracting the dimensions of the structures from the overall length and width; the resulting values are then used as the basis for determining the safe landing distance for the drone. And for... Figure 1 For the houses on the right side of the middle section, where the roof edges are almost unobstructed, the roof dimensions can be obtained directly from the house size prediction model after automatic screening and simple confirmation by the operator. This greatly simplifies the calculation process and realizes the automatic calculation of rural house dimensions.
[0124] To reveal the nonlinear mapping structure of the residential building size prediction model of this invention, training image samples are obtained by tilting the model after establishing a rural residential building proportion model. (Reference) Figure 5A As shown, when a drone takes aerial photos of the ground, the angle between the optical axis of the airborne lens and the horizontal plane is α, which can be a commonly used angle such as 45 degrees. Since the actual dimensions of rural dwellings vary greatly, this invention uses cuboids to create various sizes of dwelling models for ease of modeling, such as... Figure 5B As shown.
[0125] In offline mode, a drone is used to collect images of a model of a traditional dwelling and form a sample set. This sample set is then used to train the dwelling size prediction model. The dwelling model is rectangular, and the size of the cuboid is a first preset multiple S1 of the actual dwelling size. The multiple S1 is set according to the drone's rural cruise flight altitude and camera resolution. The size range of different cuboids used in the dwelling model corresponds to S1 times the common size range of rural dwellings, and the size difference between different cuboids corresponds to S1 times the actual dwelling size difference SD. For example, when the flight altitude is 100 meters, the sampling altitude of the model is 2 meters, thus ensuring that the dwelling model occupies a significant portion of the image space during imaging. This coordinates the image resolution with the recognition accuracy of the dwelling size prediction model and eliminates the need for the drone to specifically lower its altitude to predict dwelling sizes, thereby improving the applicability of the prediction.
[0126] Preferably, the first preset multiple S1 is between 1 / 50 and 1 / 20, and the actual residential building size difference SD is between 0.5 and 3 meters; the length, width and height of each cuboid used as the residential building model can preferably form an arithmetic sequence with a tolerance of multiple S1 of SD.
[0127] Using a model of a residential dwelling, three different actual drone cruising altitudes were simulated under laboratory conditions, with the camera shooting at a 45° downward angle. Actual aerial photography altitude h real Aerial photography altitude h under laboratory conditions model The correspondence is: when h real When h is 30m, 40m, and 50m, model The sizes are 667mm, 889mm, and 1111mm, with the following scaling ratios:
[0128]
[0129] refer to Figure 6A As shown, due to projective transformation, lines that are parallel in the original 3D cube become non-parallel in the 2D image. The residential building model is a three-dimensional model; the long, wide, and high sides of the building belong to different spatial locations and are combined in pairs to form different planes. This invention uses two opposite edges perpendicular to and sandwiching a target edge on a certain plane to express geometric feature constraints. Therefore, feature extraction is performed from the edges on different planes.
[0130] Because the images of the residential building models were sampled from oblique angles, to ensure the diversity of features representing the residential building models in the images and to improve the generalization ability of the network model predicting residential building sizes, the residential building models were photographed from 0° to 180°. Simultaneously, to highlight the distance relationship between the drone and the actual location of the residential buildings, the residential building models were arranged in different areas of the ground view, so that the residential building models in the sample set were distributed within each grid cell of the nine-square grid formed by the top, middle, bottom, left, middle, and right sides of the image plane. For simplicity, [the text continues with further details]. Figure 6B It indicated Figure 6A The top surface of the rectangular prism is located at different positions in image space during sample collection. To improve model accuracy, a black ground background was used for photography; for ease of display, the images have been inverted.
[0131] This invention subdivides the dimensions of residential building models into an interval sequence and takes pictures from various relative angles during the sampling process, ensuring sufficient samples, thereby improving the generalization ability of the neural network and reducing the prediction error of residential building dimensions.
[0132] For images of rural dwellings based on a residential building model, the acquired images are filtered and converted to grayscale, then binarized to form the target images of the dwellings. Edge detection is then used to obtain the edge images of the dwellings. Taking an image of the top plane of a residential building model at a certain location as an example, the preprocessing results at each stage are as follows: Figure 6C As shown, from left to right, the results correspond to grayscale conversion, binarization, and edge detection, respectively. A Hough line transform is applied to the Canny edge-detected residential edge image, and the transformed image is shown below. Figure 6D As shown, the left and right figures represent the Hough line transformation results of a pair of side lines on both sides of the wider side. Then, a heuristic search is performed on the transformed lines to select the corresponding two side lines.
[0133] To ensure the accurate detection and marking of residential building boundaries, a manual verification step can be added to the boundary search process. Specifically, during the process of collecting straight line features representing residential buildings and establishing a training sample library, a human-computer interface is used to eliminate interference from non-residential building boundary lines. A graphical user interface is established using the HighGUI module provided by OpenCV, and the mouse callback function setMouseCallback() is used. In the callback function, left mouse button click events and right mouse button click events are created. Each left mouse button click draws a straight line detected by Hough line transform in the image, which is then confirmed by comparison with the original residential building image.
[0134] Referring to the above Figure 7AThe description states that for each model residential building image, four geometric features are calculated for the boundary lines on both sides of a boundary, and the corresponding actual boundary dimensions (length, width, or height) are recorded as the true dimensions of the residential building. These geometric features and true dimensions are then used to form a sample set, which is then input into a neural network-based residential building size prediction model for offline training. Preferably, the sample set can be stored in a data table, with each true input and output value as a record.
[0135] The BP neural network used in the residential building size prediction model in this embodiment is as follows: Figure 4A As shown, its model is a three-layer structure including one hidden layer, wherein,
[0136] The output of the j-th node in the hidden layer is
[0137] The output of the first node in the output layer is
[0138] In the formula, f() is taken as the tansig function, w ij and v j These are the connection weights from the input layer to the hidden layer and the connection weights from the hidden layer to the output layer, respectively. j b and are the thresholds for the hidden layer and output layer, respectively. m = 4 is the number of input variables, and k is the number of hidden layer nodes. Gradient descent is used for network training.
[0139] When predicting the size of residential buildings online, a pair of sidelines corresponding to the length, width, and height of the residential building are obtained from the residential building model image. Based on the four feature quantities of each pair of sidelines, the length, width, and height of the residential building model are predicted. The prediction model also maps the actual residential building size back to the actual residential building size according to the ratio between the model and the actual size as the actual output value.
[0140] Preferably, multiple neural network-based residential building size prediction models are created in the host unit, each residential building size prediction model corresponding to a flight altitude and an overhead view angle of the image acquisition unit; when performing residential building size prediction online, the geometric features obtained after current acquisition and processing are input into the residential building size prediction model corresponding to the flight altitude and overhead view angle at the time of acquisition, and the model output value is obtained as the residential building size prediction value.
[0141] The sample datasets for three different heights were randomly divided into training and test sets, both of which included various residential building models. The geometric features of the test set were used as input to the trained network model to obtain predicted values. The fitting curves of the actual and predicted values of the residential building width at different heights are shown below. Figure 8 As shown; correspondingly, the fitting curves of the actual and predicted values of the residential building length model are as follows: Figure 9As shown; the fitting curves of the actual and predicted values of the residential building height model are as follows. Figure 10 As shown.
[0142] The regression results were analyzed at three different laboratory heights, and the coefficient of determination R0 in the width dimension was calculated for each of the three different heights. 2 The coefficients of determination were 0.99815, 0.99863, and 0.99649, respectively. The closer the coefficient of determination is to 1, the better the regression performance, indicating that the model size prediction model has high reliability for all three heights. The root mean square errors (RMSE) for the three different heights were 2.6277, 1.632, and 1.8552, respectively, indicating that the predicted values are close to the actual values, and the network model has good generalization ability. The R-squared values for length and height dimensions... 2 Both RMSE and RMSE indicate that the model has high reliability and good generalization ability.
[0143] Example 2
[0144] Based on the previous embodiment, such as Figure 2 As shown, this embodiment provides a rural dwelling identification system 1000 under the view of an unmanned aerial vehicle (UAV), which includes: an image acquisition unit 500 for image sensing and acquisition of rural scenes, a user interface unit 200 for parameter input and operation and display of rural dwelling information, a server 400 for data management and information transmission, a handheld terminal 300 for remote operation and interaction with rural dwelling information, and a host unit 100 connected to the user interface unit 200, the server 400 and the image acquisition unit 500. The host unit 100 includes an input / output module 130, a processor 110, and a storage module 120, and is configured to: acquire images of rural dwellings; determine ROI regions containing dwellings in the rural dwelling images; extract dwelling regions from the ROI regions as dwelling target images; obtain dwelling edge images by edge detection on the dwelling target images; perform Hough line transform on the dwelling edge images; obtain dwelling edge image including three pairs of side lines in the outline edge of the dwellings through heuristic search, wherein the first pair of side lines in the three pairs of side lines are two opposite side lines of the top and bottom of the dwellings sandwiched by a pair of side edges, and the second and third pairs of side lines in the three pairs of side lines are two opposite side lines sandwiched by a pair of horizontal long sides and a pair of horizontal wide sides, respectively; obtain the geometric features of the two opposite side lines relative to the image center of any pair of side lines in the dwelling edge image, input the geometric features into a trained neural network-based dwelling size prediction model to obtain dwelling size prediction values, and output the size prediction values.
[0145] Preferably, the UAV-based rural dwelling identification system 1000 may not include the server 400 and handheld terminal 300. Instead, during application, the UAV-based rural dwelling identification system is connected to the server 400 and handheld terminal 300 via a communication network. The UAV-based rural dwelling identification system 1000 may also include a positioning unit 700 for coordinate acquisition.
[0146] Preferably, the user interface unit 200 includes a display screen 210 and an operation panel 220. The UAV-based rural dwelling identification system 1000 also includes a flight control platform 600, which executes commands and flies toward the designated residential location in response to user operations on the operation panel 220 or instructions from the host unit 100.
[0147] The UAV-based rural dwelling identification system 1000 continuously detects and predicts the size of multiple frames of rural dwelling images. When the safe distance between the dwelling sites meets the requirements, it compares the sizes of multiple dwelling sites and determines the site with the largest safe distance as the target landing point. The system outputs the coordinate position information of the target landing point to the flight control platform 600 through the output unit. The flight control platform then manipulates the power unit to make the UAV fly towards the target landing point.
[0148] Preferably, when drones are used in group systems such as logistics and transportation, the system can have multiple handheld terminals and drone platforms. Through a server, the host unit can send messages at residential scale to remote clients deployed on handheld terminals, etc. Preferably, an MQTT server can be deployed to push messages to clients subscribed to the topic. Multiple devices can subscribe to multiple topics simultaneously, and when a message is published, those subscribed to the same topic will all receive the same message. The publish-subscribe pattern is a messaging pattern that decouples the client sending the message (publisher) from the client receiving the message (subscriber), so that they do not need to establish direct contact or be aware of each other's existence.
[0149] The server provides subscription and publish services. The host unit is also configured to: subscribe to the topic of obtaining residential information from the handheld terminal, and in response to the topic message, transmit the detected residential dimensions and corresponding coordinates and location information to the handheld terminal through the server; or transmit the residential dimensions and corresponding coordinates and location information to the server in the form of publishing a residential dimension event topic message when the residential dimensions are detected to reach a set safe distance threshold.
[0150] Preferably, the user interface unit is provided with an operation panel for parameter input, and a display screen for assisting parameter input and displaying the residential building size recognition results in the form of a list; the host unit also displays the recognized residential building size in the form of a curve on the display screen, and marks the coordinates and positions corresponding to the recognized size on the curve.
[0151] Preferably, the positioning unit 700 in the system also acquires heading information, and records and identifies the orientation identified based on the heading, such as due north, and the orientation corresponding to the roof axis of the residential building in the output location information.
[0152] Alternatively, the host unit can also communicate directly with the handheld terminal via a wireless network.
[0153] The host unit can be a standalone local control unit deployed on the ground, or it can be a control unit deployed on a drone platform. When deployed on a drone platform, the host unit on the drone platform performs edge computing for online residential size prediction. Due to the limited resources of edge computing units, the processing and computing of the host unit are optimized.
[0154] In a typical Hough line transform, for each edge point (x, y) in the image space, the value of ρ = xcosθ + ysinθ under different θ values is calculated, and the corresponding position in the accumulator array is voted. This process of calculating ρ requires two floating-point multiplication operations.
[0155] Therefore, based on the characteristic that the boundary straight lines in rural dwellings are relatively obvious, the pixel checking and accumulation process of the Hough line transform is optimized. Specifically, when performing the Hough line transform on each foreground pixel (u, v) in the edge image of the dwelling, in the accumulation statistics, it is determined whether the pixel is adjacent to a straight line element l. i,j (ρ i,j θ i,j The corresponding judgment method is: take a fulcrum H(ρ) on the straight line. i,j ·cosθ i,j , ρ i,j ·sinθ i,j Let H(u′, v′) be the variable, and calculate J = |vv′-(uu′)·tanθ. i,j If |J|≤ε, then the pixel is determined to be on the straight line and the points are accumulated, where i and j are the row and column indices of the ρ-θ plane matrix, and ε is a preset small positive number.
[0156] By setting a fulcrum, a tangent operation is performed once before traversal judgment. Thus, each time a foreground pixel is judged, only two subtraction operations and one multiplication operation are needed to perform the judgment and accumulation. This avoids the need to perform two time-consuming floating-point operations for each pixel, further saving computational costs. Therefore, it is suitable for edge computing deployed on drone platforms.
[0157] As a preferred approach, in the heuristic search for obtaining the edge lines of residential buildings, when the slope represents the angle in a rectangular coordinate system, if the absolute value of the difference in the slope of the edge lines is less than a very small first slope threshold, the two edge lines are merged. At the same time, if the tangent of the angle or the slope is greater than a very large second slope threshold, the edge line is determined to be a vertical line perpendicular to the horizontal axis. This prevents the slight differences in slope from causing the edge lines to be judged as different straight lines, thus avoiding the misdetection interference of short edge lines on the straight lines of residential building edges.
[0158] like Figure 4A As shown, in this preferred embodiment, during sample collection, samples can be recorded in groups from the two sets of opposite edges on the front and top surfaces of the cube, respectively. The output layer corresponds to the height, length, and width, respectively. Based on the subdivided sample sets, different neural networks can be established for the height, length, and width of the dwellings, and trained based on the corresponding sample sets. Preferably, the width and length on the horizontal plane can also be sampled and recorded together, without distinguishing between length and width.
[0159] Unlike Embodiment 1, this embodiment further includes an artificial neural network for the residential building size prediction model. The input layer receives six input parameters from the general processing module: the slopes k1 and k2 of the straight lines on both sides of the residential building, the perpendicular distances D1 and D2 from the image center point to the straight lines on both sides of the residential building, and the flight altitude and overhead angle during data acquisition. The output layer outputs the residential building size. After establishing this neural network structure, the input parameters for each sample in the training set are supplemented. After supplementing with a sufficient number of samples from different altitudes and overhead angles, the new training sample set is used for training. In online applications, the new model further simplifies the image sampling requirements for rural residential building prediction. Utilizing the generalization performance of neural networks, it reduces the number of models required for rural residential building size prediction and simplifies the system structure, enabling a single model to cover different flight sampling scenarios, thus further improving the system's applicability.
[0160] Example 3
[0161] Based on the above embodiments, this embodiment provides a controller for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV). Combined with... Figure 2 , Figure 3A , Figure 4BAs shown, the UAV-based rural dwelling identification controller 100 includes a processor 110, a storage module 120, and an input / output module 130. The processor 110 includes a general processing module 111, an acquisition and preprocessing module 112, a feature extraction module 113, an iterative learning module 114, and a dwelling size prediction module 115. The general processing module 111 responds to messages and user interface unit operations, initiates the identification of dwellings in the acquired images, schedules the identification processing flow, inputs the real-time acquired rural dwelling images into the acquisition and preprocessing module 112, and outputs the size data identified and predicted by the dwelling size prediction module 115.
[0162] Combination Figure 3A , Figure 4B As shown, the image of rural dwellings is input to the controller via the input unit 131 in the input / output module 130; the dwelling size prediction model established in the dwelling size prediction module 115 receives the geometric features of the dwellings in the image of rural dwellings extracted by the feature extraction module 113 from the general processing module 111, and the output layer is transmitted to the iterative learning module 114 and the general processing module 111 through the first connection array 1141 and the second connection array 1142, respectively.
[0163] When training the residential building size prediction model offline, the iterative learning module 114 adjusts the connection weights of the network based on the actual residential building size values input by the general processing module 111 and the residential building size prediction model through the first connection array 1141 and the output values of the neural network in the model, respectively. In the field environment, the first connection array 1141 is disconnected, and the residential building size prediction model neural network predicts the size of the residential buildings in the current rural residential building image and outputs it to the general processing module 111 through the second connection array 1142. After being processed and analyzed by the general processing module 111, it is output through the output unit 132 in the input-output module 130.
[0164] The iterative training steps include: STEP1, initializing the network weights and biases; STEP2, reading the network parameters and the training sample set previously written to the document; STEP3, normalizing the sample data; STEP4, calculating the error between the predicted output value and the expected value for each sample; STEP5, back-calculating and correcting the network weights and biases; STEP6: if the conditions for ending network training are met, end the training; otherwise, return to STEP4 to continue training; STEP7, result analysis and output.
[0165] The general processing module 111 in the controller also obtains the coordinate and orientation information of the current location of the dwelling through the positioning unit 700. After identifying the size of the dwelling, it outputs the size of the dwelling and the corresponding coordinate and orientation information through the output unit 132.
[0166] like Figure 15 As shown, after identifying the dimensions of the dwellings, the dimensions are also marked within the corresponding extracted edge segments of the dwellings in the rural dwelling image, such as using arrows to mark the corresponding directional line segments. Preferably, the acquired coordinates and location information, as well as the dwelling dimensions, are stored in a list in the storage module 120, and the list information can be stored in the cloud server 400.
[0167] When the length and width in the size data exceed a preset safety distance threshold, the coordinates of the center point of the house roof are given; otherwise, an alarm is issued, indicating that the safety distance to the house roof is insufficient. During the drone's flight, images are continuously acquired iteratively. The rural house identification controller 100 in the drone's field of view performs size prediction processing based on the acquired images. When a house with both length and width exceeding the preset safety distance threshold appears within the field of view, an alarm is issued indicating that a house with sufficient safety distance has been found. Preferably, the processed images are updated according to the flight speed and frame rate, so that houses in the field of view are identified one by one.
[0168] After identifying a site with a sufficient safe distance, the safe distance can be verified by acquiring images of residential buildings from multiple angles and predicting their sizes. When flying towards the target residential building, the anchor frame of the rural scene in the new field of view is updated after the rural scene image is identified and processed based on the rural scene target classification model. Flight control is then performed based on the position of the anchor frame of the residential building in the field of view, so that the residential building continuously approaches the center area of the field of view.
[0169] In the heuristic search for boundary lines of residential buildings, some buildings have multiple stepped structures, resulting in multiple pairs of parallel edges. In this case, there may be more than two candidate line groups. The safe distance is determined by the minimum length or width corresponding to the boundary line formed by the pair; for height, the safe distance is determined by the maximum height.
[0170] Preferably, if the ground is flat as the landing point, then in one of the candidate straight line groups, the two opposite edge lines of the candidate straight lines of residential ground rather than roofs sandwiched between a pair of other straight line segments are used to form the second pair of edge lines and the third pair of edge lines, so as to further obtain the geometric features of the ground.
[0171] Example 4
[0172] Unlike the above embodiments, this embodiment provides another method for obtaining the dimensions of the long and wide sides of a dwelling.
[0173] See Figure 1 and combined Figure 7EAs shown, in the boundary of the three-dimensional dwelling, in addition to the horizontal parallel line segments, there are also vertical side edges that can enclose the long or wide side of the dwelling. Therefore, in this embodiment, based on the obtained side edges of the dwelling, two opposing side edges of the dwelling with a pair of other straight line segments sandwiched in between are used to form the second pair of side lines and the third pair of side lines.
[0174] In this case, as a preferred option, the slope of the boundary line, i.e. the side edge, in the input of the neural network is replaced with the angle of rotating counterclockwise from the 0-degree direction to the boundary line, thereby preventing errors caused by using the slope due to the angle being close to 90 degrees.
[0175] This invention is applied to identify the dimensions of rural dwellings in the image space under the view of an unmanned aerial vehicle (UAV). After segmenting the dwelling area, extracting the dwelling edges, and determining the boundary lines, four geometric features are input into the established neural network-based dwelling size prediction model. These features include the slopes k1 and k2 of a pair of boundary lines corresponding to the length, width, and height sides, as well as the perpendicular distances D1 and D2 from the image center point to the boundary lines. The model can accurately predict the actual dimensions of rural dwellings in the three dimensions of length, width, and height, thus providing a decision-making basis for UAV landing and flight, and improving the safety and convenience of UAV operation.
[0176] The foregoing has described several embodiments of the present invention, but these embodiments are merely illustrative examples and do not limit the scope of the invention. These embodiments can be implemented in various other ways, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments or their variations are included within the scope or spirit of the invention, and are similarly included within the scope of the invention as described in the claims and its equivalents.
Claims
1. A device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV), comprising a host unit and a user interface unit, wherein the host unit includes an input / output module, a processor, and a storage module, and is configured as follows: A rural dwelling image is acquired, and a Region of Interest (ROI) containing the dwellings is determined within the rural dwelling image. The dwelling areas are extracted from the ROI as the dwelling target image. Edge detection is performed on the dwelling target image to obtain dwelling edge images. A Hough transform is applied to the dwelling edge images, and a heuristic search is used to obtain dwelling edge image images including three pairs of edges in the dwelling contour. The first pair of the three pairs of side lines are the two opposite side lines of the top and bottom of the house with a pair of side edges sandwiched in the middle. The second and third pairs of the three pairs of side lines are the two opposite side lines with a pair of horizontal long sides and a pair of horizontal wide sides sandwiched in the middle, respectively. From the image of the residential building's edge lines, for any one of the three pairs of edge lines, obtain the geometric features of the two opposite edge lines relative to the image center. Input the geometric features into a trained neural network-based residential building size prediction model to obtain the predicted residential building size value, and output the predicted size value.
2. The device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV) according to claim 1, characterized in that, The processor includes a general processing module, an acquisition and preprocessing module, a feature extraction module, an iterative learning module, and a residential building size prediction module. The acquisition and preprocessing module acquires images of rural dwellings from the image acquisition unit. After filtering, the location of the dwellings in the image is marked as a Region of Interest (ROI) by a trained target classification model based on deep learning features. After removing the portion outside the ROI region from the original rural house image to form the house target image, edge detection is performed to form a binarized house edge image. After performing Hough transform on the binarized house edge image, the vertical and horizontal line segments with the longest lengths are selected as candidate line segments based on the polar coordinate space parameter features in the Hough transform. Based on the orientation features of the candidate line segments, two opposite lines with a pair of long sides, wide sides, and high sides sandwiched in the middle are determined as a pair of side lines for the length, width, and height of the house, respectively, and together they form three pairs of side lines for the house. For each edge of any pair of the three pairs of edge lines, the feature extraction module combines the corresponding binarized residential edge image, describes it using the coordinates of the two endpoints of the line, and calculates the slope of the line and the perpendicular distance from the image center point to the line based on the two endpoints; the slope of the line on both sides of the residential building and the perpendicular distance from the image center point to the line are used as the geometric features of the residential building in the rural residential building image. The iterative learning module trains the neural network of the residential building size prediction model in the residential building size prediction module based on the collected training samples. The general processing module responds to the message, initiates the identification of dwellings in the acquired images, schedules the identification processing flow, and outputs the size data identified and predicted by the dwelling size prediction module.
3. The device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV) according to claim 2, characterized in that, The acquisition and preprocessing module includes a target detection unit. Within this unit, a target classification model based on the ASFF-YOLOv5s network is established. After capturing rural scene images using a top-down camera in the image acquisition unit, the network is trained offline using an image sample set consisting of these rural scene images labeled with rural residential areas to obtain the rural scene target classification model. During online operation, the acquisition and preprocessing module uses the rural scene target classification model to identify and process the rural scene images (including residential buildings) to obtain anchor boxes for rural residential areas. The areas within these anchor boxes are designated as the Regions of Interest (ROIs) for rural residential areas. The ASFF-YOLOv5s network adds three layers, 24, 25, and 26, between the 23rd layer of the YOLOv5s network and the layer below it. These three layers receive the outputs of the 17th, 20th, and 23rd layers as inputs, respectively. After feature fusion at different scales by the ASFF adaptive spatial feature fusion module, they successively replace the original 17th, 20th, and 23rd layers as inputs to the original output layer, the 24th layer, which is the 27th layer of the new network.
4. The device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV) according to claim 1, characterized in that, The residential building size prediction model uses an artificial neural network. Its input layer receives four geometric feature inputs from the general processing module: the slopes k1 and k2 of a pair of sides corresponding to the length, width, and height of the residential building, and the perpendicular distances D1 and D2 from the image center point to the pair of sides. The output layer outputs the length, width, and height of the residential building corresponding to the pair of sides.
5. The device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV) according to claim 1, characterized in that, The host unit is further configured to: when performing the Hough transform on each edge pixel in the residential edge image, accumulate to obtain the corresponding ρ-θ plane matrix of the image, where ρ represents the distance of the perpendicular segment of the line from the origin, and θ is the angle of the perpendicular line to the horizontal axis; in the accumulation, the foreground pixel (u, v) in the image is connected to a straight line element l i,j (ρ i,j θ i,j The corresponding judgment method is: Take a fulcrum H(ρ) on this straight line i,j ·cosθ i,j , ρ i,j ·sinθ i,j Let H(u′, v′) be the variable, and calculate J = |vv′-(uu′)·tanθ. i,j If |J|≤ε, then the pixel is determined to be on the straight line and the sum is accumulated, where ε is a preset small positive number.
6. The device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV) according to claim 1, characterized in that, The host unit is also configured to: when performing the Hough line transform on each edge pixel in the residential edge image, accumulate to obtain the ρ-θ plane matrix corresponding to the image, where ρ represents the distance of the perpendicular line segment from the origin to the line, and θ is the angle of the perpendicular line to the horizontal axis. Based on the accumulated ρ-θ plane matrix, within a range close to the vertical direction, search for the two or three longest edges as the side edges of the dwellings; Within the range of non-side edge directions, the line segments are sorted from largest to smallest according to the values of the plane matrix elements, i.e., the lengths of the line segments. Starting from the main line segment with the largest length, two secondary line segments that are close to its direction are searched in turn. The main line segment and the two corresponding secondary line segments form a candidate line group. For any line segment in the candidate line group, calculate its distance to the upper and lower endpoints of the side edge of the residential building. If the minimum value of the distance corresponds to the distance between the line segment and the upper endpoint of a side edge, then the line segment is marked as a candidate line for the roof; otherwise, it is marked as a candidate line for the ground.
7. The device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV) according to claim 6, characterized in that, The host unit is further configured to: in one of the candidate straight line groups, form the first pair of side lines by two opposing side lines: a candidate straight line of a residential roof sandwiched between a pair of side edges and a candidate straight line of the ground. In one of the candidate straight line groups, the second pair of side lines and the third pair of side lines are formed by two candidate straight lines of residential roofs sandwiched between two other straight line segments.
8. The device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV) according to claim 1, characterized in that, The host unit is further configured to calculate and extract geometric features for any pair of the three pairs of sides. Assuming the coordinates of the two endpoints of a residential building's boundary line are (x1, y1) and (x2, y2), then its slope is k = y1 - y2 / x1 - x2; The perpendicular distance from the center of the image to the edge of the house is In the formula, (x0, y0) represents the coordinates of the center of the image, A and B are the coefficients of the general form of the equation of a straight line, and C is a constant term and: A = y2 - y1, B = x1 - x2, C = -x1(y2 - y1) + y1(x2 - x1); Based on the calculation formulas for k and d, the slopes k1 and k2 of any pair of edges of the dwellings, as well as the perpendicular distances D1 and D2 from the image center point to the two edges, are calculated as four quantities, which are used as the geometric features of the dwellings in the rural dwelling image.
9. The device for identifying rural dwellings under the view of an unmanned aerial vehicle (UAV) according to claim 1, characterized in that, The residential building size prediction model uses a backpropagation neural network, which has a three-layer structure including one hidden layer. The output of the j-th node in the hidden layer is The output of the first node in the output layer is Where f() is taken as the tansig function, w ij and v j These are the connection weights from the input layer to the hidden layer and the connection weights from the hidden layer to the output layer, respectively. j b and are the thresholds for the hidden layer and output layer, respectively. m = 4 is the number of input variables, and k is the number of hidden layer nodes. Gradient descent is used for network training.
10. A rural dwelling identification system based on UAV vision, comprising: An image acquisition unit for sensing and acquiring images of rural scenes; a user interface unit for parameter input and operation, and display of rural residential information; a server for data management and information transmission; a handheld terminal for remote operation and interaction with rural residential information; and a host unit connected to the user interface unit, server, and image acquisition unit. The host unit includes an input / output module, a processor, and a storage module, and is configured as follows: A rural dwelling image is acquired, and a Region of Interest (ROI) containing the dwelling is determined in the rural dwelling image. The dwelling area is extracted from the ROI as the dwelling target image. The dwelling target image is then processed by edge detection to obtain a dwelling edge image. The dwelling edge image is then subjected to Hough transform, and a heuristic search is used to obtain a dwelling edge image including three pairs of edges in the outline edge of the dwelling. The first pair of edges in the three pairs of edges are the two opposite edges of the top and bottom of the dwelling sandwiched between a pair of side edges. The second and third pairs of edges in the three pairs of edges are the two opposite edges sandwiched between a pair of horizontal long edges and a pair of horizontal wide edges, respectively. From the image of the residential building's edge lines, for any one of the three pairs of edge lines, obtain the geometric features of the two opposite edge lines relative to the image center. Input the geometric features into a trained neural network-based residential building size prediction model to obtain the predicted residential building size value, and output the predicted size value.
Citation Information
Patent Citations
Space measuring and calculating method based on laser and images
CN110879057A
Building Height Monitoring Method Based on High-Resolution Optical Remote Sensing Satellite Imagery with Corner Points
CN113139994B
Method for identifying, acquiring and calculating ratio of length unit to pixel value of CAD (Computer Aided Design) drawing
CN116343253A
On-site survey method for photovoltaic power station
CN116659469A
An edge connectivity algorithm that can be implemented in parallel starting from a breakpoint
CN102270299A