A land change evidence generation method based on a tower unmanned aerial vehicle and a nest
Through the coordinated scheduling of tower drones and drone nests, automated evidence collection for land change surveys has been achieved, solving the problems of low efficiency, resource waste, and unstable quality in complex terrain, and generating high-quality map patch evidence results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-22
- Publication Date
- 2026-06-26
Smart Images

Figure CN122289989A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of land change survey technology, and more specifically, to a method for generating evidence of land change based on a tower-mounted drone and its nest. Background Technology
[0002] In land use change surveys, field verification is required for surface change patches identified by satellite remote sensing to confirm their actual land use information. Currently, this field verification work mainly relies on manual on-site photography and recording of land use information.
[0003] However, in complex terrain conditions (such as fragmented farmland in mountainous areas), a large number of land parcels are scattered and have poor accessibility, leading to the following technical problems with traditional manual evidence collection methods: First, field personnel need to visit each land parcel individually, resulting in high resource investment and low evidence collection efficiency; second, the shooting angle, location, and content of evidence photos depend on personnel experience, leading to inconsistent quality of results; third, the selection of land parcels, flight path planning, and image analysis lack automated information technology support, and data transfer between each stage relies on manual operation, resulting in insufficient overall process coordination, which leads to low efficiency and difficulty in guaranteeing the quality of land change survey evidence collection. Summary of the Invention
[0004] This invention provides a method for generating evidence of land change based on tower drones and drone nests, solving the technical problems of low efficiency, limited coverage, and difficulty in resource allocation in land change surveys in related technologies.
[0005] This invention discloses a method for generating evidence of land use changes based on tower-mounted drones and their nests, comprising the following steps: Obtain the data of the map features to be used as evidence, filter and classify them based on the map feature attributes and spatial location, and generate a set of target map features for drone evidence collection; Obtain the spatial coordinates and coverage parameters of each tower site within the target area, calculate the effective coverage area of each tower site, spatially match the spatial coordinates of each patch in the UAV evidence target patch set with the effective coverage area of each tower site, assign each target patch to the tower site covering the patch location, and generate the evidence task set corresponding to each site. For each tower site's evidence collection task set, obtain the spatial coordinates and geometric shape data of each target patch within the evidence collection task set. Using the tower site location as the starting and ending point, under the constraint of the UAV's maximum flight range, sort and group the flight routes of each shooting point to generate evidence collection flight routes covering all shooting points. All evidence collection flight routes constitute the flight route scheme for that tower site. The flight control and dispatch platform sends the flight path plan to the drone nest at the target tower site, and the drone is dispatched to perform fixed-point and directional shooting according to the flight path plan to collect evidence image data. The evidence image data is sequentially subjected to photo quality screening, land use information extraction, and image spatialization processing to generate evidence data for map features.
[0006] Furthermore, the filtering and classification based on map patch attributes and spatial location includes: Based on the exclusion rules for monitoring types, facility agricultural plots and greenhouse plots that require indoor confirmation are excluded from the scope of drone evidence. Based on the exclusion rules of spatial location, construction land parcels located in densely built-up areas and land parcels located within military land management areas are excluded from the scope of drone evidence. Based on the filtering rules for patch area, patches with an area lower than the area threshold are marked as patches that require manual confirmation.
[0007] Furthermore, the effective coverage area of the tower site is defined as a circular area with the spatial coordinates of the tower site as the center and the effective coverage radius as the radius. The effective coverage radius is determined by a combination of the height of the tower site, the maximum range of the deployed UAV nest, and terrain obstruction factors. When the spatial coordinates of a target patch are simultaneously located within the effective coverage area of multiple tower sites, the patch is sorted according to the Euclidean distance from the center point of the patch to each candidate tower site, and the patch is assigned to the nearest tower site.
[0008] Furthermore, the step of sorting and grouping the flight paths of each shooting point under the constraint of the drone's maximum flight range includes: The total flight distance of a single evidence-providing route is constrained by the sum of the inter-segment distances between each shooting point, the flight distance of the drone from the tower station to the first shooting point, and the flight distance of the drone from the last shooting point back to the tower station, which shall not exceed the maximum endurance of the drone in a single flight. The order of visiting each shooting point is obtained through a traveling salesman problem algorithm. The input is the set of spatial coordinates of each shooting point, and the output is the order of visiting points that minimizes the total flight distance. When a single evidence-providing route cannot cover all shooting locations within the range constraint, the remaining shooting locations are allocated to the next evidence-providing route according to a greedy strategy until all shooting locations are covered.
[0009] Furthermore, the method for determining the shooting location includes: For patches whose area is smaller than the observation area threshold, the geometric center of the target patch is used as the single shooting point; For patches whose area is greater than or equal to the observation area threshold, multiple shooting points are set at equal intervals along the long axis of the target patch so that the images captured by each shooting point can cover the entire patch area.
[0010] Furthermore, the scheduling of the drone to perform fixed-point and directional shooting according to the route plan includes: The drone flies sequentially according to the waypoint sequence in the flight path plan. After arriving at each shooting point, it performs fixed-point and directional shooting according to the shooting parameters preset in the flight path plan. The shooting parameters include shooting altitude, lens pitch angle and heading angle. The shooting height is calculated based on the ratio between the camera focal length and the sensor size in the flight direction, as well as the actual ground length of the target patch in the flight direction. The lens tilt angle is adjusted according to the terrain slope. In flat areas, a vertical overhead angle is used, while in sloping areas, the lens deflects towards the slope normal. The amount of deflection is positively correlated with the terrain slope.
[0011] Furthermore, the photo quality screening includes: Each of the collected evidence images was evaluated for quality. The quality evaluation indicators included image sharpness score, exposure uniformity score, and target area coverage score. The overall quality score is obtained by weighted summation of three scoring indicators. Images with an overall quality score below the preset quality threshold are marked as unqualified and removed, while qualified images are retained to form a set of qualified evidence images. Among them, the image sharpness score is obtained by calculating the variance of the Laplacian operator response of the image grayscale image, the exposure uniformity score is obtained by calculating the ratio of the mean to the standard deviation of the image pixel brightness values, and the target area coverage score is the proportion of the number of pixels in the target patch area to the total number of pixels in the image.
[0012] Furthermore, the extraction of land use information includes: Each image in the qualified evidence image set is input into the pre-trained land use identification model, and the class label of the land features in each image and the corresponding confidence score are output. The land use identification model is an image classification model based on a convolutional neural network. The input layer receives the pixel matrix of a single evidence image, extracts the local texture features and spatial structure features of the image through multiple convolutional and pooling layers, and then maps the features to the category space of the preset land use category system through a fully connected layer. The output layer uses the Softmax function to convert the scores of each category into probability distribution vectors. The land cover identification results of multiple images corresponding to the same target patch are comprehensively judged, and the category label with the highest frequency and average confidence exceeding the preset confidence threshold is taken as the land cover identification result of the target patch.
[0013] Furthermore, the image spatialization processing includes: Obtain the shooting location coordinates, shooting height, lens parameters and attitude parameters of each qualified evidence image, and calculate the ground coverage of each qualified evidence image based on the interior and exterior orientation elements of the image. Spatial overlay analysis is performed between the ground coverage area and the vector boundary of the target patch to calculate the proportion of the spatial intersection area between the ground coverage area and the vector boundary of the patch to the total area of the patch. When the ratio reaches the preset coverage threshold, a spatial association is established between the qualified evidence image and the target patch, and the land use identification result, the spatially associated qualified evidence image and the patch attribute information are encapsulated into structured patch evidence result data.
[0014] This invention provides a land use change evidence generation system based on tower-mounted drones and their nests, comprising: The image patch filtering module is used to acquire image patch data to be used as evidence, filter and classify them based on image patch attributes and spatial location, and generate a set of target image patches for drone evidence. The task allocation module is used to obtain the spatial coordinates and coverage parameters of each tower site, calculate the effective coverage area of each tower site, allocate the target patch to the corresponding tower site, and generate the evidence task set corresponding to each site. The flight path planning module is used to obtain the spatial coordinates and geometric shape data of each target patch within the evidence task set, sort and group the flight paths of each shooting point under the constraint of the maximum flight range of the UAV, and generate flight path schemes. The flight control and scheduling module is used to send flight path plans to the drone nests at the target tower site, and to schedule the drones to perform fixed-point and directional shooting according to the flight path plan to collect evidence image data. The intelligent analysis module is used to sequentially perform photo quality screening, land use information extraction, and image spatialization processing on the evidence image data to generate map patch evidence results data.
[0015] This invention coordinates the use of tower infrastructure resources and the automated operation capabilities of drones, and introduces automated screening, planning, and intelligent analysis processes in the data processing stage, forming a complete data processing link from map patch screening to evidence output. This solves the technical problems of low efficiency, high resource investment, and insufficient information technology support in land change surveys, achieving the following technical effects: First, automated screening and classification avoids redundant drone operations on invalid map patches, reducing resource investment in evidence collection tasks; second, using tower sites as drone take-off and landing nodes enables automated evidence collection under unattended conditions, overcoming the time consumption and traffic limitations of traditional manual evidence collection where field personnel arrive at each point, thus improving evidence coverage efficiency; third, obtaining consistent evidence images through fixed-point, directional automated shooting ensures stable image acquisition quality; and fourth, outputting evidence data containing land type information and spatial correspondences through quality screening, land type identification, and spatial image processing meets the verification requirements of land change surveys. Attached Figure Description
[0016] Figure 1 This is a flowchart of the land change survey evidence collection method provided in the embodiments of the present invention. Detailed Implementation
[0017] At least one embodiment of the present invention discloses a method for generating evidence of land use change based on a tower-mounted drone and its nest, such as... Figure 1 As shown, it includes the following steps: In land use change surveys, field verification is required for surface change patches identified by satellite remote sensing to confirm their actual land use information. Currently, field verification mainly relies on manual on-site photography and recording of land use information. However, in complex terrain (such as fragmented farmland in mountainous areas), a large number of patches are scattered and inaccessible, leading to the following technical problems with traditional manual verification methods: First, field personnel must visit each patch individually, resulting in high resource consumption and low verification efficiency; second, the shooting angle, location, and content of the verification photos depend on personnel experience, leading to inconsistent quality of results; third, the screening of patches, flight path planning, and image analysis lack automated information technology support, with data transfer between stages relying on manual operation and insufficient overall process coordination. Therefore, a collaborative verification method is needed that can utilize existing infrastructure resources to automatically complete patch screening, flight path planning, image acquisition, and intelligent analysis to improve the efficiency and quality of land use change survey verification.
[0018] According to an embodiment of this method, the hardware environment on which this method relies includes: multiple communication tower base stations distributed within the target area, each communication tower base station deploying a drone nest, a high-position video probe, and IoT sensor equipment; the drone nest is equipped with automatic charging, automatic take-off and landing, and data transmission functions; the flight control scheduling platform, as a central scheduling node, maintains data connection with the drone nests and IoT sensor equipment on each tower base station through a communication network. It should be understood that the tower base stations in this method simultaneously provide power supply and communication transmission channels, enabling the drone nests and IoT sensor equipment to operate continuously.
[0019] Step 1: Obtain the data of the map features to be used as evidence, filter and classify them based on the map feature attributes and spatial location, and generate a set of target map features for drone evidence collection; The system acquires the map parcel data to be used for evidence collection from the change investigation and monitoring system. Each map parcel data includes attribute fields such as parcel number, monitoring type, parcel area, spatial coordinates, and land use code. Based on a preset set of filtering rules, the system iterates through and judges the map parcel data to be used for evidence collection. Map parcels that meet the conditions for drone-based evidence collection are marked as target map parcels, while map parcels that do not meet the conditions are marked as manually collected map parcels. The system then outputs a set of target map parcels for drone-based evidence collection.
[0020] It should be noted that the aforementioned set of preset screening rules refers to a combination of multiple predefined judgment rules based on the attributes and spatial location characteristics of the map features. Specifically, the screening rules include at least the following types: First, exclusion rules based on monitoring type, excluding agricultural facility map features and greenhouse map features requiring indoor confirmation from the scope of drone evidence collection; Second, exclusion rules based on spatial location, excluding construction map features located in densely built-up areas and map features located within military land management areas from the scope of drone evidence collection; Third, filtering rules based on map feature area, marking map features with an area below a certain threshold as map features requiring manual confirmation, where the area threshold is a pre-set fixed value used to distinguish the minimum map feature size suitable for drone aerial photography. In other words, after screening, map features that do not meet the conditions for drone evidence collection remain in the system and are subsequently processed by the manual evidence collection process.
[0021] Step 2: Based on the coverage area of the tower sites and the spatial distribution of the target patches, assign the target patches to the corresponding tower sites and generate the evidence task set for each site; Obtain the spatial coordinates and coverage parameters of each tower site within the target area, and calculate the effective coverage area of each tower site. Spatially match the spatial coordinates of each patch in the UAV evidence target patch set with the effective coverage area of each tower site, assign each target patch to the tower site covering the patch location, and generate the evidence task set corresponding to each site.
[0022] Among them, the tower site Effective coverage area Defined as a circle centered on the spatial coordinates of the tower site, with an effective coverage radius. A circular region with radius , covering an area of . , For tower site indexing, The determination is based on a combination of factors, including the height of the tower site, the maximum range of the deployed drone nests, and terrain obstruction.
[0023] It should be noted that in the above spatial matching process, when the spatial coordinates of a target patch are simultaneously located within the effective coverage area of multiple tower sites, the patch is sorted according to the Euclidean distance from its center point to each candidate tower site, and the patch is assigned to the nearest tower site. Specifically, for the target patch... Its distribution site The method for determining it is as follows:
[0024] in, To cover the patch The collection of candidate tower sites, For the image From the center point to the tower site The Euclidean distance between them Index the target image patch.
[0025] In this embodiment of the application, to handle situations where some target patches are not within the effective coverage area of any tower site, the following processing is included in addition to step 2: Patches in the drone-provided evidence target patch set that are not covered by any tower site are marked and added to the pending evidence list. For patches in the pending evidence list, the distance difference between them and the nearest tower site is calculated. When the distance difference is within a preset range extension threshold, the patch is included in the evidence task set of the nearest tower site, and the range margin is increased accordingly in subsequent route planning; when the distance difference exceeds the range extension threshold, the patch is transferred to the manual evidence process. The range extension threshold is a pre-set fixed value, representing the maximum distance allowed to extend beyond the drone's maximum range, used to determine whether the boundary patch is feasible for inclusion in drone-provided evidence.
[0026] Step 3: Based on the spatial distribution of map features in the evidence collection task set of each site and the drone's endurance constraints, plan the drone evidence collection route and generate the route plan; For each tower site's evidence collection task set, acquire the spatial coordinates and geometric shape data of each target patch within the task set. Based on the geometric center of the patch or preset shooting observation points, using the tower site location as the start and end point, within the maximum flight range of the UAV... Under the constraints, flight routes are sorted and grouped for each shooting location, generating several evidence flight routes covering all shooting locations. Each evidence flight route includes an ordered sequence of waypoints, shooting parameters corresponding to each waypoint, and expected flight distance information. All evidence flight routes constitute the flight route plan for the tower site.
[0027] Among them, for the tower site The burden of proof set includes The set of target image patches and corresponding shooting locations is as follows: ,in The first The first to the second One shooting location. The total flight distance constraint for a single evidence-providing route is:
[0028] in, Arranged in the order of visits to the shooting locations. The number of locations covered by the evidence-providing route. To sum and traverse the indices, The flight distance between the two points. This represents the maximum range of a single drone flight.
[0029] Furthermore, in the above range constraint formula, the summation term on the left side... This represents the sum of the distances between segments as the drone flies sequentially between each shooting location. Indicates that the drone departed from the tower site. Flight distance from departure to the first shooting location, This indicates that the drone returned to the tower site from its last shooting location. The flight distance, the sum of the three factors constitutes the total round-trip distance of the evidence route, and this total round-trip distance must not exceed the maximum endurance limit of a single flight of the drone. .
[0030] It should be noted that the above method of determining shooting points refers to determining one or more shooting observation points for each target patch based on its geometry and area. For patches with an area smaller than the observation area threshold, the geometric center of the target patch is used as a single shooting point; for patches with an area greater than or equal to the observation area threshold, multiple shooting points are set at equal intervals along the long axis of the target patch, so that the images captured by each shooting point can cover the entire patch area. The observation area threshold is a pre-set fixed value used to distinguish the size boundary of patches requiring single-point shooting from those requiring multi-point coverage shooting.
[0031] It should be noted that the order of visiting each shooting point in the above route sequencing is obtained through a Traveling Salesman Problem algorithm. The input is the set of spatial coordinates of each shooting point, and the output is the order of visiting points that minimizes the total flight distance. When a single evidence route cannot cover all shooting points within the flight range constraint, the remaining shooting points are allocated to the next evidence route according to a greedy strategy until all shooting points are covered.
[0032] Furthermore, the specific execution process of the above greedy strategy is as follows: starting from the last shooting point of the currently planned evidence-gathering route, select from the remaining unassigned shooting points the closest to the current shooting point and whose addition does not exceed the range constraint. The selected shooting locations will be added to the current evidence-providing flight path; if no candidate shooting locations meet the endurance constraints, the current evidence-providing flight path will be closed back to the tower station. With the tower site Start a new evidence collection route from the beginning, and repeat the above process until all shooting locations are assigned to a certain evidence collection route.
[0033] In this embodiment of the application, to improve the efficiency of evidence collection for multiple adjacent patches in a single flight, the following processing is included in addition to step 3: Cluster analysis is performed on spatially adjacent patches within the evidence collection task set. The input is the spatial coordinates of each patch, and the output is several spatial clusters. The maximum distance between patches within each spatial cluster does not exceed a cluster distance threshold, which is a pre-set fixed value used to control the maximum allowable spacing between patches within the same spatial cluster. During flight route planning, patches within the same spatial cluster are preferentially arranged on the same evidence collection flight route, thereby reducing the round-trip flight distance of the UAV between adjacent patches.
[0034] In this embodiment of the application, in order to adapt to the shooting requirements of irregularly distributed patches, the following processing is also included in step 3: For target patches with irregular geometric shapes, the boundary contour coordinate sequence of the target patch is obtained, the minimum bounding rectangle of the target patch is calculated, parallel scan lines are generated based on the major axis direction of the minimum bounding rectangle, the spacing between scan lines is determined according to the field of view of the camera on the UAV and the preset flight altitude, the intersection of the scan line and the boundary of the target patch is used as the turning waypoint of the evidence flight path, and a scanning flight path segment covering the entire patch range is generated and embedded into the evidence flight path of the shooting point corresponding to the target patch.
[0035] Step 4: Send the flight path plan to the drone nest at the target tower site through the flight control and dispatch platform, and dispatch the drone to perform fixed-point and directional shooting according to the flight path plan to collect evidence image data; Based on the flight path plans of each tower site, the flight control and dispatch platform sends takeoff commands and flight path data to the corresponding UAV nests at the tower sites according to task priority and time windows. Upon receiving the takeoff commands and flight path data, the UAV nests automatically release and take off. The UAVs fly sequentially according to the waypoint sequence in the flight path plan. After reaching each shooting location, they perform fixed-point and directional shooting according to the preset shooting parameters (including shooting altitude, lens pitch angle, and heading angle) in the flight path plan to acquire evidence images for each shooting location. After shooting is completed, the UAVs return to their nests along the evidence-gathering flight path and land automatically. The collected evidence-gathering image data is transmitted back to the flight control and dispatch platform via the communication link.
[0036] It should be noted that the shooting height and lens tilt angle in the above shooting parameters are pre-calculated based on the area of the target patch and the terrain features. For cultivated land patches, the shooting height is determined based on the patch area so that the ground coverage of a single image can include the main area of the patch; the lens tilt angle is adjusted according to the terrain slope information, using a vertical overhead angle in flat areas and an oblique shooting angle in sloping areas to obtain more complete surface feature information.
[0037] Furthermore, the specific calculation method for the aforementioned shooting height is as follows: Let the size of the camera sensor in the flight direction be... The camera focal length is The length of the target patch in the flight direction is Then the shooting height Determined by the following formula:
[0038] That is, while ensuring that the ground coverage of a single image can include the main area of the patch, the camera focal length is used as a reference. Dimensions of the sensor in the direction of flight Calculate the required flight altitude using proportional relationships ,in This represents the actual ground length of the target patch along the flight direction. (Camera pitch angle) According to the terrain slope Make adjustments in flat areas ( (Camera tilt angle) Take a vertical overhead shot (i.e.) (Camera pitch angle in sloping areas) Based on the vertical direction, it deflects towards the slope normal direction; the amount of deflection is related to the terrain slope. Positive correlation is used to make the image plane as parallel as possible to the slope, thereby obtaining more complete surface feature information.
[0039] In this embodiment, to ensure data acquisition quality under complex field conditions (such as cloud cover, obstacles, etc.), in addition to step 4, the flight control scheduling platform continuously receives real-time flight status data and real-time image preview data transmitted by the UAV during its flight path execution. When the flight control scheduling platform detects that the occlusion ratio in the real-time image preview data of the current shooting point exceeds a preset occlusion threshold, it automatically triggers a reshooting strategy, adjusting the shooting position or angle within a preset offset range of the current shooting point and re-executing the shooting. The occlusion threshold is a pre-set fixed value, representing the upper limit of the proportion of the occluded area in the total image area. When the actual occlusion ratio exceeds the threshold, the current image is determined not to meet the evidence requirements, and a reshooting strategy is triggered. Simultaneously, the flight control scheduling platform provides a remote intervention interface, allowing the operating terminal to remotely issue correction commands for the flight altitude and shooting parameters of specific shooting points along the evidence-providing flight path, forming a data acquisition mode combining automatic acquisition and remote intervention.
[0040] In this embodiment, to supplement the ground-view information of the UAV evidence images, the following processing is included in addition to step 4: Simultaneously with the UAV arriving at the shooting location and performing the shooting, the flight control and dispatch platform sends a coordinated shooting command to the high-position video probe deployed at the tower site where the image patch is located. The high-position video probe adjusts the gimbal turning angle and focal length parameters according to the spatial coordinates of the image patch, and synchronously acquires images of the same image patch from the high-position viewpoint of the tower. The acquired high-position video images and the UAV evidence images are each accompanied by a timestamp and spatial coordinate label, and are transmitted back to the flight control and dispatch platform together, forming a multi-view evidence image set for the same image patch.
[0041] Step 5: Perform intelligent analysis and spatialization processing on the evidence image data, extract land use information, and generate evidence results; The flight control and dispatch platform receives the evidence image data and performs three sub-steps in sequence: photo quality screening, land use information extraction, and image spatialization processing, to generate evidence results data that meet the evidence requirements.
[0042] Step 501: Perform quality screening on the evidence image data to obtain a set of qualified evidence images. Each of the acquired evidence images is evaluated for quality, with evaluation indicators including image sharpness score, exposure uniformity score, and target area coverage score. An image quality evaluation algorithm is used to calculate the overall quality score for each image. The input is a single evidence image, and the output is the overall quality score. The overall quality score is obtained by weighted summation of the above three scoring indicators. Specifically, let the image sharpness score be... Exposure uniformity score is The target area coverage score is The corresponding weights are respectively , , ,and Then the overall quality score for:
[0043] Among them, image sharpness score The variance of the Laplacian operator response of the image grayscale is calculated. A larger variance value indicates a clearer image. This is then normalized and mapped to the range. Range; Exposure uniformity score This is obtained by calculating the ratio of the mean to the standard deviation of the image pixel brightness values. A larger ratio indicates more uniform exposure. This is then normalized and mapped to... Interval; Target area coverage score The target region coverage score is calculated as the proportion of pixels in the target region to the total number of pixels in the image. Within the specified range, no additional normalization processing is required. Images with a comprehensive quality score below a preset quality threshold are marked as unqualified and removed, while qualified images constitute the set of qualified evidentiary images. The quality threshold is a pre-set fixed value, with a range of [range missing]. It is used to distinguish between images that meet the requirements for providing evidence and those that do not.
[0044] Step 502: Extract land use information from qualified evidence images to generate land use identification results for each patch. Input each image from the qualified evidence image set into a pre-trained land use identification model, and output the category label and corresponding confidence score of the land features in each image. The land use identification model is an image classification model based on a convolutional neural network. The input layer of the land use identification model receives the pixel matrix of a single evidence image, and extracts the local texture features and spatial structure features of the image step by step through multiple convolutional layers and pooling layers. Then, the features are mapped to the category space of the preset land use category system through a fully connected layer. The output layer of the land use identification model uses the Softmax function to convert the scores of each category into a probability distribution vector. The land use name corresponding to the category with the highest probability is used as the land use prediction result for the image. During training, evidence image samples with land use labels are used as training data. The cross-entropy loss function is used to measure the difference between the predicted probability distribution and the true label. The Adam optimization algorithm is used to update the parameters of the land use identification model. The land cover identification results of multiple images corresponding to the same target patch are comprehensively evaluated. The category label with the highest frequency and an average confidence score exceeding a preset confidence threshold is used as the land cover identification result for that target patch. The confidence threshold is a pre-set fixed value, with a range of [range missing]. It is used to screen and identify land use classification results that meet the reliability requirements.
[0045] Step 503: Based on GIS spatial analysis, perform spatialization processing on the qualified evidence images to generate map patch evidence results. Obtain the shooting location coordinates, shooting height, lens parameters, and attitude parameters of each qualified evidence image. Calculate the ground coverage area of each qualified evidence image based on its interior and exterior orientation elements. Perform spatial overlay analysis between the ground coverage area and the vector boundary of the target map patch. Calculate the ratio of the spatial intersection area of the ground coverage area and the map patch vector boundary to the total area of the map patch. When this ratio reaches a preset coverage threshold, establish a spatial association between the qualified evidence image and the target map patch. Encapsulate the land use identification results, the spatially associated qualified evidence images, and the map patch attribute information into structured map patch evidence result data. The coverage threshold is a pre-set fixed value, with a range of [value missing]. It is used to determine whether the spatial coverage of the target patch by the qualified evidence image meets the evidence requirements.
[0046] It should be noted that the ground coverage calculation in the above image spatialization processing refers to the calculation based on the three-dimensional coordinates of the shooting location. Camera focal length Image sensor size and camera tilt angle The inverse transformation of the collinearity equation is used to calculate the ground coordinates corresponding to the four corners of the image, thereby determining the projection coverage of the image on the ground. The x-coordinate of the shooting location, The vertical coordinate of the shooting location is . The elevation of the shooting location, This refers to the width dimension of the image sensor. This refers to the height dimension of the image sensor.
[0047] In this embodiment of the application, to further enrich the auxiliary decision-making information of the evidence results, the following processing is included in addition to step 5: obtaining environmental parameter data, including temperature, humidity, and soil moisture content data, collected by IoT sensors deployed at the tower site where the map patch is located within the evidence time window. The environmental parameter data is then appended as an auxiliary attribute field to the map patch evidence result data of the corresponding map patch, forming a multi-source fusion evidence result that includes image information, land cover identification results, and environmental perception data.
[0048] In this embodiment of the application, in order to incorporate the evidence results into a unified survey data management system, the following processing is also included in step 5: converting the map patch evidence result data into a format according to the data interface specification of the land survey cloud platform, and uploading the format-converted map patch evidence result data to the Zhejiang branch of the land survey cloud through the data transmission interface to complete the access and collection of survey evidence information.
[0049] This implementation method addresses the problems of low efficiency in field evidence collection, high resource input, and insufficient information technology support in land change surveys through the following technical approach: First, because the data of the map patches to be used for evidence were automatically filtered and classified based on the map patch attributes and spatial location in step 1, map patches that are not suitable for UAV operations were separated in advance, thus avoiding redundant UAV operations on invalid map patches, reducing the number of UAV flights, and thus reducing the overall resource input of the evidence collection task.
[0050] Secondly, because steps 2 and 3 involved automated task allocation and flight path planning based on the coverage area of the tower sites and the spatial distribution of the map features, utilizing existing tower base stations as UAV take-off and landing nodes, the UAVs could cover surrounding map features from the nearest tower site without human intervention. This overcomes the time consumption and transportation limitations of traditional manual evidence collection, which requires field personnel to reach each point individually, thus improving the efficiency of evidence collection and coverage. Especially for scattered and inaccessible map features in fragmented mountainous farmland areas, the wide-area distribution of tower sites combined with the aerial mobility of UAVs means that evidence collection for such map features is no longer constrained by ground transportation conditions.
[0051] Third, because a fixed-point and directional automated shooting method is adopted in step 4, and combined with the linkage acquisition of high-position video probes, evidence images with multiple perspectives and consistent parameters can be obtained for the same patch, overcoming the problem of inconsistent shooting angles and content caused by differences in experience in manual shooting, thus making the acquisition quality of evidence images tend to be stable.
[0052] Fourth, because the evidence image data underwent quality screening, land use identification, and image spatialization processing in step 5, and the land use identification results were spatially correlated with the vector boundaries of the map patches, the output map patch evidence data simultaneously contains land use information and spatial correspondence. This overcomes the problem of unclear spatial relationships between photos and map patches in traditional evidence collection, thus enabling the map patch evidence data to directly meet the verification requirements of change investigation.
[0053] In summary, this implementation method coordinates the use of tower infrastructure resources and the automated operation capabilities of drones, and introduces automated screening, planning, and intelligent analysis processes in the data processing stage, forming a complete data processing chain from map patch screening to the output of evidence results. This improves the overall efficiency and quality of land change survey evidence collection.
[0054] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
[0055] This invention provides a method for generating evidence of land use changes based on tower-mounted drones and their nests, and also includes the following implementation methods: In complex terrain conditions (such as fragmented farmland in mountainous areas), a large number of land parcels are scattered and inaccessible, leading to the following technical problems with traditional manual evidence collection methods: First, field personnel need to visit each land parcel individually, resulting in high resource investment and low efficiency; second, the shooting angle, location, and content of evidence photos depend on personnel experience, leading to inconsistent quality of results; third, the selection of land parcels, flight path planning, and image analysis lack automated information technology support, and data transfer between each stage relies on manual operation, resulting in insufficient overall process coordination, which leads to low efficiency and difficulty in guaranteeing the quality of land change survey evidence collection.
[0056] Step 100: Obtain time-series images of the target patch and extract construction progress features; Long-term satellite imagery of the target patch is acquired, and construction progress features are extracted from the visible areas in each temporal phase image to obtain a time series of construction progress features. These features include building outline area, number of structural floors, and completion rate indicators.
[0057] It should be noted that the building outline area refers to the projected area of the building on a horizontal plane extracted from satellite imagery. The structural floor count refers to the estimated number of floors corresponding to the building's height, calculated based on the building's shadow length and solar altitude angle in the imagery. The completion index refers to the percentage of construction completion comprehensively assessed based on visual characteristics such as roof closure and facade integrity.
[0058] It should be noted that the extraction of construction progress features can be achieved using the following methods: identifying and extracting building areas and boundaries from satellite imagery based on a semantic segmentation model, and calculating the building outline area; then, based on the shadow geometric relationship formula... Calculate building height ,in The length of the shadow. The solar altitude angle is used to estimate the number of structural floors by dividing the building height by the standard floor height; the completion index is output through a pre-trained construction status classification model.
[0059] Furthermore, the standard floor height ranges from 2.8 meters to 3.6 meters, with the specific value determined according to the building type. The standard floor height for residential buildings is usually 2.8 meters to 3.0 meters, while the standard floor height for public buildings is usually 3.3 meters to 3.6 meters.
[0060] Furthermore, the semantic segmentation model employs a convolutional neural network (CNN), taking multispectral image data from satellite imagery as input and outputting a pixel-level building region segmentation mask. The CNN includes an encoder and decoder structure. The encoder extracts multi-scale feature representations of the image through multiple convolutional and pooling layers, while the decoder restores the feature representations to the original image resolution through upsampling and convolutional layers, outputting the probability value of each pixel belonging to a building category. The output layer uses... The convolutional layer transforms the feature maps of the decoder into a class probability distribution for each pixel. During training, a cross-entropy loss function is used, and supervised learning is performed using satellite imagery samples labeled with building boundaries.
[0061] Furthermore, the construction status classification model employs a convolutional neural network (CNN). The input is an image patch of the building area, and the output is a classification label or regression value for the completion index. The CNN extracts texture and structural features of the building area through multiple convolutional layers, performs dimensionality reduction via a global average pooling layer, and then outputs the completion prediction result through a fully connected layer. During training, either the classification cross-entropy loss function or the mean squared error loss function is used, and supervised learning is performed using building image samples labeled with completion levels.
[0062] Illegal construction has been found on plot D327 in a development zone on the western outskirts of a city. The plot is designated for green space, but materials indicate that large-scale construction is underway. Regulatory authorities retrieved monitoring data from the TW-089 monitoring station located near the plot. This station is equipped with fixed surveillance cameras and a drone nesting system, capable of continuously monitoring land use changes within a 3-kilometer radius.
[0063] The regulatory authorities first extracted time-series imagery data of plot D327 from satellite imagery archives over the past 120 days. Because the plot is located in a suburban area with frequent satellite overpasses, 12 effective image phases were acquired, with an interval of approximately 10 days. Semantic segmentation models were used to process the images of each phase, extracting building outlines from the visible areas and calculating construction progress features.
[0064] In the first time-phase image on February 5, 20XX, there were traces of building foundation excavation covering an area of approximately 580 square meters on the south side of plot D327. By the fifth time-phase image on March 17, 20XX, the building outline area in this area had expanded to 1240 square meters, and the building height was estimated to be approximately 8.64 meters based on the length of the shadow.
[0065] The solar altitude angle in the image is Length of building shadow Meters, substitute into the formula to calculate the building height Based on a standard residential building floor height of 2.9 meters, the estimated number of structural floors is [number missing]. layer.
[0066] Step 200: Perform trend decomposition on the construction progress characteristic time series and predict the current construction progress; The construction progress characteristic time series is decomposed to separate the long-term trend component and the periodic component. A construction speed curve is fitted based on the long-term trend component, and a time series forecasting model is used to extrapolate the long-term trend component to calculate the expected construction progress value at the current moment.
[0067] It should be noted that trend decomposition can employ the seasonal trend decomposition method, which decomposes the original time series... Decomposed into trend terms Seasonal items and residuals ,Right now ,in For the first Construction progress characteristics observations at specific times. For time index, For the first Trend of time, For the first Seasonal items of the moment, For the first The residual term at time step. Long-term trend component. Characterizing the overall growth pattern of construction progress over time, with periodic components. It reflects the periodic fluctuations in construction activities due to factors such as weather and holidays.
[0068] Furthermore, the seasonal trend decomposition method extracts the trend term using the moving average method. The specific steps are as follows: apply a centered moving average window to the original time series to calculate the smoothed trend term, with the moving average window length set to the seasonal cycle length; subtract the trend term from the original series to obtain the detrended series; group the detrended series by seasonal cycle to obtain the seasonal term; and subtract the seasonal term from the detrended series to obtain the residual term.
[0069] Furthermore, the length of the seasonal cycle is determined based on the periodic characteristics of construction activities, and is usually set to 7 days or 30 days. A 7-day cycle corresponds to the regular fluctuations of weekly construction activities, while a 30-day cycle corresponds to the phased changes in monthly construction progress.
[0070] It should be noted that the construction speed curve can be fitted using polynomial fitting or spline interpolation methods to characterize the variation of construction speed with different construction stages. The time series forecasting model can use an autoregressive integral moving average model or an exponential smoothing model, with the input being the historical long-term trend component series and the output being the expected construction progress value at the current moment.
[0071] Furthermore, time series forecasting models can also employ Long Short-Term Memory (LSTM) networks. The input is a historical long-term trend component sequence, and the output is the expected construction progress value at the current moment. The LSM network consists of an input layer, an LSM layer, and a fully connected output layer. The input layer receives the long-term trend component sequence within a historical time window. The LSM layer extracts the long-term dependency features of the time series through forgetting gates, input gates, and output gates. The fully connected output layer maps the hidden states of the LSM layer to the expected construction progress value. During training, a mean squared error loss function is used, and supervised learning is performed using historical project construction progress time series data.
[0072] Furthermore, the length of the historical time window ranges from 30 days to 180 days, with the specific value determined based on the cyclical characteristics of the construction project. Shorter window lengths are used for fast-paced construction projects, while longer window lengths are used for long-cycle construction projects.
[0073] In this embodiment, to improve the stability of construction progress prediction, outlier detection and correction can be performed on the original time series before trend decomposition. Outliers may be caused by cloud cover, imaging quality issues, or construction interruption events. They are identified using a median filtering method based on a sliding window, replacing observations that deviate from the median within the window by more than a preset multiple of the standard deviation with the median within the window, thereby obtaining a smoother input sequence.
[0074] Furthermore, the preset multiple ranges from 2 to 5 times the standard deviation. The specific value is determined based on the fluctuation characteristics of the time series. For series with larger fluctuations, a larger multiple is used to avoid misjudging normal fluctuations as outliers.
[0075] Furthermore, the length of the sliding window ranges from 5 to 15 time points, with the specific value determined based on the sampling frequency of the time series. Sequences with higher sampling frequencies use a longer window length to improve the stability of outlier detection.
[0076] Technical personnel from the regulatory department conducted trend decomposition analysis on the extracted time series of structural layers. Due to the influence of weather and weekends on construction schedules, the seasonal cycle length was set to 7 days. The moving average method was applied to extract the long-term trend component from the structural layer data for time phases T01 to T10, with a window length set to 3 time phases.
[0077] After trend decomposition, the long-term trend components for each time phase are obtained. For example, the original observed number of structural layers for time phase T05 is 3 layers, and the long-term trend components are... Layers, seasonal items Layer, residual term Layer, satisfying layer.
[0078] Based on the extracted long-term trend component sequences, technicians used a Long Short-Term Memory (LSTM) network for time series forecasting. The network input was a trend sequence from time phases T01 to T10, with a historical time window length of 100 days, covering all 10 valid time phases. After training, the network extrapolated and predicted the structural layer trend as of May 15, 20XX, outputting an expected trend value of 11.8 layers, with an upper bound of 12.5 layers and a lower bound of 11.1 layers.
[0079] Meanwhile, the same processing was applied to the time series of completion indicators, predicting that the current completion trend is 78.5%, with a prediction range of 75.2% to 81.8%. The building outline area remained stable after phase T06, with the current outline area predicted to be 1240 square meters.
[0080] Step 300: Reconstruct the 3D point cloud of the current image and detect void regions; The system reconstructs 3D point clouds from the current multi-view images collected by the tower drone, detects void areas in the point cloud with significantly lower density than the surrounding areas, identifies the types of temporary facilities causing occlusion through the occlusion detection algorithm, and extracts the boundary contours of the voids.
[0081] It should be noted that the 3D point cloud reconstruction uses the motion recovery structure algorithm or the multi-view stereo matching algorithm. The input is multi-view images and their shooting pose parameters, and the output is dense point cloud data.
[0082] Furthermore, the motion reconstruction algorithm recovers the camera pose and sparse point cloud through feature point matching and bundle adjustment. The specific steps are as follows: extract feature points from multi-view images and perform cross-view matching; estimate the relative pose between cameras based on the matched point pairs through epipolar geometric constraints; optimize the camera pose parameters and 3D point coordinates using bundle adjustment; and generate a dense point cloud based on the sparse point cloud through a multi-view stereo matching algorithm.
[0083] It should be noted that the detection of void regions refers to the meshing of the reconstructed 3D point cloud, the calculation of the point cloud density value within each mesh cell, and the marking of connected regions where the point cloud density is lower than a preset density threshold and the density difference with adjacent mesh cells exceeds a preset difference threshold as void regions. The void boundary contour is obtained by edge detection and vectorization processing of the void regions.
[0084] Furthermore, the size of the grid cells ranges from 0.5 meters to 2 meters, with the specific value determined based on the spatial resolution of the point cloud reconstruction. When the point cloud density is high, a smaller grid size is used to improve the accuracy of hole detection.
[0085] Furthermore, the preset density threshold is set to 10% to 30% of the average point cloud density of the surrounding normal area, and the preset difference threshold is set to 50% to 80% of the average point cloud density of the surrounding normal area. The specific threshold is determined based on the point cloud reconstruction quality and the degree of occlusion.
[0086] Furthermore, the range of the surrounding normal area is defined as a ring-shaped area 5 to 20 meters away from the hole boundary. The point cloud density of the ring-shaped area meets the integrity requirements and is not affected by occlusion, and can represent the point cloud density level under normal reconstruction quality.
[0087] It should be noted that the occlusion detection algorithm identifies typical construction occlusion types such as scaffolding, dust nets, and temporary fences by analyzing the texture and geometric features of the corresponding locations of holes in the original image. The identification method can employ an image classification model based on a convolutional neural network, with the input being the corresponding image patch of the hole area in the original image and the output being a label for the type of occlusion facility.
[0088] Furthermore, the image classification model in the occlusion detection algorithm employs a convolutional neural network (CNN). The input is the corresponding image patch of the hole region in the original image, and the output is the classification probability distribution of the occlusion type. The CNN extracts texture and edge features of the image patch through multiple convolutional layers, performs dimensionality reduction via global average pooling layers, and then outputs the classification probability of each occlusion type through fully connected layers. During training, a cross-entropy loss function is used, and supervised learning is performed using construction site image samples labeled with occlusion types.
[0089] On the morning of May 15, 20XX, the drone nesting system at tower site TW-089 automatically activated. The drone followed a preset flight path to collect multi-view images of plot D327. The drone circled the target building at an altitude of 80 meters, taking one image every 15 degrees, for a total of 24 multi-view images with a resolution of 0.05 meters. Simultaneously, the tower's fixed monitoring camera took three supplementary images from a 45-degree downward angle.
[0090] Technicians input 27 acquired images into a 3D point cloud reconstruction system, used the structure-of-motion (SOMO) algorithm to extract feature points and perform cross-view matching, identifying a total of 128,547 matching point pairs. After optimizing the camera pose using bundle adjustment, a dense point cloud containing 1,856,320 points was generated.
[0091] The reconstructed point cloud was meshed, with a grid cell size of 1 meter. The point cloud density within each grid cell was calculated, and the average point cloud density in the surrounding normal area was 2340 points per cubic meter. However, in the central area of the building, the point cloud density of large-area grid cells was less than 700 points per cubic meter, only 30% of the normal density, and the density difference with adjacent grid cells exceeded 1400 points, reaching 60% of the normal density.
[0092] Through connectivity analysis, a cavity area of 856 square meters was identified, corresponding to the core part of the main building. The boundary outline of the cavity was extracted; the outline is an irregular rectangle with a major axis of approximately 42 meters and a minor axis of approximately 26 meters.
[0093] Occlusion detection was performed on image patches corresponding to holes in the original image. The classification results output by the convolutional neural network showed that the occlusion type was a closed dustproof net, with a classification confidence level of 0.92. Green dense mesh netting can be seen covering the entire building facade in the image, with clear mesh texture and scaffolding support structure visible at the edges.
[0094] Step 400: Extract visible state parameters around the cavity region; Analyze the current construction status of the visible parts around the cavity area, extract the floor height parameters, span parameters, and material parameters of the completed structure, and obtain the surrounding visible status parameters.
[0095] It should be noted that the floor height parameter refers to the vertical distance between floors of the surrounding visible structure, measured from the 3D point cloud. The span parameter refers to the horizontal distance between adjacent load-bearing components in the surrounding visible structure. The material parameter refers to the type of building material identified through image spectral and texture features, including categories such as concrete, steel structure, and brick masonry.
[0096] Furthermore, the floor height parameter is measured by identifying floor boundary features in the 3D point cloud. The specific steps are: slicing the point cloud vertically, statistically analyzing the point cloud density distribution for each floor height, identifying the local minimum of the point cloud density at each floor boundary, and determining the vertical distance between adjacent floor boundaries as the floor height parameter. The span parameter is measured by identifying the spatial location of load-bearing components in the 3D point cloud. Load-bearing components are identified through the geometric shape and vertical continuity features of the point cloud, and the horizontal distance between the centerlines of adjacent load-bearing components is the span parameter.
[0097] Furthermore, the thickness of the vertical slice ranges from 0.1 meters to 0.3 meters, with the specific value determined based on the point cloud density. When the point cloud density is high, a smaller slice thickness is used to improve the recognition accuracy of the floor boundary position.
[0098] Furthermore, the identification of load-bearing components is achieved through the following characteristics: the point cloud presents a columnar or wall-like geometric shape on the horizontal cross section; the point cloud has continuity in the vertical direction and spans multiple floors in height; the point cloud density is higher than that of the surrounding area; and the surface flatness meets the requirements of structural components.
[0099] In this embodiment of the application, in order to improve the accuracy of extracting visible state parameters in the surrounding area, a preset buffer distance can be extended outward from the contour of the cavity boundary. Within the buffer range, sampling points with point cloud density that meet the integrity requirements are selected for parameter measurement. The measurement results of multiple sampling points are then weighted and averaged, with the weight being negatively correlated with the distance from the sampling point to the cavity boundary.
[0100] Furthermore, the preset buffer distance ranges from 3 meters to 10 meters, with the specific value determined based on the typical span dimensions of the building structure, ensuring that the buffer zone can cover at least one complete structural unit. Meeting the point cloud integrity requirement means that the point cloud density of the grid cell where the sampling point is located is not less than 70% of the average point cloud density of the surrounding normal area.
[0101] Furthermore, the calculation method for the negative correlation between the weight and the distance from the sampling point to the cavity boundary is as follows: normalize the distance from the sampling point to the cavity boundary to the interval between 0 and 1, and the weight value is equal to 1 minus the normalized distance. The closer the sampling point is to the cavity boundary, the smaller the weight, and the farther the sampling point is from the cavity boundary, the larger the weight.
[0102] Technicians extended an 8-meter buffer zone outward from the cavity boundary, identifying visible structural portions of the building's east and west edges within the buffer zone. The point cloud density in these areas reached 1850 to 2120 points per cubic meter, meeting the integrity requirements.
[0103] In the visible area on the east side, a 0.2-meter-thick slice analysis was performed vertically, identifying 11 floor boundary locations. The vertical distances between the floors were 2.85 meters, 2.92 meters, 2.88 meters, 2.91 meters, 2.89 meters, 2.87 meters, 2.90 meters, 2.86 meters, 2.89 meters, and 2.88 meters, respectively. A weighted average of the 10 floor height measurements was calculated, with the weight determined based on the distance from the sampling point to the cavity boundary: a sampling point 5.2 meters from the boundary had a weight of 0.65, and a sampling point 7.8 meters from the boundary had a weight of 0.98. The weighted average floor height parameter was 2.88 meters.
[0104] In the visible area on the west side, four vertical load-bearing columns were identified. The column cross-section dimensions are approximately 0.6 meters × 0.6 meters, and the column surfaces have high point cloud density and good flatness. The horizontal distances between the centerlines of adjacent load-bearing columns were measured to be 6.8 meters, 6.7 meters, and 6.9 meters, respectively, with an average span parameter of 6.8 meters.
[0105] Texture analysis of visible structures on the east and west sides of the original image revealed that the surface of the load-bearing column is grayish-white, and its spectral reflectance and rough texture characteristics are consistent with those of concrete. The material identification model outputs a classification result of concrete with a confidence level of 0.88.
[0106] Step 500: Infer the current construction status of the shading area based on time series prediction and surrounding constraints; The expected construction progress value and surrounding visible state parameters are preprocessed and combined into a feature vector. This feature vector is then input into the construction progress inference model, which outputs a predicted construction state of the obstructed area. The predicted construction state includes the predicted number of floors, the predicted structure type, and the predicted degree of completion.
[0107] It should be noted that data preprocessing eliminates differences in dimensions and numerical ranges between different features. For data categorized into material parameters, one-hot encoding is used to convert them into numerical vectors.
[0108] Furthermore, the one-hot encoding method converts each material category into a binary vector with a length equal to the total number of material categories. The vector is set to 1 only at the corresponding category position and 0 at the rest. For example, when the material categories include concrete, steel structure and brick masonry, concrete is encoded as [1, 0, 0], steel structure is encoded as [0, 1, 0], and brick masonry is encoded as [0, 0, 1].
[0109] It should be noted that the construction progress inference model employs either a gradient boosting decision tree algorithm or a random forest algorithm. The input is a feature vector containing the expected construction progress value and surrounding visible state parameters, and the output is a construction status prediction vector. The input feature vector includes the expected building outline area, the expected number of structural floors, the expected completion rate index, floor height parameters, span parameters, and material parameters.
[0110] Furthermore, the construction progress inference model can also employ a multilayer perceptron. The input is a feature vector containing the expected construction progress value and surrounding visible state parameters, and the output is a construction status prediction vector. The multilayer perceptron consists of an input layer, multiple hidden layers, and an output layer. The input layer receives the feature vector, the hidden layers extract a nonlinear combination representation of the features through fully connected layers and activation functions, and the output layer outputs the prediction layer number, predicted structure type, and predicted degree of completion through fully connected layers. During training, a mean squared error loss function is used, and supervised learning is performed using paired samples of historical construction project progress features and actual construction status.
[0111] Furthermore, the number of hidden layers in a multilayer perceptron ranges from 2 to 5, and the number of neurons in each hidden layer ranges from 64 to 512, with the specific configuration determined based on the scale of the training data and the complexity requirements of the model. The activation function used is the ReLU function or a variant thereof, to introduce nonlinear transformation capabilities.
[0112] In this embodiment of the application, to improve the spatial consistency between the predicted construction status and the surrounding area, the construction progress inference model may also introduce a consistency constraint term. The consistency constraint term requires that the difference between the predicted number of floors and the number of floors of the surrounding visible structures does not exceed a preset floor threshold, and that the predicted structure type matches the structure type indicated by the surrounding material parameters, thereby avoiding significant contradictions between the prediction results and the actual surrounding conditions.
[0113] Furthermore, the preset number of floors threshold ranges from 1 to 3 floors. The specific value is determined based on the differences in construction progress at the construction site. The differences in construction progress between different areas of the same construction site usually do not exceed the preset number of floors threshold range.
[0114] Furthermore, the consistency constraint term is added to the loss function as a regularization term during model training. The weight coefficient of the constraint term ranges from 0.1 to 0.5. The larger the weight coefficient, the more the construction progress inference model tends to output a prediction result consistent with the surrounding state.
[0115] Technicians preprocessed the expected construction progress value obtained in step 2 and the surrounding visible status parameters extracted in step 4. The expected building outline area is 1240 square meters, the expected number of structural floors is 11.8, the expected completion rate is 78.5%, the floor height is 2.88 meters, the span is 6.8 meters, and the material is concrete.
[0116] The material parameters are uniquely encoded, with a total of 3 material categories (concrete, steel structure, and brick masonry). Concrete is encoded as [1, 0, 0]. All features are combined into a feature vector: [1240, 11.8, 78.5, 2.88, 6.8, 1, 0, 0].
[0117] The feature vectors are input into the construction progress inference model, which employs a multilayer perceptron structure with three hidden layers containing 128, 256, and 128 neurons respectively, and ReLU activation function. A consistency constraint term is introduced during model training, with a constraint weight coefficient set to 0.3, and a preset layer threshold of 2 layers.
[0118] The model outputs the following construction status prediction results: the predicted number of floors is 11, the predicted structure type is a concrete frame structure, and the predicted completion rate is 76%. The predicted number of floors (11 floors) is completely consistent with the 11 floors of the surrounding visible structures, satisfying the consistency constraint. The predicted completion rate of 76% is slightly lower than the 78.5% predicted in the time series, reflecting the calibration effect of the surrounding measured data on the prediction results.
[0119] Step 600: Infer the elevation distribution of the cavity area; Historical topographic data of the cavity area is obtained, and combined with construction status prediction and effective point cloud around the cavity boundary, the current elevation distribution of the cavity area is inferred using the Kriging interpolation algorithm or the radial basis function interpolation algorithm to obtain the filling point cloud data of the cavity area.
[0120] It should be noted that historical topographic data refers to digital elevation model data of the cavity area before construction, providing a baseline ground elevation for the cavity area. The inference of the current elevation distribution is based on the following calculation: the baseline elevation provided by the historical topographic data is added to the incremental construction height obtained by multiplying the predicted number of floors in the construction status prediction by the floor height parameter, to obtain the predicted elevation value for each location in the cavity area.
[0121] Furthermore, the formula for calculating the predicted elevation value is as follows: ,in To predict elevation values, As the reference ground elevation, To predict the number of layers, This refers to the floor height parameter.
[0122] It should be noted that the Kriging interpolation algorithm or radial basis function interpolation algorithm uses the elevation values of the effective point cloud around the cavity boundary as known control points and the predicted elevation values as constraints. The input is the coordinates and elevation values of the control points, and the output is continuous elevation distribution data of the cavity area. The elevation distribution data is then converted into fill point cloud data.
[0123] Furthermore, the Kriging interpolation algorithm establishes an interpolation model through spatial autocorrelation. The specific steps are as follows: calculate the semivariance function between control points, fit the theoretical model of the semivariance function, calculate the spatial weight between the interpolation point and each control point based on the theoretical model, and obtain the elevation estimate of the interpolation point by weighted summation of the elevation values of the control points. When calculating the weights, the predicted elevation value is introduced as an additional constraint condition so that the interpolation result satisfies both spatial smoothness and construction status prediction.
[0124] Furthermore, the theoretical model for the semivariogram function can be a spherical model, an exponential model, or a Gaussian model. The model parameters include nugget value, sill value, and range, which are determined by fitting the experimental semivariogram function. The nugget value reflects measurement error and microvariability, the sill value reflects the total magnitude of spatial variation, and the range reflects the effective distance of spatial autocorrelation.
[0125] Furthermore, the control points are selected within a ring-shaped area 1 to 15 meters from the cavity boundary. The point cloud density of this ring-shaped area meets the integrity requirements and can represent the elevation change characteristics near the cavity boundary. The number of control points ranges from 20 to 200, with the specific number determined based on the area of the cavity and the complexity of its boundary.
[0126] In this embodiment of the application, in order to make the filling point cloud data and the surrounding valid point cloud transition smoothly at the boundary, a transition zone can be set at the boundary of the hole. Within the transition zone, the interpolated elevation and the surrounding measured elevation are mixed by distance weighting. The closer to the center of the hole, the greater the weight of the interpolated elevation. The closer to the boundary of the hole, the greater the weight of the surrounding measured elevation.
[0127] Furthermore, the width of the transition zone ranges from 1 meter to 5 meters, with the specific value determined based on the size of the cavity area. The width of the transition zone is typically 5% to 15% of the minimum side length of the cavity area. The distance-weighted hybrid method uses a linear interpolation method. The weight of the interpolated elevation decreases linearly with the distance from the cavity boundary, while the weight of the measured elevation around the cavity increases linearly with the distance from the cavity boundary. The sum of the two weights is 1.
[0128] Furthermore, the formula for calculating the distance-weighted mixture is as follows: ,in The elevation value after mixing. For interpolating elevation, The measured elevation of the surrounding area For interpolation elevation weights, The measured elevation of the surrounding area is used as the weight, and Weight The calculation formula is ,in Let the distance be the point from the boundary of the cavity. This refers to the width of the transition zone.
[0129] Technicians retrieved the pre-construction digital elevation model of plot D327 from the topographic database. This model, acquired in January 20XX, shows the baseline ground elevation at the corresponding location of the cavity area. It is 312.5 meters (relative to the city's elevation datum).
[0130] Based on the construction status prediction, the predicted number of floors Layer, layer height parameter Meters, the calculated increase in construction height is:
[0131] Calculate the predicted elevation of the center of the cavity area:
[0132] Within a ring-shaped area ranging from 1 to 12 meters outside the cavity boundary, 85 control points with sufficient point cloud density were selected, and their three-dimensional coordinates and elevation values were extracted. The elevation range of the control points is 313.2 meters to 343.6 meters, reflecting the height variation of the building from the ground to the top floor.
[0133] A Kriging interpolation algorithm was used to establish an elevation interpolation model. A spherical model was adopted for the semivariance function. The model parameters were determined by fitting the semivariance of the experiment: the nugget value was 0.8 square meters, the abutment value was 12.5 square meters, and the range was 18 meters. In the interpolation calculation, the predicted elevation value of 344.18 meters was used as an additional constraint condition for the cavity center to ensure that the interpolation results were consistent with the prediction of the construction status.
[0134] Interpolation calculations were performed on the cavity area using a 0.5m × 0.5m grid to generate continuous elevation distribution data. A 3m wide transition zone was set at the cavity boundary, and a distance-weighted hybrid method was used within the transition zone.
[0135] For example, for a point 1.2 meters from the boundary, the interpolation elevation weight... Weight of measured elevation of surrounding area If the interpolated elevation of this point is 338.5 meters, and the measured elevation of the surrounding area is 337.8 meters, then the combined elevation is:
[0136]
[0137] The elevation distribution data was converted into an infill point cloud. Sampling was performed according to the average point cloud density of 2340 points per cubic meter in the surrounding normal area to generate infill point cloud data containing 1,892,650 points.
[0138] Step 700: Fill the void areas with texture synthesis; Reference images matching the construction status prediction are retrieved from the construction progress image database. An image inpainting network is used to synthesize and fill the void areas with textures that match the predicted status, thereby obtaining the filling texture data for the void areas.
[0139] It should be noted that the construction progress image database stores typical construction site images of different construction stages and structural types, along with their corresponding construction status annotations. Reference image retrieval is based on the predicted number of layers, predicted structural type, and predicted degree of completion in the construction status prediction. The similarity between the construction status annotations of each image in the construction progress image database and the predicted construction status is calculated, and the image with the highest similarity is selected as the reference image.
[0140] Furthermore, similarity calculation is achieved through a weighted distance metric. The specific calculation method is as follows: the construction status label and the construction status prediction are represented as multi-dimensional vectors, and each dimension of the vector corresponds to the predicted number of layers, the predicted structure type, and the predicted completion degree, respectively. The Euclidean distance between the two vectors is calculated, and the distance components of each dimension are summed after being assigned different weights. The weights are determined according to the degree of influence of each feature on the texture appearance, with the predicted completion degree having the largest weight, followed by the predicted number of layers, and the predicted structure type having the smallest weight.
[0141] Furthermore, the formula for calculating the weighted distance is:
[0142] in For weighted distance, To predict the number of layers, The labeled layer number for the reference image. To predict the degree of difference between the structure type and the reference image structure type, To predict the degree of completion of the project, To indicate the completion level of the annotation for the reference image, , , These are weighting coefficients for the number of floors, structural type, and completion level, respectively. The range of values for these weighting coefficients is as follows: It ranges from 0.2 to 0.3. It ranges from 0.1 to 0.2. It is between 0.5 and 0.7, and .
[0143] Furthermore, structural type differences The calculation method is as follows: when the predicted structure type is exactly the same as the reference image structure type... When the two belong to different categories .
[0144] It should be noted that the image inpainting network employs an encoder-decoder structure composed of convolutional neural networks. The input consists of a mask image of the hole region, a texture image of the region surrounding the hole, and a reference image. The output is a synthesized texture image of the hole region. The encoder extracts multi-scale feature representations of the texture image of the region surrounding the hole and the reference image through multiple convolutional and pooling layers. The decoder restores the fused feature representations to the original image resolution through upsampling and convolutional layers, generating the synthesized texture of the hole region. The decoder's output layer uses... The convolutional layer converts the feature maps into RGB three-channel texture images. During training, a weighted combination of reconstruction loss and perceptual loss is used as the loss function, and supervised learning is performed using training samples containing the complete image and the corresponding mask.
[0145] Furthermore, the reconstruction loss employs a mean squared error loss function to calculate the pixel-level difference between the generated texture image and the real texture image. The perceptual loss extracts high-level feature representations of the generated and real texture images through a pre-trained image classification network, calculating the mean squared error between these feature representations to constrain the semantic consistency between the generated texture and the real texture. The overall form of the loss function is a weighted sum of the reconstruction loss and the perceptual loss, with the weight coefficients adjusted based on the convergence performance during training.
[0146] Furthermore, the formula for calculating the loss function is as follows: ,in For the total loss, To rebuild the losses, In order to perceive loss, and This refers to the weighting coefficient. The range of values for the weighting coefficient is: It ranges from 0.3 to 0.5. The value ranges from 0.5 to 0.7, with the specific value determined based on the balance between the visual quality and semantic consistency of the generated textures during the training process.
[0147] In this embodiment, the image inpainting network can employ a generative adversarial network architecture that includes a contextual attention module. The contextual attention module calculates the similarity weights between each location in the hole region and each location in the surrounding valid region using cosine similarity. Based on these similarity weights, it performs weighted aggregation of features in the surrounding region, thereby establishing a semantic correspondence between the hole region and the surrounding region and improving the visual coherence between the synthesized texture and the surrounding region.
[0148] Furthermore, when employing a generative adversarial network (GAN) architecture, the image inpainting network includes a generator and a discriminator. The generator uses the encoder-decoder structure described above, with a context attention module embedded in the decoder. The context attention module calculates the cosine similarity between the feature vectors at each location in the hole region and the feature vectors at each location in the surrounding valid region, obtaining a similarity weight matrix. This weighted aggregation of the surrounding region features is then passed to subsequent layers of the decoder. The discriminator uses a convolutional neural network, taking either a real image or a synthetic image output by the generator as input, and outputting the probability of image authenticity. The discriminator extracts image features through multiple convolutional layers, reduces dimensionality through global average pooling layers, and outputs the probability value through fully connected layers. During training, a weighted combination of adversarial loss, reconstruction loss, and perceptual loss is used as the total loss function. The generator's optimization objective is to maximize the probability that the discriminator will classify the generated image as a real image, while the discriminator's optimization objective is to maximize its ability to distinguish between real and generated images. Model optimization is achieved by alternately training the generator and discriminator.
[0149] Furthermore, the adversarial loss employs a binary cross-entropy loss function to calculate the difference between the discriminant's output probability and the true label.
[0150] The formula for calculating the total loss function is as follows: ,in For the total loss, To combat the losses, To rebuild the losses, In order to perceive loss, , , This refers to the weighting coefficient. The range of values for the weighting coefficient is: The value ranges from 0.1 to 0.3. It ranges from 0.3 to 0.4. The value ranges from 0.3 to 0.5, with the specific value determined based on training stability and generation quality.
[0151] Technicians retrieved reference images from the construction progress image database, which contains image samples of 526 construction projects in the city over the past three years. Each sample is labeled with information such as construction stage, structural type, and completion status.
[0152] Based on the construction status prediction results (predicted number of floors: 11, predicted structure type: concrete frame, predicted completion rate: 76%), the weighted distance of each candidate image is calculated. The weighting coefficients are set as follows: , , .
[0153] Image sample PROJ-20XX-0347 in the database has 10 labeled layers, a concrete frame structure, and a completion rate of 74%. The structure type of this sample is exactly the same as the prediction; therefore... Calculate the weighted distance:
[0154]
[0155] After traversing the database, the image with the smallest weighted distance, number PROJ-20XX-0347, was selected as the reference image. This image has a resolution of 0.05 meters, is taken from a 45-degree overhead angle, and shows that the main building's facade construction is complete, the window frames are installed, and waterproofing is being laid in some areas of the roof.
[0156] The network is input to inpaint the hole area using a mask image, a texture image of the surrounding visible area, and a reference image. This network employs a generative adversarial network architecture, with the generator containing an encoder-decoder structure and a context attention module, and the discriminator being a 5-layer convolutional neural network. The weight coefficients of the loss function are set as follows: , , .
[0157] The generator encoder extracts multi-scale features from the surrounding texture image and reference image. The context attention module calculates the similarity weights between the hole region and the surrounding region, and performs weighted aggregation of the surrounding features. The decoder generates a synthetic texture of the hole region, outputting an RGB image with a resolution of 0.05 meters and a size of 840 pixels × 520 pixels, corresponding to an actual hole area of 856 square meters.
[0158] The composite texture shows that the building's exterior is a gray-white concrete wall, the window frames are made of dark gray aluminum alloy, some windows on some floors have been fitted with glass, and construction materials and traces of worker activity are visible in the top floor area. The visual effect is consistent with the surrounding visible area and reference images.
[0159] Step 800: Generate an evidentiary image product with inference markers; The filled point cloud data is merged with the original valid point cloud, and the filled texture data is fused with the original texture data to obtain a complete 3D point cloud and texture model. The complete 3D point cloud and texture model are then converted into an orthophoto with geographic coordinates and a 3D model. An inference marker layer is added to the filled area, with the data source being time-series trend inference and prediction. Confidence markers are overlaid to generate a complete evidentiary image product containing metadata.
[0160] It should be noted that the generation of orthophotos refers to converting three-dimensional point clouds and texture models into two-dimensional images with a vertical downward perspective according to unified map projection parameters, eliminating geometric distortions caused by terrain undulations and shooting tilt angles, and registering them to a specified geographic coordinate system so that the coordinates of each point on the image are consistent with the coordinates on the ground.
[0161] Furthermore, the generation of orthophotos is achieved through the following steps: constructing a digital surface model based on the 3D point cloud and texture model, setting a unified vertical projection direction and map projection parameters, performing projection transformation on each texture patch in the 3D point cloud and texture model, calculating the pixel position of each texture patch on the projection plane, filling the texture information of the texture patch to the corresponding pixel position through texture mapping, selecting the texture patch closest to the projection plane when multiple texture patches are projected to the same pixel position, and finally generating a continuous orthophoto.
[0162] Furthermore, map projection parameters include projection type, central meridian, standard parallel of latitude, and coordinates of the projection origin. The projection type can be Gauss-Kruger projection or universal transverse Mercator projection. Specific parameters are determined based on the geographical location of the target patch and the national coordinate system standard.
[0163] It should be noted that the inference marker layer is stored in vector data format, recording the boundary range of the hole region, the identifier of the filling method, and the data source description text. The confidence level label is a numerical field, and its calculation formula is as follows:
[0164] in, This is the overall confidence level value. The confidence score for time series prediction is obtained by normalizing the inverse of the prediction interval width output by the time series prediction model. The spatial consistency confidence level is calculated based on the degree of matching between the construction status prediction and the surrounding visible status parameters. and These are preset weighting coefficients.
[0165] Furthermore, preset weighting coefficients and The values of all values are between 0 and 1, and satisfy the following conditions: When the time series data has a long time span and a high sampling frequency, The value is relatively large, typically between 0.6 and 0.8; when the integrity of the visible portion surrounding the cavity is high and the measurement accuracy is high, The value is relatively large, typically between 0.4 and 0.6. Overall confidence level value. The value ranges from 0 to 1, and the closer the value is to 1, the higher the reliability of the inference result.
[0166] Furthermore, time series prediction confidence The calculation method is as follows: the time series forecasting model outputs the upper and lower bounds of the forecast interval along with the predicted value. The width of the forecast interval is defined as the difference between the upper and lower bounds. The reciprocal of the forecast interval width is taken and mapped to the 0-1 interval through a linear transformation. The narrower the forecast interval, the higher the confidence level of the time series prediction. Spatial consistency confidence level. The calculation method is as follows: calculate the absolute value of the difference between the predicted number of layers and the average number of layers of the surrounding visible structures, calculate the matching degree between the predicted structure type and the structure type indicated by the surrounding material parameters, and convert the difference in the number of layers and the matching degree of the type into a confidence value in the range of 0 to 1 through a weighted combination. The smaller the difference in the number of layers and the higher the matching degree of the type, the higher the confidence of spatial consistency.
[0167] Furthermore, time series prediction confidence The calculation formula is ,in To predict the interval width, This is a scaling factor, ranging from 0.1 to 1.0. The specific value is determined based on the numerical range of the prediction interval width. The values of are distributed in the range of 0 to 1.
[0168] Furthermore, spatial consistency confidence. The calculation formula is: ,in To predict the number of layers, This represents the average number of stories in the surrounding visible structures. The preset maximum layer difference threshold, For structure type matching degree, and These are the weighting coefficients, and The range of values for the weighting coefficients is as follows: It ranges from 0.4 to 0.6. The range is 0.4 to 0.6. (Structure type matching degree) The value is: when the predicted structure type perfectly matches the structure type indicated by the surrounding material parameters. When the two do not match .
[0169] In this embodiment of the application, to improve the standardization and transparency of the evidentiary materials, a confidence assessment report can also be generated. The confidence assessment report is presented in a structured document format and includes the following: the original data collection time and spatial scope, the location coordinates and area statistics of the void area, the time span and sampling frequency of the time series data, the type and prediction interval of the time series prediction model, the various values of the construction status prediction, a description of the filling method, and a statement of applicable limitations. The confidence assessment report is output in conjunction with the evidentiary image product for the reviewers of the evidence to refer to and assess its validity.
[0170] Furthermore, the confidence assessment report is stored in PDF or XML format, containing structured metadata fields and readable text descriptions. The time series data in the confidence assessment report is recorded in days for the time span, the sampling frequency in days for the time phase interval, and the prediction interval in the form of upper and lower bound value pairs. The predicted values for the construction status include specific values for the predicted number of floors, predicted structure type, and predicted degree of completion.
[0171] Technicians merged the incomplete point cloud data (1,892,650 points) with the original valid point cloud data (1,856,320 points) to generate a complete point cloud model containing 3,748,970 points. The incomplete texture data (840×520 pixels) was then fused with the original texture data to obtain a complete texture map.
[0172] The complete 3D point cloud and texture model were registered according to the CGCS2000 urban coordinate system, using the Gauss-Kruger projection with the central meridian at 117 degrees east longitude and the projection origin at 39 degrees north latitude. A vertical projection transformation was performed on the 3D model to generate an orthophoto with a ground resolution of 0.1 meters, a size of 1850×1320 pixels, and a coverage area of 185 meters × 132 meters.
[0173] Create an inferred marker layer to store the boundaries of the hole regions in GeoJSON vector format. The boundary coordinates adopt the CGCS2000 coordinate system, and the filling method identifier is recorded as "temporal trend extrapolation + kriging interpolation + texture synthesis". The data source description text is "inferred based on the temporal analysis of satellite imagery from February to May 20XX and surrounding measured data".
[0174] Calculate the confidence level label. The upper bound of the prediction interval output by the time series prediction model is 12.5 layers, the lower bound is 11.1 layers, and the prediction interval width is... Layers. Set the scaling factor. Calculate the confidence level of time series prediction:
[0175] Predicted number of layers The average number of floors in the surrounding visible structure. Layers, absolute value of the difference in the number of layers Layers. Preset maximum layer difference threshold. The predicted structure type is a concrete frame, the surrounding material parameters are concrete, and the structure type is a perfect match. Set weighting coefficients , Calculate the spatial consistency confidence score:
[0176] Set weighting coefficients , Calculate the overall confidence level:
[0177] The overall confidence level of 0.732 is marked on the inference label layer, indicating that the inference result has a high degree of reliability.
[0178] A confidence assessment report was generated and stored in PDF format. The report includes the following: The original data collection period was from February 5th to May 15th, 20XX, the spatial scope was plot D327 and a 500-meter buffer zone, the center coordinates of the void area were 117.3426 degrees east longitude and 39.1853 degrees north latitude, and the area was 856 square meters. The time series data spanned 100 days, the sampling frequency was 10 days, the time series prediction model was a Long Short-Term Memory network, and the prediction interval was from layer 11.1 to layer 12.5. The predicted number of floors under construction was 11, the structural type was a concrete frame, and the completion rate was 76%. The filling method was described as "temporal trend extrapolation combined with Kriging interpolation and texture synthesis," and the applicable conditions were limited to scenarios where the visible parts of the surrounding area of the occluded area were complete and the time series data was continuous. The prediction results represent reasonable inferences based on historical trends and surrounding conditions and are not equivalent to actual measurement data.
[0179] The regulatory authorities submitted the generated evidence-based imagery to the enforcement review department. Reviewers examined the orthophotos and found that the main outline of building D327 on plot D was clear, and its height matched that of surrounding reference buildings. After reviewing the inference marker layer, they confirmed that the central area of 856 square meters was inferred data, with a comprehensive confidence level of 0.732. Reading the confidence assessment report, they understood that the inference was based on a 100-day time series trend analysis and measured structural parameters of the surrounding area. They determined that the evidence material had reasonable evidentiary value and could serve as a reference for assessing the progress of illegal construction. This, combined with subsequent on-site verification results, led to the final enforcement conclusion.
[0180] This invention coordinates the use of tower infrastructure resources and the automated operation capabilities of drones, and introduces automated screening, planning, and intelligent analysis processes in the data processing stage, forming a complete data processing link from map patch screening to evidence output. This solves the technical problems of low efficiency, high resource investment, and insufficient information technology support in land change surveys, achieving the following technical effects: First, automated screening and classification avoids redundant drone operations on invalid map patches, reducing resource investment in evidence collection tasks; second, using tower sites as drone take-off and landing nodes enables automated evidence collection under unattended conditions, overcoming the time consumption and traffic limitations of traditional manual evidence collection where field personnel arrive at each point, thus improving evidence coverage efficiency; third, obtaining consistent evidence images through fixed-point, directional automated shooting ensures stable image acquisition quality; and fourth, outputting evidence data containing land type information and spatial correspondences through quality screening, land type identification, and spatial image processing meets the verification requirements of land change surveys.
Claims
1. A method for generating evidence of land use changes based on tower-mounted drones and their nests, characterized in that, Includes the following steps: Obtain the data of the map features to be used as evidence, filter and classify them based on the map feature attributes and spatial location, and generate a set of target map features for drone evidence collection; Obtain the spatial coordinates and coverage parameters of each tower site within the target area, calculate the effective coverage area of each tower site, spatially match the spatial coordinates of each patch in the UAV evidence target patch set with the effective coverage area of each tower site, assign each target patch to the tower site covering the patch location, and generate the evidence task set corresponding to each site. For each tower site's evidence collection task set, obtain the spatial coordinates and geometric shape data of each target patch within the evidence collection task set. Using the tower site location as the starting and ending point, under the constraint of the UAV's maximum flight range, sort and group the flight routes of each shooting point to generate evidence collection flight routes covering all shooting points. All evidence collection flight routes constitute the flight route scheme for that tower site. The flight control and dispatch platform sends the flight path plan to the drone nest at the target tower site, and the drone is dispatched to perform fixed-point and directional shooting according to the flight path plan to collect evidence image data. The evidence image data is sequentially subjected to photo quality screening, land use information extraction, and image spatialization processing to generate evidence data for map features.
2. The method according to claim 1, characterized in that, The filtering and classification based on map patch attributes and spatial location includes: Based on the exclusion rules for monitoring types, facility agricultural plots and greenhouse plots that require indoor confirmation are excluded from the scope of drone evidence. Based on the exclusion rules of spatial location, construction land parcels located in densely built-up areas and land parcels located within military land management areas are excluded from the scope of drone evidence. Based on the filtering rules for patch area, patches with an area lower than the area threshold are marked as patches that require manual confirmation.
3. The method according to claim 1, characterized in that, The effective coverage area of the tower site is defined as a circular area with the spatial coordinates of the tower site as the center and the effective coverage radius as the radius. The effective coverage radius is determined by a combination of the height of the tower site, the maximum range of the deployed UAV nest, and terrain obstruction factors. When the spatial coordinates of a target patch are simultaneously located within the effective coverage area of multiple tower sites, the patch is sorted according to the Euclidean distance from the center point of the patch to each candidate tower site, and the patch is assigned to the nearest tower site.
4. The method according to claim 1, characterized in that, The process of sorting and grouping flight paths for each shooting location under the constraint of the drone's maximum flight range includes: The total flight distance of a single evidence-providing route is constrained by the sum of the inter-segment distances between each shooting point, the flight distance of the drone from the tower station to the first shooting point, and the flight distance of the drone from the last shooting point back to the tower station, which shall not exceed the maximum endurance of the drone in a single flight. The order of visiting each shooting point is obtained through a traveling salesman problem algorithm. The input is the set of spatial coordinates of each shooting point, and the output is the order of visiting points that minimizes the total flight distance. When a single evidence-providing route cannot cover all shooting locations within the range constraint, the remaining shooting locations are allocated to the next evidence-providing route according to a greedy strategy until all shooting locations are covered.
5. The method according to claim 1, characterized in that, The methods for determining the shooting locations include: For patches whose area is smaller than the observation area threshold, the geometric center of the target patch is used as the single shooting point; For patches whose area is greater than or equal to the observation area threshold, multiple shooting points are set at equal intervals along the long axis of the target patch so that the images captured by each shooting point can cover the entire patch area.
6. The method according to claim 1, characterized in that, The scheduling of drones to perform fixed-point and directional shooting according to the route plan includes: The drone flies sequentially according to the waypoint sequence in the flight path plan. After arriving at each shooting point, it performs fixed-point and directional shooting according to the shooting parameters preset in the flight path plan. The shooting parameters include shooting altitude, lens pitch angle and heading angle. The shooting height is calculated based on the ratio between the camera focal length and the sensor size in the flight direction, as well as the actual ground length of the target patch in the flight direction. The lens tilt angle is adjusted according to the terrain slope. In flat areas, a vertical overhead angle is used, while in sloping areas, the lens deflects towards the slope normal. The amount of deflection is positively correlated with the terrain slope.
7. The method according to claim 1, characterized in that, The photo quality screening includes: Each of the collected evidence images was evaluated for quality. The quality evaluation indicators included image sharpness score, exposure uniformity score, and target area coverage score. The overall quality score is obtained by weighted summation of three scoring indicators. Images with an overall quality score below the preset quality threshold are marked as unqualified and removed, while qualified images are retained to form a set of qualified evidence images. Among them, the image sharpness score is obtained by calculating the variance of the Laplacian operator response of the image grayscale image, the exposure uniformity score is obtained by calculating the ratio of the mean to the standard deviation of the image pixel brightness values, and the target area coverage score is the proportion of the number of pixels in the target patch area to the total number of pixels in the image.
8. The method according to claim 1, characterized in that, The extraction of land category information includes: Each image in the qualified evidence image set is input into the pre-trained land use identification model, and the class label of the land features in each image and the corresponding confidence score are output. The land use identification model is an image classification model based on a convolutional neural network. The input layer receives the pixel matrix of a single evidence image, extracts the local texture features and spatial structure features of the image through multiple convolutional and pooling layers, and then maps the features to the category space of the preset land use category system through a fully connected layer. The output layer uses the Softmax function to convert the scores of each category into probability distribution vectors. The land cover identification results of multiple images corresponding to the same target patch are comprehensively judged, and the category label with the highest frequency and average confidence exceeding the preset confidence threshold is taken as the land cover identification result of the target patch.
9. The method according to claim 1, characterized in that, The image spatialization processing includes: Obtain the shooting location coordinates, shooting height, lens parameters and attitude parameters of each qualified evidence image, and calculate the ground coverage of each qualified evidence image based on the interior and exterior orientation elements of the image. Spatial overlay analysis is performed between the ground coverage area and the vector boundary of the target patch to calculate the proportion of the spatial intersection area between the ground coverage area and the vector boundary of the patch to the total area of the patch. When the ratio reaches the preset coverage threshold, a spatial association is established between the qualified evidence image and the target patch, and the land use identification result, the spatially associated qualified evidence image and the patch attribute information are encapsulated into structured patch evidence result data.
10. A system for generating evidence of land use changes based on a tower-mounted UAV and its nest, used to execute the method described in any one of claims 1 to 9, characterized in that, include: The image patch filtering module is used to acquire image patch data to be used as evidence, filter and classify them based on image patch attributes and spatial location, and generate a set of target image patches for drone evidence. The task allocation module is used to obtain the spatial coordinates and coverage parameters of each tower site, calculate the effective coverage area of each tower site, allocate the target patch to the corresponding tower site, and generate the evidence task set corresponding to each site. The flight path planning module is used to obtain the spatial coordinates and geometric shape data of each target patch within the evidence task set, sort and group the flight paths of each shooting point under the constraint of the maximum flight range of the UAV, and generate flight path schemes. The flight control and scheduling module is used to send flight path plans to the drone nests at the target tower site, and to schedule the drones to perform fixed-point and directional shooting according to the flight path plan to collect evidence image data. The intelligent analysis module is used to sequentially perform photo quality screening, land use information extraction, and image spatialization processing on the evidence image data to generate map patch evidence results data.