Low-altitude unmanned aerial vehicle video frame extraction and orthophoto change detection method based on deep learning
By using a deep learning-based method for low-altitude UAV video frame extraction and orthophoto change detection, the uncertainty problem in change detection in existing technologies is solved, and accurate localization and clear representation of changed areas are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XINGTAI JIRUI ENGINEERING TECHNOLOGY CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-06-05
AI Technical Summary
Existing technologies struggle to distinguish between real changes and imaging differences in low-altitude UAV video frame extraction and orthophoto change detection, especially when the surface type is complex or the changes are evolving in stages, leading to blurred change boundaries or inaccurate spatial positioning.
By using a deep learning-based method, a sequence of low-altitude UAV video frames is acquired, pixel categories are identified and grids are divided, surface category distribution features are extracted, spatial mapping relationships are constructed, and pixel-level transition processing is performed in conjunction with orthophotos to form a continuous coverage sequence, thereby achieving accurate positioning of changing areas.
It improves the temporal directionality and spatial determinism of change detection, and enhances the stable expression of change information and the clarity of result interpretation.
Smart Images

Figure CN122157039A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image change detection technology, and in particular to a method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning. Background Technology
[0002] The field of image change detection technology includes technologies that compare and analyze image data acquired at different times to identify changes in the state of spatial targets or ground features. This field is based on multi-source image acquisition and standardized processing, and revolves around core content such as image registration, geometric correction, temporal consistency processing, and pixel-level or object-level difference analysis. It typically involves the preprocessing and comparative analysis of remote sensing images, aerial images, or UAV images, and is used to support the identification and labeling of changed areas in application scenarios such as monitoring, inspection, and assessment. Its technical system emphasizes the consistency of images in terms of spatial location scale and expression form to ensure comparability between images from different times.
[0003] Among them, low-altitude UAV video frame extraction and orthophoto change detection refers to the technical aspects of using video data acquired by low-altitude UAVs as input, extracting static image frames from the video at set time intervals, performing geometric correction on the extracted image frames based on a reference image to form an orthophoto, and comparing and analyzing the orthophoto with the reference orthophoto at the corresponding time to identify the difference region. It covers the temporal frame extraction processing of UAV video data, spatial alignment and coordinate correction processing between the extracted image and the reference image, and the process of judging the change region based on the difference of image features. The change detection process uses the corrected static orthophoto as the analysis object, and determines the spatial distribution of the change region by performing difference analysis on two or more images at different time phases.
[0004] Existing technologies use extracted static orthophotos as the main analysis object. Change identification relies on overall differences in the image or local pixel anomalies. In cases where the surface type is complex or changes are in stages, it is difficult to distinguish between real changes and imaging differences. At the same time, the alignment process of multi-temporal images is highly dependent on data quality. Once there is a viewpoint shift or scale inconsistency, it is easy to cause blurred change boundaries or inaccurate spatial positioning, which affects subsequent inspection judgment and application of results. Summary of the Invention
[0005] To address the technical problems existing in the prior art, embodiments of the present invention provide a method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning. The technical solution is as follows: A deep learning-based method for low-altitude UAV video frame extraction and orthophoto change detection includes the following steps: S1: Acquire video frame sequences of low-altitude drone patrols, identify pixels and assign them to corresponding artificial ground, water bodies, vegetation or bare ground, cluster similar pixels and fix grid division, extract the spatial distribution features of grid surface category labels, and form a set of image frame surface category distribution descriptions. S2: Based on the land surface category distribution results of each image frame in the image frame land surface category distribution description set, select adjacent image frames, compare them in the same grid area, identify the location of land surface category change, and extract the image frame number and time, mark them as change-related image frames, and form a video change-related frame sequence. S3: For the change-related image frames in the video change-related frame sequence, obtain orthophotos, extract and match edge features, construct spatial mapping relationships, project the fixed grid of the image frames onto the geographic coordinate system of the orthophotos, and form a description set of spatial correspondence between image frames and orthophotos; S4: Based on the spatial correspondence between the image frames and the orthophoto spatial correspondence, the spatial position correspondence of the fixed grid regions in the set is described. The grid image content corresponding to the changing associated image frames is selected, spliced in spatial order, and pixel-level transition processing is performed on adjacent edges to obtain a continuous orthophoto coverage sequence.
[0006] As a further aspect of the present invention, the image frame land surface category distribution description set includes grid cell identifiers, dominant land surface categories for each grid, land surface category area proportion characteristics, and land surface category spatial consistency indicators. The video change-related frame sequence specifically includes a set of change frame numbers, a change occurrence time identifier, and a set of change grid area indexes. The image frame and orthophoto spatial correspondence description set includes grid area geographic coordinate positioning results, image coordinate and geographic coordinate correspondence indexes, and spatial mapping consistency parameters. The orthophoto continuous coverage sequence includes continuous coverage geographic range, temporal image arrangement order identifiers, and stitching area fusion boundary results.
[0007] As a further aspect of the present invention, the step of obtaining S1 is as follows: S101: Acquire the sequence of video image frames continuously collected by the low-altitude UAV during the patrol flight, read the image content frame by frame in chronological order, call the preset semantic segmentation model for each frame, generate the corresponding category label number for each pixel in the image, record the pixel location area of the four categories of labels (artificial ground, water body, vegetation, and bare ground) in the image, and generate the image frame pixel label distribution result. S102: Based on the pixel label distribution results of the image frame, perform a clustering operation on pixels with the same label number in the image space according to their coordinate adjacency, connect continuously distributed pixel regions, classify and organize all clustered regions according to label category, extract the position coverage relationship of each type of label in the image, and obtain a set of label cluster coverage ranges. S103: Call the tag aggregation coverage set, divide the image into multiple static region units according to the fixed grid division method, read the pixel position covered by the corresponding tag in each unit, and summarize all tag categories in the same unit, sort out the spatial distribution of surface categories in each unit, and obtain the image frame surface category distribution description set.
[0008] As a further aspect of the present invention, the step of obtaining S2 is as follows: S201: Based on the distribution results of the surface category of each grid area corresponding to each image frame in the image frame surface category distribution description set, select any two adjacent images in the order of image frame time, obtain all grid numbers and corresponding surface category label contents in each image frame, and record the frame number and time number for each group of image frames to generate a continuous frame grid label list. S202: Call the continuous frame grid label list, perform a one-to-one matching operation on the surface category labels in the two frames of images at the positions where the grid numbers are consistent, identify the grid numbers where the label content is inconsistent, perform marking processing on the frame numbers corresponding to the areas where the matching fails, extract the sequence numbers of all image frames with label differences, and obtain the set of label difference frame numbers. S203: Based on the set of tag difference frame numbers, read the time sequence number corresponding to each image frame, filter the frame pairs with time intervals within a set range by combining the frame number sequence, organize and classify the image frames that meet the frame sequence continuity condition into change record samples, summarize the frame numbers and time sequence numbers of all samples, and establish a video change associated frame sequence.
[0009] As a further aspect of the present invention, the step of obtaining S3 is as follows: S301: For the change-related image frames in the video change-related frame sequence, obtain the orthophotos generated by the flight area corresponding to the image frames at different time periods, call the edge pixel sets in the image frames and orthophotos, read the brightness gradient values and coordinate positions on the contour boundaries respectively, perform relative position comparison processing on the two sets of edge point sets, and generate edge position matching comparison relationship. S302: Based on the edge position matching relationship, the edge points in the image frame coordinate system are paired with the corresponding points in the orthophoto coordinate system. The center point coordinates of the fixed grid area within the image frame are called and the transformation displacement value in the edge matching result is called. Coordinate mapping processing is performed on all grid areas to obtain the image grid projection coordinate set. S303: Call the image grid projection coordinate set, read the coordinate index of the fixed grid region in the image frame, locate the corresponding projection region in the orthophoto, establish a set of spatial position relationship records for each pair of regions, organize all records according to the image frame number and region number, and establish a spatial correspondence description set between the image frame and the orthophoto.
[0010] As a further aspect of the present invention, the step of obtaining S4 is as follows: S401: Based on the spatial correspondence between the image frame and the orthophoto spatial correspondence description set, the image content corresponding to each grid region in the change-related image frame is extracted in the order of geographic coordinates. The image frame number, grid number and image content index information are read in sequence to establish an image call list based on spatial location sorting and generate a grid image sorting index sequence. S402: Call the grid image sorting index sequence, read the image content of each grid region in the image frame, splice the image segments sequentially according to the arrangement position in the sorting index, arrange each segment continuously in the image coordinate system along the spatial coordinate direction, record the position of each image boundary and the position of the splicing connection line, and obtain the continuous image splicing boundary coordinate set. S403: Based on the continuous image stitching boundary coordinate set, extract the row and column indices of pixels on adjacent stitching boundaries, read the image brightness value and color channel value on both sides of the boundary region respectively, perform average smoothing processing according to the corresponding boundary pixel position, cover the original pixels in the processing area, update the image content to a continuous and unbroken result layer, and generate a continuous coverage sequence of orthophoto images.
[0011] As a further aspect of the present invention, the method further includes: S5: Based on the image content corresponding to the fixed grid area at different times in the continuous coverage sequence of the orthophoto, the surface category results are read according to the fixed grid area division, the image categories are compared across time periods in the same grid area, the spatial location where the surface category name changes is located and the time content is summarized to obtain the low-altitude UAV video frame extraction and orthophoto change detection results. The results of low-altitude UAV video frame extraction and orthophoto change detection specifically include a set of changed geographical locations, a set of change occurrence time information, and a set of surface category change types.
[0012] As a further aspect of the present invention, the step of obtaining S5 is as follows: S501: Based on the image content corresponding to the fixed grid area at different times in the continuous coverage sequence of the orthophoto, call the combination of each grid number and the image frame number to read the covered image content, and use the geospatial coordinate system of the orthophoto as the positioning reference to establish the corresponding index between all image content and spatial coordinates to generate a grid image spatial positioning record. S502: Based on the spatial positioning record of the grid image, extract the image content corresponding to different times in each grid area with the same spatial location, read the surface category label value in the image in chronological order, compare the difference status of the label value in the previous and next images according to the grid number, identify the location of the category name change, and obtain the set of surface label difference coordinates. S503: Call the set of differential coordinates of the surface tags, extract the corresponding grid number and its associated time series number, organize the change content of the surface tags and the corresponding collection time data according to the grid position, summarize all the spatial positions and time relationships containing change records, and obtain the low-altitude UAV video frame extraction and orthophoto change detection results.
[0013] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this invention, by introducing a gridded temporal description based on surface semantics, change identification shifts from simple image difference judgment to category evolution analysis with semantic meaning. The change screening scope focuses on key areas and key moments where surface attribute transfers occur. Combined with a continuous image organization method under a unified geographic coordinate benchmark, the content collected at different time periods maintains a coherent expression in spatial location, thereby strengthening the temporal orientation and spatial certainty of change results and improving the stable expression ability of change information and the clarity of result interpretation in patrol applications. Attached Figure Description
[0014] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart illustrating the acquisition process of S1 in this invention; Figure 3 This is a flowchart illustrating the acquisition process of S2 in this invention; Figure 4 This is a flowchart illustrating the acquisition process of S3 in this invention; Figure 5 This is a flowchart illustrating the acquisition process of S4 in this invention; Figure 6 This is a flowchart of the acquisition process for S5 of the present invention. Detailed Implementation
[0015] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0016] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0017] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0018] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0019] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0020] Please see Figure 1 This invention provides a technical solution: a method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning, comprising the following steps: S1: Acquire the video image frame sequence continuously collected by the low-altitude UAV during the patrol flight, read the image content frame by frame in chronological order, call the pre-trained semantic segmentation neural network model to perform pixel-level semantic recognition processing on each frame image, assign each pixel to one of four land surface category labels: artificial ground, water body, vegetation or bare ground, spatially cluster pixels with the same land surface category label in the image space, divide the image into regions according to a fixed grid method, extract the spatial distribution features of each land surface category label in each grid region, and form an image frame land surface category distribution description set; S2: Based on the surface category distribution results of each image frame in the image frame surface category distribution description set, adjacent image frames are selected in chronological order, and the surface category labels are read at the corresponding grid area positions. The surface category labels of the same grid area in consecutive image frames are compared region by region to identify the image frame positions where the surface category labels change. The image frame numbers and chronological order information of the image frames where the surface category labels change are extracted. The image frames where the surface category labels change are defined as change-related image frames, forming a video change-related frame sequence. S3: For the change-related image frames in the video change-related frame sequence, obtain the orthophotos of the same flight area corresponding to the change-related image frames at different time periods. By extracting and matching edge features of the change-related image frames and orthophotos, construct a spatial mapping transformation relationship based on the matched edge spatial correspondence. Project the fixed grid area in the change-related image frame onto the geographic spatial coordinate system of the orthophoto. Determine the spatial positional correspondence between the fixed grid area in the change-related image frame and the corresponding fixed grid area in the orthophoto. Extract the position of the spatially corresponding area in the image coordinate system to form a spatial correspondence description set between the image frame and the orthophoto. S4: Based on the spatial correspondence between image frames and orthophotos, describe the spatial correspondence of fixed grid regions in the set, select the image content of the corresponding fixed grid region in the change-related image frames according to the spatial position order, and perform spatial sequence stitching on the selected image content so that the change-related image frames collected from different times are arranged continuously in geographical location. Perform pixel-level transition processing on the edge position of adjacent image content in the continuous coverage area to obtain the continuous coverage sequence of orthophotos. S5: Based on the image content corresponding to fixed grid areas at different times in the continuous coverage sequence of orthophotos, the geospatial coordinate system where the orthophoto is located is used as the spatial reference. The surface category results are read according to the fixed grid area division. The content of the old and new images is compared across time periods and regions at the same grid area location to determine the spatial location where the surface category name changes. The acquisition time information and surface category change content corresponding to the change location are summarized to obtain the low-altitude UAV video frame extraction and orthophoto change detection results.
[0021] The image frame land surface category distribution description set includes grid cell identifiers, dominant land surface categories in each grid, land surface category area proportion characteristics, and land surface category spatial consistency indicators. The video change-related frame sequence specifically includes a set of change frame numbers, change occurrence time identifiers, and a set of change grid area indexes. The image frame and orthophoto spatial correspondence description set includes grid area geographic coordinate positioning results, image coordinate and geographic coordinate correspondence indexes, and spatial mapping consistency parameters. The orthophoto continuous coverage sequence includes continuous coverage geographic range, temporal image arrangement order identifiers, and stitching area fusion boundary results. The low-altitude UAV video frame extraction and orthophoto change detection results specifically include a set of change geographic locations, a set of change occurrence time information, and a set of land surface category change types.
[0022] Please see Figure 2 The steps to obtain S1 are as follows: S101: Acquire the sequence of video image frames continuously collected by the low-altitude UAV during the patrol flight, read the image content frame by frame in chronological order, call the preset semantic segmentation model for each frame, generate the corresponding category label number for each pixel in the image, record the pixel location area of the four categories of labels (artificial ground, water body, vegetation, and bare ground) in the image, and generate the image frame pixel label distribution result. Acquire a sequence of continuously captured video image frames from a low-altitude UAV during its patrol flight. Read the video stream data file recorded in the onboard storage device (such as a high-speed SD card or SSD) of the UAV's high-definition optoelectronic pod or digital camera. Set the video stream sampling frame rate parameter to 30fps. Extract single-frame image data at set frame rate intervals. Use an image acquisition card to convert each extracted frame image data into a digital signal in the form of a three-dimensional RGB matrix (dimension: ...). For each frame of the image, the 3D matrix is input into a pre-constructed and trained deep convolutional neural network structure (such as a U-Net network based on the ResNet backbone). The convolutional layers in the network perform convolution operations on the input matrix to extract high-dimensional feature maps. These high-dimensional feature maps are then input into a Softmax classification layer. For each feature vector at a spatial coordinate position, the posterior probability value corresponding to one of four preset categories—artificial ground, water body, vegetation, and bare ground—is calculated. For example, for a pixel at coordinates (100, 200) in a 3840×2160 image, the calculated probability set for the four categories is 0.1, 0.05, 0.8, 0.05 (summing up to 1.0). A maximum index lookup operation is performed on this probability set, identifying the category index corresponding to the maximum probability value of 0.8 as the vegetation category. This category index is then assigned as the label number to the pixel at that coordinate position. This probability calculation and maximum index lookup operation is repeated for all pixel coordinates in the image. A category confidence threshold is then set. If the maximum classification probability of a pixel is lower than the threshold, it is marked as an unclassified region. If it is higher than or equal to the threshold, it is confirmed as the corresponding category. The pixel coordinates of all confirmed categories are counted. The set of pixel coordinates corresponding to artificial ground is denoted as set A, the set of pixel coordinates corresponding to water bodies is denoted as set B, the set of pixel coordinates corresponding to vegetation is denoted as set C, and the set of pixel coordinates corresponding to bare ground is denoted as set D. All two-dimensional coordinate point data contained in each set are recorded to generate the pixel label distribution results of the image frame.
[0023] S102: Based on the pixel label distribution results of the image frame, perform a clustering operation on pixels with the same label number in the image space according to their coordinate adjacency, connect continuously distributed pixel regions, classify and organize all clustered regions according to label category, extract the position coverage relationship of each type of label in the image, and obtain a set of label cluster coverage ranges. Based on the image frame pixel label distribution results, read the pixel two-dimensional coordinate data from sets A, B, C, and D generated in step S101. Taking the vegetation category pixels in set C as an example, select any starting pixel coordinate from set C. Using either Depth-First Search (DFS) or Breadth-First Search (BFS) algorithms, we search for whether a set C contains a coordinate of... , , or If four neighboring pixels exist in the set C, they are considered to belong to the same connected component. The adjacent pixels are then used as new center points to continue the coordinate search until no new neighboring pixels can be found. All pixel coordinates connected by this process are defined as an independent cluster region. The connected component search is repeated for the remaining unvisited pixels in set C until all pixels in set C are assigned to specific cluster regions. For each generated cluster region, the total number of pixels it contains is calculated. A minimum area threshold of 50 pixels is set (corresponding to approximately 0.05 square meters of actual ground surface area, depending on flight altitude and resolution). Cluster regions with less than 50 pixels are identified as salt-and-pepper noise regions and removed. Cluster regions with a total number of pixels greater than or equal to 50 are retained. For each retained cluster region, the x and y coordinates of all pixels within it are traversed, and the minimum x-coordinate is selected. With the maximum value and the minimum value of the ordinate. With the maximum value ; Using formula Calculate the bounding rectangle pixel coverage of the vegetation cluster. For example, if the x-coordinate range of a vegetation cluster is [100, 300] and the y-coordinate range is [200, 400], then its coverage is determined as follows: For pixel regions, the above connected component search, area filtering, and boundary extreme value extraction operations are performed sequentially on the pixel coordinate data in sets A, B, and D. The boundary coordinate parameters of all retained regions are then categorized and summarized to obtain the label cluster coverage set.
[0024] S103: Call the tag aggregation coverage set, divide the image into multiple static region units according to the fixed grid division method, read the pixel position covered by the corresponding tag in each unit, and summarize all tag categories in the same unit, sort out the spatial distribution of surface categories in each unit, and obtain the image frame surface category distribution description set. The system calls the tag aggregation coverage set, sets the fixed grid size parameters, for example, setting the grid cell size to 64×64 pixels. Based on the total image resolution of 3840×2160, the image space is divided horizontally into 60 intervals and vertically into 33.75 intervals. After rounding down, a complete 60×33 grid matrix is formed (edges less than one grid cell are discarded or padded). For each static region cell in the grid matrix, the image coordinate range of the four vertices of that cell is obtained. For example, the image coordinate range of the four vertices of the first static region cell is obtained. Each grid cell covers a pixel x-coordinate range of ; The range of the vertical axis is Traverse every pixel within the coordinate range, read the category label number of that pixel generated in step S101, and set four counters corresponding to artificial ground, water bodies, vegetation, and bare ground respectively. If the label of the currently scanned pixel is vegetation, the vegetation counter value is incremented by 1. After traversing all 4096 pixels in the grid cell, obtain the pixel count value of each category in the cell, and use the formula... Calculate the first The proportion of land-like surfaces within this grid, of which For the first The pixel count value of the class. Given a total of 4096 pixels within the grid, the percentage of vegetation types can be calculated. If the value is greater than the preset dominant category determination threshold of 0.5, then the main surface attribute of the grid cell is marked as vegetation. If the proportion of all categories does not exceed 0.5, then it is marked as "mixed land cover". The four category proportion values calculated for each grid cell are used to form a feature vector. And the eigenvector is compared with the matrix index of the grid cell. Associative storage is performed, and the feature vectors of all grid cells in the entire map are serialized and arranged in row priority order to obtain the image frame surface category distribution description set.
[0025] Please see Figure 3 The steps to obtain S2 are as follows: S201: Based on the distribution results of land surface categories in the grid area corresponding to each image frame in the image frame land surface category distribution description set, select any two adjacent images in the order of image frame time, obtain all grid numbers and corresponding land surface category labels in each image frame, and record the frame number and time number for each group of image frames to generate a continuous frame grid label list. Based on the land surface category distribution results of each image frame in the image frame land surface category distribution description set, the timestamp information of all frames stored in the description set is extracted. With frame index: According to timestamp The incrementing logic sorts and organizes the frame sequence, and sets a loop variable. Traverse from 1 to , retrieve the first Frame and the Frame 1 is a set of adjacent comparison samples, for the 1st frame Frame image, read its contents The data structure of each grid cell, where , Traverse each grid cell Extract the main surface category label of the grid. ,For example At the same time, read the unique frame number corresponding to that frame. and timestamp values , constructing a form like The data record format is also for the first... Each frame performs the same data extraction action to retrieve its corresponding record: , will the Frame and the The data records of each frame are merged into a tuple pair. This operation is repeated until all adjacent frame pairs have been processed, generating a tuple containing... The list structure of the group data generates a list of grid labels for consecutive frames.
[0026] S202: Call the continuous frame grid label list, perform a one-to-one matching operation on the surface category labels in the two frames of images at the positions where the grid numbers are consistent, identify the grid numbers where the label content is inconsistent, perform marking processing on the frame numbers corresponding to the areas where the matching fails, extract the image frame numbers with label differences, and obtain the set of label difference frame numbers. Call the list of consecutive frame grid labels and iterate through each pair of adjacent frames in the list. For each frame pair, set a change grid counter. According to grid coordinates from to Perform a nested loop, reading the data in each iteration. In coordinates Surface category label and At the same coordinates Surface category label Perform a string or number equality check. If the result is... ,For example and If a semantic change has occurred in the grid region, the grid counter will be adjusted accordingly. The value is increased by 1, and the grid coordinates are recorded. After comparing all 1980 grids within the set of frames in the temporary change set, the formula is used. Calculate the rate of change for the frame pair and set a threshold for the significance of inter-frame changes. (That is, 5% of the area changes), if the calculated... > Then determine Compared to For frames with significant changes, Frame number If marked as a difference frame ≤ If a frame is considered redundant or has slight jitter, it is not marked. After traversing all frame pairs, the frame numbers of all frames marked as difference frames are collected, duplicate numbers are removed and sorted in ascending order to obtain a set of labeled difference frame numbers.
[0027] S203: Based on the set of tag difference frame numbers, read the time sequence number corresponding to each image frame, filter the frame pairs with time intervals within the set range by combining the frame number order, organize and classify the image frames that meet the frame sequence continuity condition into change record samples, summarize the frame numbers and time sequence numbers of all samples, and establish a video change associated frame sequence. Based on the set of tag-differential frame numbers, let the sequence of frame numbers contained in the set be... Iterate through each number in the sequence. According to the number Index the corresponding timestamp in the original video data Calculate the time difference between adjacent difference frames. Set minimum time interval threshold Seconds and maximum time interval threshold Seconds, judgment Is it within the range? inside, if The excessively frequent changes suggest that the camera's high-frequency shaking noise might be caused by airflow, and this should be eliminated. Frames are used to ensure sequence sparsity, if This indicates that there are undetected long-term changes or video discontinuities (such as dropped frames), then in and Forced insertion of frames from the middle of the original video sequence As supplementary keyframes, update the sequence. For the final sequence after screening and supplementation Extract the frame number of each frame in the sequence. and its corresponding timestamp Construct a tuple An ordered list, for example, a sequence record as: The ordered list is defined as a summary sequence describing key changes in land surface categories in the video, and a sequence of video change-related frames is established.
[0028] Please see Figure 4 The steps to obtain S3 are as follows: S301: For the change-related image frames in the video change-related frame sequence, obtain the orthophotos generated by the flight area corresponding to the image frames at different time periods, call the edge pixel sets in the image frames and orthophotos, read the brightness gradient values and coordinate positions on the contour boundaries respectively, perform relative position comparison processing on the two sets of edge point sets, and generate edge position matching comparison relationship. For change-related image frames in a video change-related frame sequence, all image frame data in the sequence, along with their corresponding timestamps and latitude / longitude positioning information recorded by the RTK-GPS module on the UAV, are read. Based on the GPS positioning information, a Geographic Information System (GIS) database is queried to obtain orthophoto (DOM) files of that latitude / longitude range acquired through satellite remote sensing or aerial photography in a historical period (e.g., one year ago). For each change-related image frame, it is converted into a grayscale image matrix. Simultaneously, the corresponding orthophoto cropped area is converted into a grayscale matrix. , respectively targeting and Edge detection is performed using the Sobel operator; specifically, it utilizes convolutional kernels. Calculate the brightness gradient of each pixel in the image in the horizontal direction. and the brightness gradient in the vertical direction Using the formula Calculate the gradient magnitude (dimensionless grayscale change rate), set the gradient magnitude threshold to 100, and retain pixels with gradient magnitudes greater than 100 as edge feature points, thus constructing image frame edge point sets respectively. and the set of edge points of the orthophoto ,in and For pixel coordinates, and To utilize Calculate the gradient direction angle, traverse Each edge point in ,exist The search term is in the middle and its Euclidean distance is less than 50 pixels and the gradient direction angle difference is... For each candidate pair of points found, the normalized cross-correlation coefficient (NCC) of the pixel grayscale values within its neighborhood (e.g., an 11×11 window) is calculated. If the NCC value is greater than the matching threshold of 0.85, it is confirmed as a valid matching point pair. All valid matching point pairs are then processed. The coordinate correspondence is stored in a list to generate an edge position matching correspondence.
[0029] S302: Based on the edge position matching relationship, the edge points in the image frame coordinate system are paired with the corresponding points in the orthophoto coordinate system. The coordinates of the center point of the fixed grid area within the image frame are called and the transformation displacement value in the edge matching result. Coordinate mapping is performed on all grid areas to obtain the image grid projection coordinate set. Based on the edge position matching relationship, four sets of non-collinear matching point pairs were randomly selected from the list. Construct a perspective transformation matrix to solve the system of equations, and then iteratively calculate the homography matrix using the least squares method or the RANSAC (random sample consensus) algorithm. The matrix This is a 3×3 transformation matrix used to describe the planar projection mapping from the image frame coordinate system (pixel units) to the orthophoto coordinate system (pixel units). The inlier distance threshold of the RANSAC algorithm is set to 3.0 pixels, and the matrix is optimized through multiple iterations. The parameters are adjusted until the proportion of interior points exceeds 70%, confirming the optimal homography matrix. For each fixed grid region within an image frame (such as the 64×64 grid divided in step S103), the pixel coordinates of its geometric center point are extracted. Represent the coordinates in homogeneous coordinate form. Using the formula Perform matrix multiplication to calculate the projected coordinates of the center point in the orthophoto coordinate system, and then... = / , Normalization is performed (to eliminate scale factors) This yields the precise geographic pixel coordinates of the grid center in the orthophoto. , The above coordinate mapping calculation is performed on the center point and four vertices of all grid regions within the image frame to obtain the image grid projection coordinate set.
[0030] S303: Call the image grid projection coordinate set, read the coordinate index of the fixed grid area in the image frame, locate the corresponding projection area in the orthophoto, establish a set of spatial position relationship record items for each pair of areas, organize all records according to the image frame number and area number, and establish a spatial correspondence description set between the image frame and the orthophoto. The image grid projection coordinate set is called, the projection data of each grid cell in the set is traversed, and the grid index in the image frame is read. and its corresponding orthophoto projection coordinate range (These represent the projected coordinates of the four vertices: top left, top right, bottom right, and bottom left, respectively.) In the large coordinate system of the orthophoto image, the polygon scan transformation algorithm (Scan-line Algorithm) or the bounding box (AABB) detection algorithm is used to locate... The covered orthophoto pixel area is marked as being aligned with the image frame grid. For corresponding regions with spatial co-location relationships, the pixel coordinates are converted into geographic coordinates (longitude and latitude) by reading the Geo Transform six-parameter model (including the latitude and longitude of the top left corner, pixel resolution, etc.) from the orthophoto file header. This results in a data structure containing a quadruple of "image frame ID - grid index - orthophoto coordinate range - geographic coordinate range". For example, the record item might be: The above association operation is repeated for each frame in the video change association frame sequence and all grid cells within it. All generated record items are stored in a hash table or database table and organized with the image frame ID as the first-level index and the grid index as the second-level index to establish a description set of the spatial correspondence between image frames and orthophotos.
[0031] Please see Figure 5 The steps to obtain S4 are as follows: S401: Based on the spatial correspondence between image frames and orthophotos, describe the spatial location correspondence of fixed grid regions in the set, extract the image content corresponding to each grid region in the changed associated image frames in the order of geographic coordinates, read the image frame number, grid number and image content index information in sequence, establish an image call list based on spatial location sorting, and generate a grid image sorting index sequence. Based on the spatial correspondence between image frames and orthophotos, and the spatial location correspondence of fixed grid regions in the description set, the orthophoto projection coordinate range information stored in each record in the description set is extracted. Read the geographic pixel coordinates of the top-left vertex of each rectangular region. As spatial positioning anchor points for this grid region, a structured metadata list containing "image frame index - grid index - anchor point coordinates - original image cropping range" is constructed, with the primary key for spatial sorting set as the anchor point ordinate. The secondary key is the x-coordinate of the anchor point. The metadata list is sorted in ascending order using the quicksort algorithm. Specifically, for any two records in the list... and Comparing its numerical value, if (in (For row height tolerance, set to 5 pixels), then determine Located in space Above, if Then further comparison numerical value, if Then determine Located in space On the left, by recursively performing the above comparison and swap operations, the grid image data, originally arranged in random order by time or frame ID, is reorganized into an ordered sequence that conforms to the geographical spatial distribution pattern from north to south and from west to east. For each element in the sorted sequence, a unique linear sequence index is assigned. For example, the first element in the sequence corresponds to the grid area at the northwesternmost geographical corner, and the second element... Each element corresponds to the southeasternmost grid region. This ordered sequence and its corresponding index mapping are written into the memory buffer to generate a grid image sorted index sequence.
[0032] S402: Call the grid image sorting index sequence, read the image content of each grid region in the image frame, stitch the image segments sequentially according to the arrangement position in the sorting index, arrange each segment continuously in the image coordinate system along the spatial coordinate direction, record the position of each image boundary and the position of the stitching connection line, and obtain the continuous image stitching boundary coordinate set. The grid image sorting index sequence is invoked. Based on the total coverage recorded in the sequence, a 3D zero matrix is initialized as the panoramic stitching canvas, with its size set to [size missing]. ; For example Iterate through each record in the index sequence. Based on the "Original Image Cropping Range" parameter in the record item From the corresponding source image frame Perform matrix slicing operations and read the pixel matrix. Based on the "anchor point coordinates" in the record item. Calculate the target pasting area of the tile on the panoramic canvas, and then... The pixel value is assigned to the corresponding position on the canvas. Before performing the assignment operation, it checks whether there are already non-zero pixel values in the current target area. If there are non-zero pixels, it is determined to be an overlapping area, and the current patch is calculated. With existing tiles Overlapping rectangular range For example, the width of the overlapping region is For each pixel, extract the center line of the overlapping area in the canvas coordinate system as the stitching line, and record the set of all pixel coordinates along the stitching line: The set is then associated with the index numbers of adjacent tiles and stored. After traversing and laying all tiles, the coordinate data of the stitching lines generated in all overlapping areas are summarized to obtain the coordinate set of the continuous image stitching boundary.
[0033] S403: Based on the coordinate set of the continuous image stitching boundary, extract the row and column indices of the pixels on the adjacent stitching boundary, read the image brightness value and color channel value on both sides of the boundary area respectively, perform average smoothing according to the corresponding boundary pixel position, cover the original pixels in the processing area, update the image content to a continuous and unbroken result layer, and generate a continuous orthophoto overlay sequence. Based on the set of continuous image stitching boundary coordinates, traverse each stitching seam line in the set. Set the width parameter of the smooth transition region. The pixel area, defined as a 5-pixel range extending to both sides of the seam line, is used as the blending zone. For each row of pixels within the blending zone (assuming the seam line is vertical), the horizontal coordinate range of that row within the blending zone is obtained. For each pixel coordinate within this range Read the corresponding tile on the left side of the image. raw pixel values and in the right-hand panel raw pixel values Calculate the normalized distance weight of the current pixel from the left non-fusion boundary. ; in are dimensionless coefficients and ; Using linear weighted fusion formula Calculate the merged pixel values, for example, for the leftmost pixel in the merged region. The pixel values are taken entirely from the left image; for the rightmost pixel... The pixel values are taken entirely from the right image, for the pixels at the center position. The pixel value is the average of the two images. Write back to the corresponding position on the panoramic canvas, perform the above weighted operation on the three RGB color channels respectively, eliminate brightness abrupt changes and crack effects at the stitching point, and after all boundary areas have been processed, output the final smooth and complete panoramic matrix to generate a continuous orthophoto overlay sequence.
[0034] Please see Figure 6 The steps to obtain S5 are as follows: S501: Based on the image content corresponding to the fixed grid area at different times in the continuous coverage sequence of orthophotos, call the combination of each grid number and the image frame number to read the covered image content, and use the geospatial coordinate system of orthophotos as the positioning reference to establish the corresponding index between all image content and spatial coordinates to generate grid image spatial positioning records. Based on the image content corresponding to fixed grid areas at different times in a continuous orthophoto cover sequence, each stitched orthophoto file stored in the sequence is read to obtain the acquisition timestamp associated with each image. In addition to the spatial resolution parameters of the image (e.g., 0.1 m / pixel), for each orthophoto, it is divided into grids according to the grid division rules set in step S103. The grid is divided into 100×100 grids within a 1000m×1000m coverage area, and each grid cell is traversed. Calculate its geographic center coordinates Read the surface category attribute of the grid in the current orthophoto. Simultaneously, the corresponding baseline orthophoto of the area is retrieved from the historical database (e.g., data from one year ago), and the same geographic coordinates are read. Historical surface category attributes Construct a data entry containing four attributes: a grid unique identifier. Geographic center coordinates Historical Moments Category Current data collection time category and the current time stamp For example, the record item is: The above information extraction and association operations are performed on all grid cells and all mosaic images. The resulting tens of thousands of records are stored in a memory index table with geographic coordinates as hash keys, generating grid image spatial positioning records.
[0035] S502: Based on the spatial positioning record of grid images, extract the image content corresponding to different times in each grid area with the same spatial location, read the surface category label value in the image in chronological order, compare the difference status of the label value in the previous and next images according to the grid number, identify the location of the category name change, and obtain the set of surface label difference coordinates. Based on the spatial positioning records of the grid imagery, traverse each grid record item in the index table and extract its historical category. Compared with the current category Define the logic for determining the state of the land surface category: If and If the text strings are completely identical, the status of that position is determined to be "unchanged" and the record is skipped; if the strings are inconsistent; For example and If a change in land cover has occurred at that location, the confidence parameter for that grid (if any) is further extracted, and a confidence threshold is set. If the confidence level of the current classification If the change is confirmed to be valid, it is not; otherwise, it is marked as a "suspected change" and added to the list of changes to be reviewed. Calculate the change vector and the geographic center coordinates of the grid Set of vertex coordinates of the mesh boundary The change type labels are stored in a difference set list, and the number of different change types is counted. For example, "vegetation to artificial ground" has 50 grids, and "water body to bare ground" has 20 grids. All difference grids in the list are clustered according to geographical proximity, and the Euclidean distance formula is used to determine if the center distance between adjacent difference grids is less than 1. If so, they are considered as part of the same change event, assigned the same event ID, and the resulting set of surface label difference coordinates is obtained.
[0036] S503: Call the set of differential coordinates of surface labels, extract the corresponding grid number and its associated time series number, organize the change content of surface labels and the corresponding collection time data according to the grid position, summarize all spatial positions and time relationships containing change records, and obtain the low-altitude UAV video frame extraction and orthophoto change detection results. Call the set of surface label difference coordinates and iterate through each aggregated change event in the set. Extract a list of all grid numbers contained in the event. and the corresponding timestamp information Calculate the total area covered by the event. ,in Generate detailed change description text for the actual geographic area of a single grid cell (e.g., calculated as 100 square meters based on resolution), formatted as "in time". Located at coordinates region, area The surface area of square meters is composed of Become This description text, along with the corresponding orthophoto screenshot, the original video frame screenshot, and the binarized masks before and after the transformation, is packaged into a standardized JSON output file, for example: {"event_id":101,"time": Finally, the JSON objects of all events are summarized and sorted by timeline to obtain the low-altitude UAV video frame extraction and orthophoto change detection results.
[0037] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning, characterized in that, Includes the following steps: S1: Acquire video frame sequences of low-altitude drone patrols, identify pixels and assign them to corresponding artificial ground, water bodies, vegetation or bare ground, cluster similar pixels and fix grid division, extract the spatial distribution features of grid surface category labels, and form a set of image frame surface category distribution descriptions. S2: Based on the land surface category distribution results of each image frame in the image frame land surface category distribution description set, select adjacent image frames, compare them in the same grid area, identify the location of land surface category change, and extract the image frame number and time, mark them as change-related image frames, and form a video change-related frame sequence. S3: For the change-related image frames in the video change-related frame sequence, obtain orthophotos, extract and match edge features, construct spatial mapping relationships, project the fixed grid of the image frames onto the geographic coordinate system of the orthophotos, and form a description set of spatial correspondence between image frames and orthophotos; S4: Based on the spatial correspondence between the image frames and the orthophoto spatial correspondence, the spatial position correspondence of the fixed grid regions in the set is described. The grid image content corresponding to the changing associated image frames is selected, spliced in spatial order, and pixel-level transition processing is performed on adjacent edges to obtain a continuous orthophoto coverage sequence.
2. The method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning according to claim 1, characterized in that: The image frame land surface category distribution description set includes grid cell identifiers, dominant land surface categories for each grid, land surface category area proportion characteristics, and land surface category spatial consistency indicators. The video change-related frame sequence specifically includes a set of change frame numbers, a change occurrence time identifier, and a set of change grid area indexes. The image frame and orthophoto spatial correspondence description set includes grid area geographic coordinate positioning results, image coordinate and geographic coordinate correspondence indexes, and spatial mapping consistency parameters. The orthophoto continuous coverage sequence includes continuous coverage geographic range, temporal image arrangement order identifiers, and stitching area fusion boundary results.
3. The method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning according to claim 1, characterized in that: The steps for obtaining S1 are as follows: S101: Acquire the sequence of video image frames continuously collected by the low-altitude UAV during the patrol flight, read the image content frame by frame in chronological order, call the preset semantic segmentation model for each frame, generate the corresponding category label number for each pixel in the image, record the pixel location area of the four categories of labels (artificial ground, water body, vegetation, and bare ground) in the image, and generate the image frame pixel label distribution result. S102: Based on the pixel label distribution results of the image frame, perform a clustering operation on pixels with the same label number in the image space according to their coordinate adjacency, connect continuously distributed pixel regions, classify and organize all clustered regions according to label category, extract the position coverage relationship of each type of label in the image, and obtain a set of label cluster coverage ranges. S103: Call the tag aggregation coverage set, divide the image into multiple static region units according to the fixed grid division method, read the pixel position covered by the corresponding tag in each unit, and summarize all tag categories in the same unit, sort out the spatial distribution of surface categories in each unit, and obtain the image frame surface category distribution description set.
4. The method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning according to claim 1, characterized in that: The steps for obtaining S2 are as follows: S201: Based on the distribution results of the surface category of each grid area corresponding to each image frame in the image frame surface category distribution description set, select any two adjacent images in the order of image frame time, obtain all grid numbers and corresponding surface category label contents in each image frame, and record the frame number and time number for each group of image frames to generate a continuous frame grid label list. S202: Call the continuous frame grid label list, perform a one-to-one matching operation on the surface category labels in the two frames of images at the positions where the grid numbers are consistent, identify the grid numbers where the label content is inconsistent, perform marking processing on the frame numbers corresponding to the areas where the matching fails, extract the sequence numbers of all image frames with label differences, and obtain the set of label difference frame numbers. S203: Based on the set of tag difference frame numbers, read the time sequence number corresponding to each image frame, filter the frame pairs with time intervals within a set range by combining the frame number sequence, organize and classify the image frames that meet the frame sequence continuity condition into change record samples, summarize the frame numbers and time sequence numbers of all samples, and establish a video change associated frame sequence.
5. The method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning according to claim 1, characterized in that: The steps for obtaining S3 are as follows: S301: For the change-related image frames in the video change-related frame sequence, obtain the orthophotos generated by the flight area corresponding to the image frames at different time periods, call the edge pixel sets in the image frames and orthophotos, read the brightness gradient values and coordinate positions on the contour boundaries respectively, perform relative position comparison processing on the two sets of edge point sets, and generate edge position matching comparison relationship. S302: Based on the edge position matching relationship, the edge points in the image frame coordinate system are paired with the corresponding points in the orthophoto coordinate system. The center point coordinates of the fixed grid area within the image frame are called and the transformation displacement value in the edge matching result is called. Coordinate mapping processing is performed on all grid areas to obtain the image grid projection coordinate set. S303: Call the image grid projection coordinate set, read the coordinate index of the fixed grid region in the image frame, locate the corresponding projection region in the orthophoto, establish a set of spatial position relationship records for each pair of regions, organize all records according to the image frame number and region number, and establish a spatial correspondence description set between the image frame and the orthophoto.
6. The method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning according to claim 1, characterized in that: The steps for obtaining S4 are as follows: S401: Based on the spatial correspondence between the image frame and the orthophoto spatial correspondence description set, the image content corresponding to each grid region in the change-related image frame is extracted in the order of geographic coordinates. The image frame number, grid number and image content index information are read in sequence to establish an image call list based on spatial location sorting and generate a grid image sorting index sequence. S402: Call the grid image sorting index sequence, read the image content of each grid region in the image frame, splice the image segments sequentially according to the arrangement position in the sorting index, arrange each segment continuously in the image coordinate system along the spatial coordinate direction, record the position of each image boundary and the position of the splicing connection line, and obtain the continuous image splicing boundary coordinate set. S403: Based on the continuous image stitching boundary coordinate set, extract the row and column indices of pixels on adjacent stitching boundaries, read the image brightness value and color channel value on both sides of the boundary region respectively, perform average smoothing processing according to the corresponding boundary pixel position, cover the original pixels in the processing area, update the image content to a continuous and unbroken result layer, and generate a continuous coverage sequence of orthophoto images.
7. The method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning according to claim 1, characterized in that: The method further includes: S5: Based on the image content corresponding to the fixed grid area at different times in the continuous coverage sequence of the orthophoto, the surface category results are read according to the fixed grid area division, the image categories are compared across time periods in the same grid area, the spatial location where the surface category name changes is located and the time content is summarized to obtain the low-altitude UAV video frame extraction and orthophoto change detection results. The results of low-altitude UAV video frame extraction and orthophoto change detection specifically include a set of changed geographical locations, a set of change occurrence time information, and a set of surface category change types.
8. The method for low-altitude UAV video frame extraction and orthophoto change detection based on deep learning according to claim 7, characterized in that: The steps for obtaining S5 are as follows: S501: Based on the image content corresponding to the fixed grid area at different times in the continuous coverage sequence of the orthophoto, call the combination of each grid number and the image frame number to read the covered image content, and use the geospatial coordinate system of the orthophoto as the positioning reference to establish the corresponding index between all image content and spatial coordinates to generate a grid image spatial positioning record. S502: Based on the spatial positioning record of the grid image, extract the image content corresponding to different times in each grid area with the same spatial location, read the surface category label value in the image in chronological order, compare the difference status of the label value in the previous and next images according to the grid number, identify the location of the category name change, and obtain the set of surface label difference coordinates. S503: Call the set of differential coordinates of the surface tags, extract the corresponding grid number and its associated time series number, organize the change content of the surface tags and the corresponding collection time data according to the grid position, summarize all the spatial positions and time relationships containing change records, and obtain the low-altitude UAV video frame extraction and orthophoto change detection results.