Edge-based remote sensing image target detection precision evaluation framework and construction method

Through the edge-based accuracy evaluation framework and dynamic Epsilon edge accuracy function, the problem of distinguishing position error and subject error in remote sensing images is solved, and the detailed evaluation of target edge accuracy and error source analysis are achieved, which improves the reliability of the detection results.

CN120071101APending Publication Date: 2025-05-30ZHEJIANG UNIV

Patent Information

Application Number
CN202510028520.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively distinguish and quantify the position error and subject error of target detection in remote sensing images, resulting in the inability to accurately evaluate the accuracy of the target edge.

Method used

Edge-based accuracy evaluation framework is adopted to identify edge position errors through dynamic Epsilon band edge accuracy function (DEAF), and edge error scores, omission errors and F1-edge scores are calculated by approaching stationary points and intermediate points for detailed geometric and topic accuracy analysis.

Benefits of technology

The detailed evaluation of the edge accuracy of remote sensing image object detection is realized, the algorithm parameters can be optimized, the reliability of the detection results can be improved, and the area error can be decomposed into subject errors and position errors, and the error source can be analyzed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071101A_ABST
    Figure CN120071101A_ABST
Patent Text Reader

Abstract

The invention discloses an edge-based remote sensing image target detection precision evaluation framework and a construction method, and belongs to the field of satellite remote sensing product precision evaluation. The method comprises the following steps: constructing a DEAF function, and calculating a curve of an edge hit rate changing along with bandwidth; key shape points of the DEAF curve are extracted and are used for generating edge-based precision indexes and area precision components; a remote sensing image is input into the evaluation framework, dynamic accuracy is calculated by gradually increasing the width of a buffer band of an object edge, and comprehensive evaluation of the object edge position and the target object quality is achieved. According to the method, the edge position error can be effectively identified, more refined geometric and theme precision analysis is provided, optimization of algorithm parameter selection of remote sensing image target detection is facilitated, and the reliability of a detection result is improved. The method can be widely used for evaluating target detection application products such as agricultural plots and buildings, and can play an important role in the fields of deep learning hyper-parameter optimization, map quality evaluation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of satellite remote sensing product accuracy evaluation, and specifically relates to an accuracy evaluation framework and construction method for remote sensing image target detection based on edges. Background Art

[0002] Target detection in the field of remote sensing is a binary land use and land cover (LULC) classification task for mapping specific geographical features from remote sensing images, including farmland, roads, buildings, etc. Target detection simplifies the comprehensive LULC classification problem into a two-class problem, marking the target object pixels as the foreground category and other object pixels as the background category, saving the workload and budget for training sample collection, verification, and fine processing after classification.

[0003] For remote sensing image mapping projects, accuracy evaluation is an important link for guiding parameter selection and model design. Based on the error sources, pixel-based mapping errors are divided into two categories: thematic errors and positional errors. When the mapping label is inconsistent with the reference label, a thematic error occurs, which is usually caused by insufficient classifier accuracy. Traditional area-based accuracy metrics, such as F1 score, omission error, and misclassification error, are mainly used to measure thematic errors. These metrics rely on the confusion matrix, and evaluate the mapping quality by cross-comparing the model classification results with the reference dataset and quantifying the errors pixel by pixel. This calculation method based on the confusion matrix only reveals the consistency of the evaluation and the reference dataset in terms of regional labels, and it is difficult to reflect the geometric quality and edge accuracy of the target object.

[0004] Positional error is used to measure the distance between the position of the target object on the map and its reference position or the actual ground spatial range, usually characterized by the root mean square error (RMSE). Positional error is affected by various factors such as image resolution, geometric distortion, sensor limitations, and terrain, and is mainly limited by the quality of the source image, especially the quality of high-resolution images. For area-based accuracy evaluation (omission error and misclassification error), there is currently no effective method to distinguish the influence of positional error from that of thematic error.

[0005] Accurately evaluating the target edge is of great significance for studying boundary changes, such as the movement of the coastline, the melting of snow cover, and the expansion of cities. By measuring the edge accuracy, researchers can determine whether the detected boundary changes significantly exceed the uncertainties caused by image distortion or scale errors in reference data annotation. In the field of GIS, directly comparing the reference edge with the evaluated edge remains an important challenge. There is still a lack of methods to separate the position error from the total mapping error, thus unable to quantify the specific impacts of different error components (such as position error and thematic error) on the overall mapping quality. This problem limits the application potential of edge accuracy assessment in remote sensing mapping. Therefore, it is crucial to construct an edge-based accuracy assessment framework to provide an overall method for target detection tasks to evaluate thematic accuracy, position accuracy, and the quality of target objects. Summary of the Invention

[0006] The object of the present invention is to overcome the defects in the prior art and provide an accuracy assessment framework and construction method for target detection of remote sensing images based on edges. The present invention can effectively identify the edge position error, provide a more refined geometric and thematic accuracy analysis, help optimize the selection of algorithm parameters for target detection of remote sensing images, and improve the reliability of detection results. The present invention can be widely used to evaluate target detection application products such as agricultural plots and buildings, accurately evaluate the spatial geometric information of object edges, and can play an important role in fields such as deep learning hyperparameter optimization and map quality assessment. Compared with the traditional pixel-based area accuracy assessment method, the edge-based accuracy assessment framework provides a more detailed geometric and thematic accuracy analysis, which can not only improve the accuracy assessment of the target object edge, but also effectively assist in the optimization of the detection algorithm, thereby improving the overall accuracy and stability of the target detection task.

[0007] The specific technical solutions adopted by the present invention are as follows:

[0008] In the first aspect, the present invention provides a construction method for an accuracy assessment framework for target detection of remote sensing images based on edges, specifically as follows:

[0009] S1: Obtain the data of the remote sensing target detection product to be evaluated, use the visual interpretation method to draw target reference labels in the area to be evaluated, and construct a data set by combining the target detection product and the reference labels;

[0010] S2: Use the target detection product and the target reference labels to be evaluated in S1 to construct a dynamic Epsilon-band edge accuracy function, and construct an edge hit percentage function curve by gradually expanding the buffer width;

[0011] S3: Based on the dynamic Epsilon - band edge precision function in S2, extract the approaching steady points and intermediate points, and calculate the edge misclassification error, edge omission error, and F1 - edge score of the object to be evaluated through the approaching steady points;

[0012] S4: Measure the edge precision of the target detection product through the F1 - edge score in S3; in deep learning, use the F1 - edge score to optimize the hyperparameters of the target detection deep - learning algorithm and select the optimal parameter combination; calculate the shift error of the object to be evaluated through the intermediate points, generate a buffer zone at the object edge in combination with the shift error, and decompose the area misclassification error and omission error into the theme error and position error for error source analysis and quality evaluation.

[0013] Preferably, the construction method of the dynamic Epsilon - band edge precision function in S2 is as follows:

[0014] Create buffer zones at the reference object edge and the object - to - be - evaluated edge respectively, gradually increase the bandwidth and calculate the percentage of the hit part in the total candidate part;

[0015] Use an exponential model to fit the hit percentage curve, and extract three parameters: Nugget, Sill, and Range. The Nugget value is used to quantitatively evaluate the part where the edge of the object to be evaluated completely matches the reference object edge, the Sill value is used to represent the maximum hit percentage, and the Range value is used to represent the distance range where the error is mainly dominated by the theme error.

[0016] Preferably, the edge misclassification error commission edge and the edge omission error omission edge are calculated as follows:

[0017] commission edge = 1 - level_off reference_based (1)

[0018] omission edge = 1 - level_off evaluation_based (2)

[0019] where level_off reference_based and level_off evaluation_based respectively represent the values of the approaching steady points of the dynamic Epsilon - band edge precision function based on the reference object and the object to be evaluated.

[0020] Furthermore, the calculation method of the F1 - edge score F1 edge is as follows:

[0021]

[0022] Preferably, in S4, the greater the shift error, the more significant the geometric misalignment between the object to be evaluated and the reference object.

[0023] Preferably, in S4, the shift error is used to decompose the area-based thematic error into a thematic misclassification error commission thematic and a thematic omission error omission thematic , and the calculation method is:

[0024]

[0025] where area ref_buffer represents the area after adding a buffer zone with the size of the shift distance at the edge of the reference object, and area evaluation_buffer represents the area after adding a buffer zone with the size of the shift distance at the edge of the object to be evaluated; area evaluation represents the area of the object to be evaluated, and area ref represents the area of the reference object;

[0026] The remaining part of the area error, i.e., the positional misclassification error commission positional and the positional omission error omission positional in the positional error, and the calculation method is:

[0027] commission positional = commission area - commission thematic (6)

[0028] omission positional = omission area - omission positional (7)

[0029] where commission area represents the overall misclassification error calculated based on the area, and omission area represents the overall omission error calculated based on the area.

[0030] In a second aspect, the present invention provides an accuracy evaluation framework for edge-based remote sensing image target detection obtained by using any one of the construction methods in the first aspect.

[0031] Preferably, the accuracy evaluation framework is applicable to high-resolution remote sensing image data with a resolution range of 0.5 m to 2.5 m.

[0032] The present invention has the following beneficial effects compared with the prior art:

[0033] The present invention develops an accuracy evaluation framework for object detection in remote sensing images based on edges, which is used to evaluate the accuracy of the edges of object products in remote sensing images. Different from the traditional area-based metric that only reflects pixel-by-pixel consistency, this new method calculates accuracy based on the edges of target objects, aiming to provide a comprehensive evaluation not only for thematic accuracy but also for geometric quality. The core concept is the Dynamic Epsilon-band Accuracy Function (DEAF), which evaluates dynamic accuracy by gradually expanding the buffer zone to cover a wide range of edge misalignments. Two key shape points, namely the approaching steady point and the middle point, are extracted from the DEAF to generate the shift distance, edge-based misclassification / omission errors, and the F1-edge score. The framework proposed by the present invention provides new insights into evaluating the accuracy of target objects from object edges, which helps to develop better object detection products. Description of the Drawings

[0034] Figure 1 is a schematic diagram of the Dynamic Epsilon-band accuracy function adopted by the edge-based object accuracy evaluation framework in the embodiment;

[0035] Figure 2 is an image of the synthesized target object in Embodiment 1, used to simulate the application of the accuracy evaluation framework of the present invention;

[0036] Figure 3 is a comparison of the F1-edge score, edge misclassification error, edge omission error in the framework of the present invention and the traditional F1 score, misclassification error, and omission error in the synthesized image in Embodiment 1;

[0037] Figure 4 is a heat map comparing the F1-edge score in the framework of the present invention with two other evaluation metrics (traditional F1 score, over-segmentation) when using deep learning to select different hyperparameter combinations for farmland object detection in Embodiment 2;

[0038] Figure 5 is a farmland map (white represents farmland) of four example regions predicted by a deep learning model with different hyperparameter combinations in Embodiment 2.

[0039] Figure 6 is a comparison graph of the composition analysis of error sources in the framework of the present invention for two datasets with different resolutions in Embodiment 3. Detailed Embodiments

[0040] The present invention will be further described and explained below in conjunction with the accompanying drawings and specific embodiments. Under the premise that there is no conflict between the technical features of each embodiment of the present invention, corresponding combinations can be made.

[0041] The target detection task in the field of remote sensing is a binary LULC classification task for drawing specific geographical features from remote sensing images, and is widely used in fields such as natural resource monitoring and urban planning. This target detection simplifies the classification problem by marking the target object pixel points as the foreground and the others as the background. However, traditional accuracy evaluation methods such as the evaluation indicators extended from the F1 score and the confusion matrix fail to fully reflect the spatial quality of the target object. Therefore, the present invention proposes a framework for evaluating the accuracy of remote sensing image target detection based on edges and its construction method, focusing on edge accuracy to more comprehensively evaluate the quality of remote sensing target detection. The construction method is as follows:

[0042] S1: Obtain the remote sensing target detection product data to be evaluated, use the visual interpretation method in QGIS software to draw target reference labels in the area to be evaluated, and construct a data set by combining the target detection product and the reference labels.

[0043] As a preferred embodiment of the present invention, the process of drawing the farmland and building labels in the data set by visual interpretation is completed in QGIS software.

[0044] S2: Use the target detection product to be evaluated and the target reference labels in S1 to construct a dynamic Epsilon-band edge accuracy function (DEAF), and construct a curve of the percentage of edge hits by gradually expanding the buffer width.

[0045] As a preferred embodiment of the present invention, the construction of the DEAF function is completed using the Python language. The specific construction method is as follows:

[0046] Create buffer zones at the edges of the reference object and the object to be evaluated respectively, gradually increase the bandwidth and calculate the percentage of the hit part in the total candidate part; use the exponential model to fit the hit percentage curve, and extract three parameters: Nugget, Sill, and Range. The Nugget value is used to quantify the part where the edge of the evaluated object completely matches the edge of the reference object, the Sill value is used to represent the maximum hit percentage, and the Range value is used to represent the distance range where the error is mainly dominated by the thematic error.

[0047] Specifically, the Epsilon band is an uncertain area around the edge of an object defined by a specific width. A dynamic Epsilon band is generated in the binary object detection maps of a pair of objects to be evaluated and reference objects to enclose the edges of the other object. The included edge parts within the Epsilon band are the hit parts, and the unincluded parts are the missed parts. The sum of these two parts is the total candidate part. A hit percentage graph is plotted according to the dynamic edge precision function, that is, the ratio of the hit edge length to the candidate length, as a function of the Epsilon band. As the bandwidth gradually increases, the framework calculates the edge hit percentage for each sampled bandwidth. If the percentage of the continuously hit parts stops changing according to a predefined stability threshold (default is 3%), the Epsilon band expansion calculation will terminate.

[0048] As the Epsilon band expands, when the bandwidth starts to expand, the hit percentage increases sharply at first and then the increasing amplitude gradually becomes smaller. As more and more edges that were not originally matched are included in the Epsilon band, the curve gradually flattens until all edge errors are covered by the Epsilon band. This trend is similar to the characteristics described by the semivariogram in geostatistics, which contains several key variables. Sill represents the point where the semivariogram tends to be stable, indicating the total variance where the feature no longer has spatial correlation. Range is the distance at which the spatial variance reaches 95% of the Sill, that is, the distance at which the data is no longer spatially autocorrelated. Nugget represents the initial value of the semivariogram. Inspired by the semivariogram, an exponential model is used to fit the increase in the percentage of the hit parts as the width of the Epsilon band increases:

[0049]

[0050] where C 0 represents the percentage of the intersection of the edge of the object to be evaluated and the edge of the reference object. C is the maximum percentage increase starting from C 0 when the bandwidth is infinite. a is the scale factor of the bandwidth x in the DEAF model. Nugget in DEAF refers to the proportion of the edges that are completely matched with the reference edge. Sill represents the maximum hit percentage at infinite bandwidth, that is, the total proportion of the position error obtained through the Epsilon band. Range is the distance at which the error source is no longer the position but becomes the subject. Using the above method, two DEAFs are created. The DEAF that establishes the Epsilon band with the edge of the reference object is used to measure the misclassification error, measuring the proportion of the edge of the object to be evaluated outside the band, that is, the misclassification error. The other DEAF establishes the Epsilon band with the edge of the object to be evaluated and calculates the proportion of the edge of the reference object outside the band, which is used to measure the omission error.

[0051] The two definitions of edge error are as follows: Edge misclassification error refers to non-reference object edge pixels being misidentified as edge pixels, and edge omission error refers to reference object edges being omitted and identified as non-edge pixels.

[0052] In practical applications, in combination with Figure 1 , a dynamic Epsilon band is generated in the binary object detection map of a pair of objects to be evaluated and reference objects. It is set that the initial Epsilon band increases by 3 pixel points each time to include the edges of the other object. The included edge parts within the Epsilon band are the hit parts, and the unincluded parts are the omitted parts. The sum of these two parts is the total candidate part. A hit percentage graph is drawn, that is, the ratio of the hit edge length to the candidate length, and thus a dynamic Epsilon band edge accuracy function (DEAF) is generated.

[0053] S3: Based on the dynamic Epsilon band edge accuracy function in S2, the level-off point and the middle point are extracted, and the edge misclassification error, edge omission error, and F1-edge score of the object to be evaluated are calculated through the level-off point.

[0054] As a preferred embodiment of the present invention, two key shape points are extracted according to the DEAF function: the level-off point and the middle point, which are used as the medium for generating the accuracy evaluation index. Specifically as follows:

[0055] The DEAF function is constructed on the edges of the reference object and the object to be evaluated respectively, and the respective level-off points are extracted. The edge misclassification error commission is represented by the distance between the hit percentage of the level-off point and 100% edge and the edge omission error omission edge . The calculation methods of the two errors are as follows:

[0056] commission edge = 1 - level_off reference_based (1)

[0057] omission edge = 1 - level_off evaluation_based (2)

[0058] where level_off reference_based and level_off evaluation_based represent the values of the level-off points of the dynamic Epsilon band edge accuracy functions based on the reference object and the object to be evaluated respectively.

[0059] The F1-edge score calculation method for evaluating the precision quality of the object to be evaluated using the generated edge-based misclassification error and omission error is as follows:

[0060]

[0061] In this embodiment, the point of 0.95*Sill is used as the approaching steady point, and the calculation formula of Sill is:

[0062] Sill = C 0 +C

[0063] The median error of the cartographic edge is represented by using the midpoint between 0.95*Sill and 0, that is, the midpoint, which represents the Epsilon bandwidth for achieving 50% of the maximum hit percentage. The bandwidth of the midpoint is named the shift distance. The larger the shift distance, the greater the difference between the object to be evaluated and the reference object, and the greater the degree of misalignment of the edge in space. The shift distance will be used to decompose the error into a theme error and a position error. The framework generates two sets of precision evaluation indicators by using the approaching steady point and the midpoint: an edge-based precision evaluation indicator and an area-based precision evaluation indicator.

[0064] Specifically, when creating an Epsilon band at the edge of the reference object, the hit percentage is the percentage of the edge of the object to be evaluated included in the Epsilon band based on the reference object. Therefore, the distance between the hit percentage at this approaching steady point and 100% represents the edge misclassification error after excluding the position error. The remaining edge error is mainly attributed to the objects that are completely over-detected, and this error is used to show the quality of the extracted objects. Similarly, by creating an Epsilon band at the edge of the object to be evaluated, the distance between the hit percentage at this approaching steady point and 100% represents the edge omission error.

[0065] S4: Measure the edge precision of the target detection product through the F1-edge score in S3. The higher the edge precision, the better the geometric quality of the target detection result. In deep learning, the F1-edge score is used to optimize the hyperparameters of the target detection deep learning algorithm and select the optimal parameter combination; calculate the shift error of the object to be evaluated through the midpoint, generate a buffer band at the edge of the object in combination with the shift error, and decompose the area misclassification error and omission error into a theme error and a position error for error source analysis and quality evaluation.

[0066] The optimization process of this step combines the F1-edge score, the edge error index, and the traditional area precision index, and optimizes the performance of the target detection algorithm by grid-searching the learning rate and weight decay parameters.

[0067] As a preferred embodiment of the present invention, the DEAF function is used to extract the midpoint, and the Epsilon bandwidth of the midpoint, i.e., the shift error, is obtained. The greater the shift error, the more significant the geometric misalignment between the object to be evaluated and the reference object. In practical applications, first, calculate the area-based error (including the misclassification error and the omission error) of the remote sensing image target detection product to be evaluated. The calculation method is as follows:

[0068]

[0069] Among them, commossopn area represents the overall misclassification error calculated based on the area, omission area represents the overall omission error calculated based on the area, area evaluation represents the area of the object to be evaluated, area ref represents the area of the reference object.

[0070] Using the shift error as the Epsilon bandwidth, buffer zones are generated at the edges of the reference object and the object to be evaluated respectively, and the thematic misclassification error (commission thematic ) and the thematic omission error (omission thematic ) are calculated. The specific calculation method is as follows:

[0071]

[0072] Among them, area ref_buffer represents the area after adding a buffer zone with the size of the shift distance at the edge of the reference object, area evaluation_buffer represents the area after adding a buffer zone with the size of the shift distance at the edge of the object to be evaluated; area evaluation represents the area of the object to be evaluated, area ref represents the area of the reference object;

[0073] The remaining part of the area-based error is used to calculate the positional misclassification error commission positional and the positional omission error omission positional in the position error, that is, the positional misclassification error and the positional omission error in the position error are calculated by eliminating the thematic error from the area-based error, and the analysis of the error source is completed. The calculation method is:

[0074] commission positional = commission area - commission thematic (6)

[0075] omission positional=omission area -omission thematic (7)

[0076] Among them, commission area represents the overall misclassification error calculated based on area, and omission area represents the overall omission error calculated based on area.

[0077] Specifically, the obtained shift error (shift distance) is added to the edge of the reference object as the buffer bandwidth to decompose the area-based thematic error into thematic misclassification error and thematic omission error. The shift distance based on the reference object is added to the edge of the reference object to generate a buffer area (i.e., area ref_buffer ), and the thematic misclassification error is calculated after adding the buffer area. Similarly, the shift distance based on the evaluation object is added to the edge of the object to be evaluated to generate a buffer area (i.e., area evaluation_buffer ), and the thematic omission error is calculated.

[0078] The accuracy evaluation framework obtained by the above construction method can be widely applied to the accuracy evaluation of target detection products such as farmland plots and buildings, and is applicable to high-resolution remote sensing image data with a resolution range of 0.5m to 2.5m.

[0079] The method and effect of the present invention will be illustrated below by examples.

[0080] Example 1

[0081] In this example, a set of synthetic maps ( Figure 2 ) are used to evaluate how effectively the framework of the present invention evaluates the target detection task. The above method of the present invention is used for analysis, and the specific steps are as follows:

[0082] Step 1) Data generation: Use Python to generate synthetic maps. The foreground pixel value of the target object is 1, and the background pixel value is 0. The size of each map is 600*600 pixels, and the width of the target object ( Figure 2 the orange area in) is 500 pixels, and the height varies. Each of the four mapping rows represents a different error scenario, where the first column represents the synthetic map of the reference object, and the second column represents the synthetic map of the object under ideal conditions, used to compare the edge-based accuracy evaluation index based on the present invention and the traditional accuracy evaluation index. The first row and the third row respectively represent (from left to right) an increase in the degree of under-segmentation and over-segmentation. In both cases, the number of edges is incorrect, but the entire target area is largely correct. The second row and the fourth row represent an increase in the error scenarios where the target object is omitted or mis-mapped (i.e., under-detection and over-detection).

[0083] Step 2) Generation of DEAF function: Taking the first column of each row as the reference object and the second to sixth columns as the objects to be evaluated, according to Figure 1 generate a dynamic Epsilon band, obtain the DEAF function for each case, extract the Level-off point of DEAF, generate the marginal misclassification error and marginal omission error for each case, calculate F1-edge, and display the results in Figure 3 .

[0084] Step 3) Result comparison: Use the area-based accuracy evaluation indices (area misclassification error, area omission error, traditional F1 score) and the edge-based accuracy evaluation indices within the framework of the present invention (marginal misclassification error, marginal omission error, F1-edge score) to conduct a comprehensive accuracy evaluation ( Figure 3 ). Display the cases of complete omission or incorrect mapping of the target object ( Figure 3 (a), (d), (f), (h)). In terms of the increasing error levels in the detection candidate map, the region-based and edge-based metrics have similar capabilities. In contrast, the region-based metrics cannot capture the errors in the cases of under-segmentation ( Figure 3 (a), (e)) or over-segmentation ( Figure 3 (c), (g)), while the edge-based accuracy evaluation indices of the present invention effectively quantify the increasing segmentation errors, showing that the traditional area-based indices are insensitive to segmentation errors because they do not show the omission and misclassification errors of object edges. In contrast, the edge-based scores of the present invention can correctly quantify the edge-related errors and have the ability to reflect the errors of the mapping theme ( Figure 3 (b), (d), (f), (h)).

[0085] Example 2

[0086] In this example, Funan County, Fuyang City, Anhui Province, China (32°24'-32°54'N, 115°16'-115°57'E) and Ruian City, Wenzhou City, Zhejiang Province (27°40'-28°0'N, 120°10'-121°15'E) are selected as the study areas, and the above method of the present invention is used for analysis, aiming to verify whether the F1-edge score within the framework proposed by the present invention can achieve more accurate results than the traditional F1 score when classifying remote sensing images using deep learning.

[0087] Step 1) Data acquisition: The GF-2 satellite remote sensing images of the study area in this embodiment are from the public dataset (https: / / github.com / cathy-xie522 / Cropland-Parcel-Dataset). The dataset divides the data from these two regions into 1761 blocks of 512×512 pixels, and then divides them into three parts: the training set (1059 blocks), the validation set (351 blocks), and the test set (351 blocks).

[0088] Step 2) Model construction and prediction: Use the Pytorch framework in Python to construct the ResUnet model. The basic architecture of ResUnet adopts the U-Net structure. The loss function used during model training is the combined loss function of cross entropy loss and Dice loss to reduce the interference caused by sample imbalance. The model is initially trained using the training set. At this stage, the validation set is used to evaluate model overfitting and provide a feedback mechanism for performance improvement. The present invention uses the predicted target detection results from the test set to compare three different accuracy evaluation metrics, namely the F1-edge score, the traditional F1 score, and the root mean square error (RMS) of the segmentation error.

[0089] Step 3) Hyperparameter grid search: Select two key parameters, the learning rate and the weight decay, for parameter optimization. The learning rate controls the update amplitude applied to the model weights during each iteration, thus controlling the convergence speed and the stability of the training process. Weight decay is a regularization technique used to prevent model overfitting, causing the model's parameters to shrink to zero during each update process, thereby limiting its complexity.

[0090] Fine-tune the learning rate and weight decay parameters for grid search and determine the best parameter combination for each evaluation method. The two-dimensional grid search space is defined as: the candidate learning rate values are [0.01, 0.001, 0.0001, 0.00001], and the candidate weight decay values are [0.01, 0.001, 0.0001, 0.00001]. Select the cell that reaches the highest accuracy as the best parameter combination. The present invention tests the F1-edge, the traditional F1 score, and the root mean square (RMS) of over-segmentation and under-segmentation errors as the comparison performance of the grid search accuracy metrics.

[0091] Step 4) DEAF function generation: Use the target detection products predicted by each different combination of learning rate and weight decay (16 kinds) as the objects to be evaluated, and the true labels in the dataset as the reference objects. Apply the DEAF framework respectively. Use F1-edge to characterize the edge accuracy, and use F1 score / RMS to characterize the area accuracy. The results are in Figure 4Shown in

[0092] Step 4) Result comparison: In terms of model hyperparameter selection, the traditional F1-score ( Figure 4 a) and RMS ( Figure 4 b) heatmaps show similar patterns and select the same combination of learning rate (0.001) and weight decay (0.00001). In contrast, the F1-edge of the present invention framework selects a higher learning rate (0.01) and a higher weight decay (0.001). By comparing the best hyperparameter combinations determined by F1-score / RMS and F1-edge, we visually show the quality of the farmland plot object map generated by the ResUnet model through the two hyperparameter combinations ( Figure 5 ). The comparison shows that the hyperparameter combination selected by F1-edge correctly depicts more edges of target objects ( Figure 5 red circles in a, b), and better distinguishes small non-agricultural targets ( Figure 5 red circles in c, d), indicating that F1-edge can help improve hyperparameter selection and thus improve the final accuracy of the target detection task.

[0093] Example 3

[0094] This example selects Pingxiang City, Guangxi Zhuang Autonomous Region, China (21°57'-22°16'N, 106°41'-106°59'E) as the study area, aiming to evaluate the analysis of the theme and location errors of the same building target detection task. The specific steps are as follows:

[0095] Step 1) Data acquisition:

[0096] The remote sensing products of the target detection to be evaluated in Example 2 come from the public building footprint datasets Open buildings (OB, version 3) (https: / / sites.research.google / open-buildings / ) and China building rooftop area dataset (CBRA) (https: / / zenodo.org / records / 7861676). The former has a resolution of 0.5m, and the latter has a resolution of 2.5m. Both datasets are generated by the deep learning Unet model method using remote sensing images of their respective resolutions.

[0097] The reference target detection building footprint map is generated by visual interpretation through the high-resolution Bing map (https: / / www.microsoft.com / en-us / maps / bing-maps / ) in QGIS software.

[0098] Step 2) Generation of DEAF function: Taking OB and CBRA as the objects to be evaluated and the reference building footprint map as the reference object, according to Figure 1 generate a dynamic Epsilon band, generate DEAF functions for two different resolution cases, extract the Level-off point and Middle point of DEAF, and decompose the area-based misclassification error and omission error into their respective thematic errors and positional errors (thematic misclassification error, thematic omission error, positional misclassification error, positional omission error).

[0099] Step 3) Result analysis: Comparison of DEAF key points of OB and CBRA as shown in Table 1. OB has a higher Level-off point than CBRA (97.2% vs 77.4%), showing a lower edge omission error, indicating that a finer spatial resolution of the dataset can more accurately depict building edges. In contrast, the value of the horizontal point of the DEAF function of OB based on the reference object is slightly lower (91.5% vs. 94.3%), perhaps because the finer edge depiction of OB leads to some mis-mapping of building edges (more misclassification errors). The shift distance quantifies the positional movement between the reference object and the object to be evaluated, highlighting the impact of their different resolutions. At a resolution of 0.5 m, the displacement distance based on the reference object is only 2.5 m, and the displacement distance based on the object to be evaluated is only 1.9 m. At a resolution of 2.5 m, these distances increase to 5.2 m and 7.3 m respectively.

[0100] Table 1

[0101]

[0102]

[0103] Table 1 shows the different error source compositions derived from the two datasets. The misclassification error ( Figure 6 d) is mainly caused by their different positional errors (CBRA vs. OB = 28%:14%). The low resolution requires larger mapping units, resulting in many small non-building objects (such as sidewalks) being mapped as buildings. The thematic error accounts for the main proportion of the omission error (CBRA vs. OB = 29%:18%, Figure 6 e). The total area misclassification error (thematic error and positional error) of CBRA is higher than that of OB (CBRA vs. OB = 58%:36%). When applying a wider Epsilon band, most of the area-based misclassification errors caused by the coarser resolution of CBRA can be eliminated, and the error source analysis function finds that the positional error type dominates the misclassification error.

[0104] The present invention proposes a new edge-based precision evaluation framework and its construction method for evaluating target detection tasks in remote sensing images. The core concept of the present invention is the Dynamic Epsilon-band Accuracy Function (DEAF), which evaluates dynamic accuracy by gradually expanding the buffer zone to cover a large range of edge misalignments. Two key points, namely the approaching steady point and the intermediate point, are extracted from the DEAF to generate the shift distance, edge-based omission, misclassification error, and F1-edge score. These metrics can be used for (1) segmentation scale evaluation, (2) deep learning hyperparameter selection, and (3) remote sensing mapping product evaluation, which cannot be achieved by traditional region-based metrics, opening up a new paradigm for evaluating key features of maps (such as object edges) and can be used as an alternative to per-pixel accuracy evaluation.

[0105] The embodiments described above are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting the means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A method for constructing an edge-based remote sensing image target detection accuracy assessment framework, characterized in that: The details are as follows: S1: Obtain the remote sensing target detection product data to be evaluated, draw the target reference labels in the area to be evaluated using the visual interpretation method, and construct a data set by combining the target detection product and the reference labels; S2: Use the target detection product to be evaluated in S1 and the target reference label to build a dynamic Epsilon band edge accuracy function, and build an edge hit percentage function curve by gradually expanding the buffer width; S3: extracting the approximate stable point and the intermediate point based on the dynamic Epsilon band edge accuracy function in S2, and calculating the edge misclassification error, edge omission error and F1-edge score of the object to be evaluated through the approximate stable point; S4: The edge accuracy of the target detection product is measured by the F1-edge score in S3; the F1-edge score is used in deep learning to optimize the hyperparameters of the target detection deep learning algorithm and select the optimal parameter combination; the displacement error of the object to be evaluated is calculated through the intermediate point, and a buffer zone is generated at the edge of the object in combination with the displacement error, and the area misclassification error and omission error are decomposed into subject error and position error to perform error source analysis and quality evaluation.

2. The method for constructing an edge-based remote sensing image target detection accuracy assessment framework according to claim 1, characterized in that: The method for constructing the dynamic Epsilon band edge precision function in S2 is as follows: Create buffer zones at the edge of the reference object and the edge of the object to be evaluated, gradually increase the bandwidth, and calculate the percentage of the hit part to the total candidate part; An exponential model is used to fit the hit percentage curve, and three parameters, Nugget, Sill, and Range, are extracted. The Nugget value is used to quantify the portion of the edge of the evaluation object that completely matches the edge of the reference object. The Sill value is used to indicate the maximum hit percentage. The Range value is used to indicate the distance range in which the error is mainly dominated by the subject error.

3. The method for constructing an edge-based remote sensing image target detection accuracy assessment framework according to claim 1, characterized in that: The edge misclassification error commission edge and edge omission error edge The calculation method is: commission edge =1-level_off reference_based (1) omission edge =1-level_off evaluation_based (2) Among them, level_off reference_based and level_off evaluation_based They represent the values ​​approaching the stable point of the dynamic Epsilon band edge accuracy function based on the reference object and the object to be evaluated respectively.

4. The method for constructing an edge-based remote sensing image target detection accuracy assessment framework according to claim 3, characterized in that: The F1-edge score F1 edge The calculation method is:

5. The method for constructing an edge-based remote sensing image target detection accuracy assessment framework according to claim 1, characterized in that: In S4, the larger the displacement error is, the more significant the geometric misalignment between the object to be evaluated and the reference object is.

6. The method for constructing an edge-based remote sensing image target detection accuracy assessment framework according to claim 1, characterized in that: In S4, the area-based theme error is decomposed into theme misclassification error Commission using the shift error thematic and subject omission error thematic , the calculation method is: Among them, area ref_buffer Represents the area after adding a buffer zone of the displacement distance at the edge of the reference object. evaluation_buffer Indicates the area after adding a buffer zone of the displacement distance at the edge of the object to be evaluated; area evaluation Indicates the area of ​​the object to be evaluated, area ref Indicates the area of ​​the reference object; The remaining part of the area error is the position error commission in the position error. positional and position omission error positional , which is calculated as: commission positional =commission area -commissionthematic (6) omission positional =omission area -omissionthematic (7) Among them, commission area Represents the overall misclassification error based on area calculation, omission area Represents the overall omission error based on the area calculation.

7. An accuracy assessment framework for edge-based remote sensing image target detection obtained using the construction method described in any one of claims 1 to 6.

8. The edge-based remote sensing image target detection accuracy assessment framework according to claim 7, characterized in that: The accuracy assessment framework is applicable to high-resolution remote sensing image data with a resolution range of 0.5m to 2.5m.

Citation Information

Patent Citations

  • No-reference remote sensing image quality evaluation method

    CN117094962A

  • Natural resource remote sensing mapping image positioning method and system

    CN117994678A

Cited By

  • Precision evaluation method and system of measurement method based on deep learning

    CN120804853A