Multi-stage air-ground collaborative data optimization method based on metadata and visual feature fusion

By adopting a multi-stage data optimization method in the spatial and earth collaborative visual group intelligence perception technology, using metadata and visual features for screening and clustering, the complex problems of data redundancy and multi-dimensional coordination are solved, efficient data transmission and processing are achieved, and system efficiency and data quality are improved.

CN120236101APending Publication Date: 2025-07-01AIQUSAIAI (SUZHOU) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510377652.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the existing air-ground collaborative visual group intelligence perception technology, high data redundancy, complex multi-dimensional coordination and low transmission efficiency, resulting in waste of storage, transmission and computing resources, affecting task real-time and system efficiency.

Method used

A multi-stage spatial and ground coordinated data optimization method based on the fusion of metadata and visual features is adopted. Through a multi-stage screening strategy, metadata is used for preliminary filtering, and redundancy is further eliminated by combining spatial direction clustering and visual feature analysis to ensure that the selected data subset is optimal in space-time coverage, perspective diversity and visual representation.

Benefits of technology

It effectively reduces data redundancy, improves data quality and system efficiency, significantly reduces transmission burden and processing complexity, adapts to edge computing environments with limited resources, and supports real-time decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236101A_ABST
    Figure CN120236101A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-stage air-ground collaborative data optimization method based on metadata and visual feature fusion, and the method comprises the steps: collecting data, and obtaining an original data set P; based on pre-screening of metadata, rapidly generating a subset Pv meeting task constraints from the original data set P; performing spatial direction similarity grouping based on spectral clustering, namely clustering a subset Pv according to a spatial position and a shooting direction to generate N image categories; and carrying out image data optimization based on visual features, and screening representative image subsets with significant visual differences from N image categories. According to the method, through a multi-stage screening strategy, primary filtering is carried out by utilizing metadata, redundancy is eliminated in combination with spatial direction clustering and visual feature analysis, and it is ensured that a screened data subset is optimal in space-time coverage, visual angle diversity and visual representativeness; and meanwhile, efficient optimization is realized before data uploading, so that the transmission burden and the processing complexity are remarkably reduced, and the data quality of tasks and the system efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of air-ground collaborative visual crowd sensing, and particularly relates to a multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features. Background Art

[0002] With the wide application of unmanned aerial vehicles (UAVs) and mobile devices, air-ground collaborative visual crowd sensing (AGCVCS) has become an important technical means for achieving multi-perspective and high-coverage data collection. In the prior art, aerial devices (such as UAVs) provide wide-area coverage, and ground devices (such as smartphones) provide fine perspectives. The two are combined to collect data through distributed collaboration, and are widely used in fields such as smart city construction and disaster monitoring. However, due to the differences in perspective, altitude, resolution, and acquisition frequency of the devices, the collected data exhibits high heterogeneity and complex redundancy. For example, UAVs may repeatedly cover areas that have already been collected by ground devices, and independent shootings by multiple users at similar locations further exacerbate data overlap, resulting in waste of storage, transmission, and computing resources. In addition, existing methods mostly rely on post-processing after comprehensive uploading and fail to perform effective screening at the data source. Especially in edge computing environments with limited bandwidth, transmitting unoptimized high-resolution data often causes network congestion and processing delays, significantly affecting task real-time performance and system efficiency.

[0003] Existing data optimization schemes for visual crowd sensing usually perform screening based on a single dimension (such as time or location). For example, some technologies eliminate duplicate data through timestamps, or filter spatially overlapping samples using GPS coordinates. However, these methods do not fully consider the multi-dimensional characteristics of the air-ground collaborative scenario (such as height, angle, and visual content differences), resulting in it being difficult to balance representativeness and redundancy in the screening results. In terms of coordinating multi-dimensional heterogeneous data, traditional methods lack a systematic framework and are unable to effectively integrate metadata and visual features, making it difficult to meet the requirements of AGCVCS tasks for full-range coverage and high-quality samples.

[0004] In addition, existing technologies lack an efficient preprocessing mechanism in resource-constrained scenarios, and directly transmitting large-scale data exacerbates the computing burden and affects real-time decision-making capabilities.

[0005] Therefore, the prior art still needs to be further improved. Summary of the Invention

[0006] Objective of the Invention: To overcome the above deficiencies, the objective of the present invention is to provide a multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features. Through a multi-stage screening strategy, preliminary filtering is carried out using metadata (such as time, location, altitude, angle), and redundant data is further eliminated by combining spatial direction clustering and visual feature analysis to ensure that the selected data subset is optimal in terms of spatio-temporal coverage, perspective diversity, and visual representativeness, and can well solve the technical problems of high data redundancy, complex multi-dimensional coordination, and low transmission efficiency in air-ground collaborative visual crowdsensing.

[0007] Technical Solution: A multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features, comprising:

[0008] S1): Collect data. Fine collection of data in the area is carried out by air equipment and ground equipment respectively to obtain an air-ground image data set, and a subset with high representativeness and low redundancy is selected from the air-ground image data set to obtain an original data set P;

[0009] S2): Preliminary screening based on metadata, that is, using lightweight metadata to quickly eliminate irrelevant data and quickly generate a subset Pv that meets the task constraints from the original data set P;

[0010] S3): Spatial direction similarity grouping based on spectral clustering, that is, the subset Pv in 2) is clustered according to spatial position and shooting direction to generate N image categories, providing a grouping basis for subsequent visual optimization;

[0011] S4): Image data optimization based on visual features. Select a representative image subset with significant visual differences from the N image categories generated in S3), with a total of no more than B images.

[0012] In the multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features described in the present invention, the preliminary screening based on metadata in S2), that is, using lightweight metadata to quickly eliminate irrelevant data and quickly generate a subset Pv that meets the task constraints from the original data set P, is specifically as follows:

[0013] Efficient filtering is carried out using metadata (such as time, location, altitude, resolution). The input data is defined based on the AGCVCS data model, including image ID, task ID, timestamp, GPS coordinates, altitude, and resolution; the screening conditions include:

[0014] (1) Format screening: Retain the formats supported by the task and eliminate incompatible data;

[0015] (2) Spatio-temporal screening: According to the time range and spatial range of the task, select samples that meet the conditions;

[0016] (3) Height screening: According to the task height range, retain the images that meet the air-ground constraints and reduce the perspective overlap;

[0017] (4) Quality screening: Eliminate the low-quality images with resolution lower than the threshold to enhance the visual performance;

[0018] The above conditions are executed in parallel during the screening process, and the intersection is taken to generate the subset Pv.

[0019] In the multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features described in the present invention, the spatial direction similarity grouping based on spectral clustering in S3), that is, the subset Pv in 2) is clustered according to the spatial position and shooting direction to generate N image categories; the specific process is as follows:

[0020] S301): First, perform feature extraction: that is, extract the shooting point features from the metadata of the subset Pv and perform transformation; convert the angle information θ between each shooting point and the target Target into periodic features (cosθ, sinθ), and standardize the spatial position to obtain the standardized spatial coordinates (x, y);

[0021] S302): Similarity calculation, use the Gaussian function to construct the similarity matrix S to capture the spatial and direction characteristics of the data in AGCVCS, and calculate the similarity between shooting points;

[0022] The element S of the similarity matrix s ij The formula is:

[0023] Among them, f i is the feature vector of the i-th shooting point, and f j is the feature vector of the j-th shooting point; ||f i -f j || represents the Euclidean distance between feature vectors; σ is the parameter of the Gaussian kernel, which controls the attenuation speed of similarity, and the similarity matrix S is a symmetric matrix;

[0024] S303): Spectral clustering: Based on the similarity matrix S in S302), construct the Laplacian matrix L = D - S, calculate the eigenvectors corresponding to the first k smallest eigenvalues, and form the matrix U;

[0025] In the position space, use K-means to cluster {u i} to generate N categories, where D is the diagonal matrix, and k is optimized by the silhouette coefficient. Through spatial, direction, and height features, it is expected to aggregate similar air-ground shooting points, reduce redundancy, and coordinate the multi-dimensional distribution;

[0026] The grouping result aggregates the images of similar shooting points into one category, reducing redundancy in space and direction.

[0027] In the multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features described in the present invention, the specific process of optimizing image data based on visual features in S4) is as follows: S401): Extract features from each image, that is, use the SIFT algorithm to extract the key points and descriptors of each image, and generate a feature vector D2;

[0028] S402): Similarity measurement, that is, according to the image features obtained in S401), use FLANN to match the descriptors to calculate the similarity between images. The formula is:

[0029]

[0030] In the formula, S ij represents the similarity score between images I i and I j , N represents the total number of matched feature points, and d i,k d j,k respectively represent the descriptors of the Kth matched point in images I i and I j . ||d i,k -d j,k || represents the Euclidean distance between the matched descriptors; σ2 is a scale parameter used to control the attenuation of the distance weight, where a smaller σ2 value emphasizes the distance difference between the matched points, while a larger σ2 value weakens this difference;

[0031] S403): Construct an undirected graph and set a threshold. Based on the similarity matrix in S402), construct an undirected graph G and set a similarity threshold τ (median) to remove low-similarity edges;

[0032] S404): Maximum independent set search. Based on the undirected graph G constructed in S403), use the greedy algorithm to solve the MIS and optimize the images with significant differences.

[0033] In the multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features described in the present invention, in S403), the undirected graph G=(V, E),

[0034] where the node set V represents the image set, and the edge set E represents the similarity relationship between images;

[0035] The weight w ij of the edge is defined as follows according to the following rules:

[0036]

[0037] The value of the similarity threshold τ is set by calculating the median of all similarity scores, effectively filtering out outliers and low-correlation edges.

[0038] In the multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features described in the present invention, during the process of constructing the undirected graph G in S404), the MIS algorithm is used to optimize the representative subset. The goal of the MIS problem is to find a subset of nodes such that there is no edge connection between any two nodes. Its mathematical expression is as follows:

[0039]

[0040] where x v is a binary variable used to represent the selected state of node v; when x v = 1, it means that node v is selected; when x v = 0, it means that node v is not selected;

[0041] The constraint condition x u + x v ≤ 1, which means that any adjacent nodes cannot be selected into the independent set at the same time, thus ensuring the properties of the independent set.

[0042] From the above technical solutions, it can be seen that the present invention has the following beneficial effects:

[0043] 1. The multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features described in the present invention uses a multi-stage screening strategy. It initially filters using metadata (such as time, location, altitude, angle), and further eliminates redundancy by combining spatial direction clustering and visual feature analysis, ensuring that the selected data subset is optimal in terms of spatio-temporal coverage, perspective diversity, and visual representativeness. At the same time, it can also achieve efficient optimization before data upload, significantly reducing the transmission burden and processing complexity, improving the data quality and system efficiency of the AGCVCS task, and providing support for real-time applications in resource-constrained environments.

[0044] 2. The multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features in the present invention significantly improves the data quality and system efficiency of air-ground collaborative visual crowdsensing. First, through metadata pre-screening, irrelevant data is quickly eliminated, reducing redundancy by 50%-70%, reducing the transmission burden at the edge, and is particularly suitable for bandwidth-constrained environments. Second, spatial direction clustering effectively coordinates the multi-dimensional heterogeneity of air-ground data, reduces spatial and perspective overlaps, and ensures the spatio-temporal and perspective diversity of sample coverage in the target area. Third, visual feature optimization screens representative images with significant differences through SIFT and MIS algorithms, further reducing visual redundancy and enhancing data representativeness. Experiments show that while maintaining a coverage rate of over 90%, this method reduces the data scale to 30%-40% of the original, significantly improving processing efficiency and task accuracy. In addition, this method has a low computational complexity, is adapted to edge computing scenarios, and supports real-time decision-making. Compared with traditional methods, the present invention has obvious advantages in resource utilization, data quality, and system response speed, providing efficient support for applications such as smart cities and disaster monitoring.

[0045] 3. Multi-stage fusion screening: Innovatively combines metadata pre-screening, spatial direction spectrum clustering, and visual feature optimization to gradually achieve data screening from coarse to fine, ensuring efficiency and representativeness. Brief Description of the Drawings

[0046] Figure 1 This is an overview of the multi-stage visual data optimization process of the present invention, showing a schematic diagram of the processing steps from the original data set to the final optimized subset;

[0047] Figure 2 This is an example of visual crowdsensing image optimization in the present invention, showing the process of screening out A and C from three images (A, B, C) through metadata and visual features;

[0048] Figure 3 This is a schematic diagram of the problems existing in the real collected images of the present invention;

[0049] Figure 4 This is the spectral clustering grouping experimental structure of the data set NPU of the present invention;

[0050] Figure 5 This is the spectral clustering grouping experimental result of the data set TOWER of the present invention;

[0051] Figure 6 This is the spectral clustering grouping experimental result of the data set NORMAL of the present invention;

[0052] Figure 7 This is an example of the optimization experimental result of the TOWER data set of the present invention, and the red frame is the optimized image;

[0053] Figure 8This is an example of the preferred experimental results of the COIL-100 dataset in the present invention. The red box indicates the preferred image. Detailed implementation manners

[0054] The present invention will be further clarified below with reference to the accompanying drawings and specific embodiments.

[0055] Embodiment 1

[0056] As Figure 1 shown, a multi-stage air-ground collaborative data selection method based on the fusion of metadata and visual features includes:

[0057] S1): Collect data. Fine collection of data in the area is performed by air equipment and ground equipment respectively to obtain an air-ground image dataset. A subset with high representativeness and low redundancy is selected from the air-ground image dataset to obtain the original dataset P.

[0058] S2): Pre-selection based on metadata, that is, using lightweight metadata to quickly eliminate irrelevant data, and quickly generating a subset Pv that meets the task constraints from the original dataset P.

[0059] S3): Spatial direction similarity grouping based on spectral clustering, that is, the subset Pv in 2) is clustered according to spatial position and shooting direction to generate N image categories, providing a grouping basis for subsequent visual selection.

[0060] S4): Image data selection based on visual features. Representative image subsets with significant visual differences are selected from the N image categories generated in S3), with a total of no more than B images.

[0061] In the multi-stage air-ground collaborative data selection method based on the fusion of metadata and visual features described in this embodiment, the pre-selection based on metadata in S2), that is, using lightweight metadata to quickly eliminate irrelevant data and quickly generating a subset Pv that meets the task constraints from the original dataset P, is specifically as follows:

[0062] Efficient filtering is performed using metadata (such as time, position, height, resolution). The input data is defined based on the AGCVCS data model, including image ID (pid), task ID (tid), timestamp (time), GPS coordinates (locat), height (heig), and resolution (resol); the screening conditions include:

[0063] (1) Format screening: Retain the formats supported by the task (such as.jpg,.png), and eliminate incompatible data;

[0064] (2) Spatiotemporal screening: According to the time range ([T_start, T_end]) and spatial range ([D_min, D_max]) of the task, samples that meet the conditions are screened.

[0065] (3) Height screening: According to the task height range (altRange), retain the images that meet the air-ground constraints and reduce the perspective overlap;

[0066] (4) Quality screening: Eliminate low-quality images with a resolution lower than the threshold (such as 360p) to improve the visual performance;

[0067] The above conditions are executed in parallel during the screening process, and the intersection is taken to generate the subset Pv. This stage runs on edge devices, significantly reducing the amount of data uploaded. For example, for the task TOWER (time: 10:00 - 18:00 on March 14, 2024, height: 0 - 20m), only the images that meet the conditions are retained, and it is expected to reduce 30% - 50% of the redundant data.

[0068] In the multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features described in this embodiment, the spatial direction similarity grouping based on spectral clustering in S3), that is, the subset Pv in 2) is clustered according to the spatial position and shooting direction to generate N image categories; the specific process is as follows:

[0069] S301): First, perform feature extraction: that is, extract the shooting point features from the metadata of the subset Pv and perform transformation; convert the angle information θ between each shooting point and the target Target into periodic features (cosθ, sinθ), and standardize the spatial position to obtain the standardized spatial coordinates (x, y); to ensure that the scales of different features are consistent and prevent a certain feature from dominating the similarity calculation, the mathematical expression is as follows:

[0070]

[0071] S302): Similarity calculation, use the Gaussian function to construct the similarity matrix S to capture the spatial and direction characteristics of the data in AGCVCS, and calculate the similarity between shooting points;

[0072] The element S of the similarity matrix S ij The formula is:

[0073] Among them, f i is the feature vector of the i-th shooting point, f j is the feature vector of the j-th shooting point; ||f i -f j || represents the Euclidean distance between feature vectors; σ is the parameter of the Gaussian kernel, controlling the attenuation speed of similarity, and the similarity matrix S is a symmetric matrix;

[0074] S303): Spectral clustering: Based on the similarity matrix S in S302), construct the Laplacian matrix L = D - S, calculate the eigenvectors corresponding to the first k smallest eigenvalues, and form the matrix U;

[0075] In the position space, use K-means to cluster {u i} to generate N categories, where D is the diagonal matrix, k is optimized by the silhouette coefficient, and through spatial, direction, and height features, it is expected to aggregate similar open-space shooting points, reduce redundancy, and coordinate the multi-dimensional distribution;

[0076] The grouping result aggregates the images of similar shooting points into one category, reducing redundancy in space and direction. For example, the images taken by multiple ground devices at the same location in the task will be clustered into one category.

[0077] To evaluate the effectiveness of the spectral clustering algorithm in grouping spatial direction similarities, the silhouette coefficient is used as the evaluation metric. The silhouette coefficient is a metric for evaluating the quality of clustering results, and its value ranges from -1 to 1;

[0078] The silhouette coefficient comprehensively considers the compactness and separation of clustering. Specifically, for each data point i, calculate its silhouette coefficient s(i):

[0079] where a(i) represents the average distance from data point i to other points within its assigned cluster, which measures the compactness of the clustering; while b(i) represents the average distance from data point i to all points within the nearest neighboring cluster, which reflects the separation of the clustering. The overall silhouette coefficient is the average of s(i) for all data points:

[0080]

[0081] The interpretation of the silhouette coefficient results is as follows:

[0082] s(i) ≈ 1: Data point i is very similar to other points within its assigned cluster and is far from points in other clusters, indicating good clustering;

[0083] s(i) ≈ 0: Data point i is located on the boundary of two clusters, indicating average clustering;

[0084] s(i) ≈ -1: Data point i is wrongly assigned outside the nearest cluster, indicating poor clustering. In this study, by calculating the silhouette coefficient under different numbers of clusters, the number of clusters with the highest silhouette coefficient is selected as the final clustering result to ensure the rationality and stability of the clustering, and finally N image categories are obtained.

[0085] In the multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features described in this embodiment, the specific process of image data optimization based on visual features in S4) is as follows:

[0086] S401): Extract features from each image, that is, use the SIFT algorithm to extract the key points and descriptors of each image, and generate a feature vector D2. That is, the SIFT algorithm effectively captures the core visual information of the image by detecting the key points in the image and generating corresponding descriptors. This algorithm has good invariance to rotation, scaling, and illumination changes, ensuring the robustness and reliability of the extracted feature vectors in complex air-ground collaborative scenarios. After feature extraction, the visual features of the image are quantified into vector form, denoted as

[0087] D = SIFT(I)

[0088] where D represents the set of feature descriptors, and I is the input image; the extracted features not only effectively capture the core visual information of the image but also provide a solid foundation for subsequent similarity measurement and maximum independent set optimization. In addition, due to the good invariance of SIFT features to rotation, scale, and illumination changes, their robustness and reliability in complex scenarios are fully guaranteed, providing important support for the efficient operation of the entire system in large-scale image processing;

[0089] S402): Similarity measurement, that is, according to the image features obtained in S401), use FLANN to match the descriptors to calculate the similarity between images. The formula is:

[0090]

[0091] In the formula, S ij represents the similarity score between images I i and I j , N represents the total number of matched feature points, d i,k d j,k respectively represent the descriptors of the Kth matched point in images I i and I j , ||d i,k -d j,k || represents the Euclidean distance between the matched descriptors; σ2 is a scale parameter used to control the attenuation of the distance weight, where a smaller σ2 value emphasizes the distance difference between the matched points, while a larger σ2 value weakens this difference; this formula weights the distances of the matched points through a Gaussian kernel function and calculates the average similarity between all matched points of two images;

[0092] S403): Construct an undirected graph and set a threshold. Based on the similarity matrix in S402), construct an undirected graph G and set a similarity threshold τ (median) to remove low-similarity edges;

[0093] S404): Maximum independent set search. Based on the undirected graph G constructed in S403), use the greedy algorithm to solve the MIS and preferentially select images with significant differences.

[0094] In the multi-stage air-ground collaborative data preference method based on the fusion of metadata and visual features described in this embodiment, it is characterized in that: in S403), the undirected graph G=(V, E),

[0095] where the node set V represents the image set, and the edge set E represents the similarity relationship between images;

[0096] The weight w of the edge ij is defined as follows according to the following rules:

[0097]

[0098] The value of the similarity threshold τ is set by calculating the median of all similarity scores, effectively filtering out outliers and low-correlation edges.

[0099] In the multi-stage air-ground collaborative data preference method based on the fusion of metadata and visual features described in this embodiment, during the process of constructing the undirected graph G in S404), the MIS algorithm is used to preferentially select the representative subset. The goal of the MIS problem is to find a subset of nodes such that there is no edge connection between any two nodes. Its mathematical expression is as follows:

[0100]

[0101] where x v is a binary variable used to represent the selected state of node v; when x v =1, it means that node v is selected; when x v =0, it means that node v is not selected;

[0102] The constraint condition x u +x v ≤1, which means that any adjacent nodes cannot be selected into the independent set at the same time, thus ensuring the properties of the independent set. This method effectively reduces redundancy and improves data quality. At the same time, through staged processing, the algorithm is expected to reduce visual redundancy in AGCVCS and improve representativeness, ensure computational efficiency, reduce memory consumption, and adapt to the efficient processing of air-ground heterogeneous scenarios.

[0103] Embodiment 2

[0104] The multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features described in this embodiment is the same as the method in Embodiment 1. In addition, when collecting data in S1), AGCVCS aims to achieve multi-perspective and efficient target perception by integrating the visual data of aerial and ground devices. However, the heterogeneity and distributed acquisition characteristics of aerial and ground devices pose significant challenges in obtaining highly representative and low-redundancy data. An ideal dataset needs to strike a balance between spatio-temporal coverage and quality to meet the application requirements of smart cities, disaster monitoring, etc. Therefore, the following key issues in data quality assurance for the AGCVCS task:

[0105] 1) The data redundancy problem under air-ground collaboration. The AGCVCS task relies on the wide-area coverage of aerial devices and the fine-grained acquisition of ground devices, and the data may overlap highly in spatio-temporal distribution. For example, drones may repeatedly cover areas that have already been collected by ground devices, and the independent shootings of multiple users in similar locations further exacerbate the redundancy. This not only increases the consumption of storage and computing resources but also affects the real-time performance of the task due to processing delays. In an edge environment with limited bandwidth, transmitting such data comprehensively will significantly reduce the system efficiency.

[0106] 2) The coordination challenges of multi-dimensional heterogeneous data. AGCVCS needs to meet multi-dimensional requirements such as space, time, and perspective to ensure the all-round perception of the target area. For example, drones provide high-altitude wide-angle views, and ground devices capture low-angle details, and the two need to cooperate to cover different heights and time points. However, the differences in resolution and acquisition frequency between aerial and ground devices make it difficult to balance data diversity and redundancy. How to optimize the data optimization strategy under resource constraints to avoid duplicate coverage while maintaining representativeness has become the core problem in data quality assurance.

[0107] 3) The efficiency requirements for transmission and processing. The high-resolution and distributed characteristics of visual data in AGCVCS pose high requirements for bandwidth and computing power. Aerial devices need to transmit large-capacity data through wireless links, and the intensive uploads of ground devices exacerbate network congestion. In an edge computing scenario, directly transmitting unfiltered data often leads to a sharp increase in latency and cost, affecting real-time decision-making. Therefore, screening the most representative samples before uploading to reduce the transmission burden and improve the processing efficiency is an urgent problem to be solved in the AGCVCS task. The above problems jointly restrict the data quality assurance of AGCVCS, involving three levels: redundancy filtering, multi-dimensional coordination, and efficiency requirements.

[0108] To solve the above problems, a task and data model applicable to AGCVCS is proposed in this embodiment. The AGCVCS task model clarifies the core elements and constraints of the task and adapts to the complex requirements of air-ground collaboration.

[0109] Formally defined as a six-tuple:

[0110] task = {tid, type, whr, whn, angInter, altRange}

[0111] In the above formula, each parameter is shown in the following table:

[0112]

[0113] This model respectively constrains the viewing angle and altitude through angInter and altRange to ensure the multi-dimensional features of the target area covered by the air-ground equipment in cooperation. For example, the task TOWER is defined as:

[0114] task = {tid = 1, type = (.jpg,.jpeg),

[0115] whr = (N34.246,E108.904),

[0116] whn = (202403141000, 202403141800),

[0117] angInter = 45°, altRange = (0, 20m)}

[0118] It is required that the executor take pictures of the landmark TOWER at an angular interval of 45° from the ground to a height of 20 meters from 10:00 to 18:00 on March 14, 2024, and the image data format is jpg or jpeg. Compared with the traditional VCS model, this model adds altRange to distinguish air and ground data, solves the coordination challenges of multi-dimensional heterogeneous data, and initially filters redundancy through angle constraints to provide a basis for subsequent optimization.

[0119] The AGCVCS data model standardizes the representation form of data and supports screening and optimization in the air-ground cooperation scenario. The formal definition is: data = {pid, tid, wid, type, time, locat, heig, resol}

[0120] The meanings of each parameter in the formula are shown in the following table:

[0121]

[0122] The model covers spatio-temporal attributes (time, locat), height (heig), and quality (resol), and supports the management of the heterogeneity of air-ground data. For example, a data record is: data = {pid = 101, tid = 1, wid = 5, type =.jpg, time = 202403141530, locat = (N34.246, E108.905), heig = 10.2m, resol = 1080p}, which means that executor 5 executed task 1 at a height of 10.2 meters at 15:30 on March 14, 2024, and captured high-definition image 101 with the data format of.jpg. The drone high-altitude data and ground data are distinguished by heig. Combining the altRange and angInter of the task model, redundant data (such as duplicate height or angle) can be filtered through metadata to handle the data redundancy problem. At the same time, recording resol and locat supports filtering samples that meet the quality requirements through metadata, improving the transmission efficiency.

[0123] The specific process of extracting the shooting point features from the metadata of subset Pv in the step S301) is as follows:

[0124] Input: dataset P, time range [Tstart, Tend], target location target_locat, distance range [Dmin, Dmax], resolution threshold resolmin, height range altRange Output: filtered subset Pv

[0125]

[0126]

[0127] The specific process of image data optimization based on visual features is as follows:

[0128] Input: N-class image set I = {I1, I2,..., I M}, similarity parameter σ, target number B;

[0129] Output: representative subset Ir (size ≤ B)

[0130]

[0131]

[0132] Embodiment 3

[0133] This embodiment is based on specific experimental data on the experimental results of the multi-stage data optimization algorithm based on the fusion of metadata and visual features, analyzes its screening efficiency and data quality in the air-ground cooperation scenario, and verifies the data quality guarantee effect.

[0134] The AGCVCS experimental dataset includes real shooting data and public datasets. The real shooting data was taken by 22 volunteers using DJI drones and iPhones on and around the campus, covering 348 photos of 3 buildings, and recording metadata (such as GPS, angle, height, time, and resolution). The resolution of the original photos ranges from 960×540 to 5280×2970. The acquisition equipment is equipped with GPS and gyroscopes to ensure metadata accuracy;

[0135] To ensure the basic quality of the dataset, basic cleaning is performed after collection to remove low-quality images, such as blurry, overexposed, and compositionally defective images, such as incomplete targets or insufficient details, combined with automated evaluation (clarity, brightness) and manual verification. Cleaning highlights air-ground collaborative features (such as eliminating high-altitude defocusing of drones or severe ground occlusions) and adapts to visual perception tasks. Figure 5 Common problems in data collection are shown, such as the target being too far away, image blur caused by device movement, the target being partially blocked, and imaging difficulties under strong light conditions. Volunteers shoot from parallel, upward, and downward perspectives to ensure multi-dimensional coverage. After the data is uploaded to the server, a high-quality real data set is formed. The following table lists key statistical information such as the target, number of images, and number of performers.

[0136]

[0137] In addition to the data specially shot for the experiment, the third phase expanded the use of the COIL-100 dataset (7200 images, 128×128, 100 objects shot every 5° rotation) to simulate the change of viewing angle and verify the universality and robustness of the visual feature optimization algorithm. The real data was combined with COIL-100 to comprehensively evaluate the performance of the algorithm in the heterogeneous scenes of AGCVCS.

[0138] The experimental scheme for each stage is as follows:

[0139] The first phase of the experimental plan, the screening criteria are based on metadata:

[0140] (1) Format check. Keep JPEG / JPG / PNG to ensure compatibility;

[0141] (2) Spatiotemporal screening. Keep time∈[Tstart,Tend], di∈[Dmin,Dmax] to ensure relevance;

[0142] (3) Altitude screening. Preserve the air-ground perspective (such as drone high-altitude images) based on heig∈altRange;

[0143] (4) Quality screening. Retain resol≥360p. The experiment is expected to reduce irrelevant data through processing, generate high-quality Pv, and adapt to the AGCVCS heterogeneous scenario. The aim is to verify the efficiency and quality of the section pre-screening algorithm on the AGCVCS dataset.

[0144] The experimental scheme for the second stage. The experiment constructs a scenario in a two-dimensional plane. Shooting points are generated from the target point positions and metadata (locat, θ) of real data. Let the target point be the origin. The shooting points are distributed based on spatial position and angle. Scale differences are eliminated through standardization. The similarity matrix is calculated through Euclidean distance and direction differences, and after mapping to a low-dimensional space, eigen-decomposition is performed. K-means clustering generates categories, and the optimal k is determined by the Silhouette Score. The results are visually displayed in the two-dimensional plane. Thus, the similarity grouping performance of the section spectral clustering algorithm is verified.

[0145] The experimental scheme for the third stage. Through steps such as feature extraction, similarity measurement, graph construction, and MIS search, a low-redundancy and highly representative image subset B is screened out from N categories in the dataset. SIFT feature extraction adopts batch processing and multi-threading to eliminate abnormal images; invalid points are filtered after FLANN matching, and the similarity score is weighted by a Gaussian kernel (σ = 1.0), which reflects the difference between the UAV perspective and the ground perspective in the air-ground scene. The median threshold τ generates a weighted graph G, and the greedy MIS algorithm with pruning and neighborhood optimization selects the subset preferably, expecting to adapt to the efficient processing of the AGCVCS heterogeneous scenario. Thus, the effectiveness of the section visual feature optimization is verified.

[0146] The experimental results of each stage are as follows:

[0147] (1) Results of the pre-screening algorithm. The first-stage experiment verifies the performance of the pre-screening algorithm on three real datasets (NPU, TOWER, NORMAL), as shown in the following table. The screening results show that the algorithm reduces the data by an average of 24% (NPU 29.2%, TOWER 18.6%, NORMAL 24.3%), and retains the images that are highly relevant to the task objectives in terms of time, space, and height. The format, time-space, height, and quality screening work together to improve data relevance (based on heig∈altRange) and clarity (resol≥360p). In terms of computational efficiency, it only takes 0.4 seconds to process 100 images per device, which is suitable for the edge side and reduces the upload burden.

[0148]

[0149] The above table shows the results of the pre-screening algorithm

[0150] The second-stage experiment of the spectral clustering grouping algorithm verified the two-dimensional grouping effect of the spectral clustering algorithm on the pre-screened subset Pv. By analyzing the spatial position and direction information of the shooting points, the relationship diagram between the number of clusters and the silhouette coefficient and the cluster scatter plot were drawn to evaluate and display the grouping effect and verify the effectiveness of the algorithm in the air-ground cooperation scenario. The experiment targeted three datasets, NPU, TOWER, and NORMAL, and analyzed the grouping results separately. Figure 4 Shows the spectral clustering experimental results of the NPU dataset. The left silhouette coefficient curve Figure 4 (a), the horizontal axis represents the number of clusters, and the vertical axis is the corresponding silhouette coefficient. The results show that when k = 4, the silhouette coefficient reaches the maximum value of 0.679, which is the optimal number of clusters. Under the other k values, the fluctuations of the silhouette coefficient reflect suboptimal solutions or over-segmentation deficiencies. Scatter Figure 4 (b), four categories are marked with different colors, and the target points are marked with red stars. The results show that the grouping algorithm based on spectral clustering can effectively distinguish the shooting points according to the spatial position and direction information of the shooting points and correctly group the shooting points with similar spatial characteristics.

[0151] Figure 5 Shows the spectral clustering experimental results of the TOWER dataset. From the silhouette coefficient curve Figure 5 (a), it can be seen that when the number of clusters k = 6, the silhouette coefficient reaches the maximum value of 0.658. Therefore, 6 is selected as the optimal number of clusters. Scatter Figure 5 (b), the algorithm successfully groups the data effectively according to the angular similarity and distance information of the shooting points. The grouping results shown in the figure are clear and have a high degree of separation, further verifying the adaptability of the spectral clustering method, which can handle more complex shooting scenarios and effectively reduce spatio-temporal redundancy.

[0152] Figure 6 Shows the experimental results of the NORMAL dataset. From the silhouette coefficient curve Figure 6 (a), it can be seen that when the number of clusters k = 5, the silhouette coefficient reaches the maximum value of 0.658, which is selected as the optimal number of clusters. Scatter Figure 6 (b) shows that the five categories accurately classify the shooting points with similar directions and spatial distributions, and the grouping consistency is strong, supporting the robustness of the algorithm.

[0153] The above data shows that for both small-scale and larger-scale image datasets, the spectral clustering algorithm can work effectively under different shooting conditions and data distributions. As an evaluation metric for clustering quality, the silhouette coefficient can accurately indicate the optimal number of clusters, making the grouping process more scientific. The classification results after clustering not only have high accuracy but also can visually display the spatial distribution of the shooting points, providing a solid foundation for subsequent optimization based on visual features. This further proves the efficiency and applicability of the spectral clustering method under complex distributions and heterogeneous perspectives, supporting the goal of data quality assurance.

[0154] The third-stage experiment of the image optimization results based on visual features aims to verify the optimization algorithm based on visual features, screen out a subset of B images with low redundancy and high representativeness from the N categories obtained in the second stage, achieve data optimization, and improve the data quality of AGCVCS. The experiment tests self-collected datasets (NPU, TOWER, NORMAL) and the public dataset COIL-100, and evaluates the robustness and efficiency of the algorithm in the air-ground collaborative heterogeneous scenario. Figure 7 Shows an example of some optimization results of the TOWER dataset. From the 6 categories (more than 200 images) in the second stage, a subset is screened based on SIFT features and the MIS algorithm, and the finally optimized images are marked with red frames. The results retain the key perspectives of the target building, reduce visual redundancy while maintaining diversity, and verify the representativeness and filtering effect of the algorithm in the complex scenario of AGCVCS. Figure 8 Shows an example of some optimization results of the target with ID = 66 in the public dataset COIL-100. Through the optimization algorithm based on visual features, 8 of the most representative images are screened out, and the results are as Figure 8 shown. The similarity redundancy of the screened image subset is significantly reduced, and the key features of the 360° rotation are covered. In addition, the experiment shows that the algorithm is efficient in processing large-scale image data. The total time taken to screen all 7200 images of COIL-100 is only 20 seconds, and only 0.2 seconds are required for 100 images. This shows that the algorithm can not only effectively screen out a highly representative image subset but also has good efficiency and can be applied to the processing requirements of large-scale datasets.

[0155] The above experimental results show that the algorithm can perform excellently in different types of datasets. From screening representative images in the multi-class grouping of the AGCVCS self-collected dataset to selecting a subset with significant visual differences from the COIL-100 dataset, the algorithm demonstrates high accuracy and efficiency. The screened image subset maintains a high degree of representativeness in visual features and significantly reduces the redundancy of the dataset. In addition, the application of the MIS algorithm ensures the visual differences between the selected images, and the high efficiency of the algorithm makes it suitable for handling the actual needs of large-scale multi-dimensional heterogeneous datasets in the AGCVCS scenario. Through experiments on different datasets, the applicability of this method is further verified, laying a solid data quality guarantee for the downstream tasks of the subsequent platform.

[0156] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, several improvements can be made without departing from the principle of the present invention, and these improvements should also be regarded as the protection scope of the present invention.

Claims

1. A multi-stage air-ground collaborative data optimization method based on the fusion of metadata and visual features, characterized by: include: S1): Collect data. Use aerial equipment and ground equipment to collect data in the area in detail to obtain an air-ground image dataset. Select a highly representative and low-redundant subset from the air-ground image dataset to obtain the original dataset P. S2): Metadata-based pre-screening, that is, using lightweight metadata to quickly eliminate irrelevant data and quickly generate a subset Pv that meets the task constraints from the original data set P; S3): Spatial direction similarity grouping based on spectral clustering, that is, the subset Pv in 2) is clustered according to spatial position and shooting direction to generate N image categories, providing a grouping basis for subsequent visual optimization; S4): Based on the image data optimization of visual features, a representative image subset with significant visual differences is selected from the N image categories generated in S3), with a total of no more than B images.

2. The multi-stage space-ground collaborative data optimization method based on metadata and visual feature fusion according to claim 1 is characterized by: The metadata-based pre-screening in S2) is to use lightweight metadata to quickly remove irrelevant data and quickly generate a subset Pv that meets the task constraints from the original data set P, as follows: Using metadata for efficient filtering, the input data is defined based on the AGCVCS data model, including image ID, mission ID, timestamp, GPS coordinates, altitude, and resolution; Filters include: (1) Format screening: retain the formats supported by the task and eliminate incompatible data; (2) Spatiotemporal screening: Screen samples that meet the conditions based on the time and space scope of the task; (3) Altitude screening: Based on the mission altitude range, retain images that meet the space constraints and reduce perspective overlap; (4) Quality screening: Eliminate low-quality images with resolution below the threshold to improve visual expression; The screening process executes the above conditions in parallel and takes the intersection to generate the subset Pv.

3. The multi-stage space-ground collaborative data optimization method based on metadata and visual feature fusion according to claim 1 is characterized by: The spatial direction similarity grouping based on spectral clustering in S3), that is, the subset Pv in 2) is clustered according to the spatial position and shooting direction to generate N image categories; the specific process is as follows: S301): first perform feature extraction: extract shooting point features from the metadata of the subset Pv and convert them; convert the angle information θ between each shooting point and the target Target into a periodic feature (cosθ, sinθ), and standardize the spatial position to obtain the standardized spatial coordinates (x, y); S302): Similarity calculation, using Gaussian function to construct a similarity matrix S to capture the spatial and directional characteristics of the data in AGCVCS and calculate the similarity between the shooting points; The elements S of the similarity matrix s ij The formula is: Among them, f i is the feature vector of the i-th shooting point, f j is the feature vector of the jth shooting point; ||f i -f j || represents the Euclidean distance between feature vectors; σ is the parameter of the Gaussian kernel, which controls the decay rate of similarity. The similarity matrix S is a symmetric matrix; S303): Spectral clustering: construct a Laplace matrix L=DS based on the similarity matrix S in S302), calculate the eigenvectors corresponding to the first k smallest eigenvalues, and form a matrix U; In the status space, use K-means to i } Clustering generates N categories, where D is the corner matrix and k is optimized by the silhouette coefficient; The grouping results aggregate images of similar shooting points into one category, reducing redundancy in space and direction.

4. The multi-stage space-ground collaborative data optimization method based on metadata and visual feature fusion according to claim 1 is characterized by: The specific process of optimizing the image data based on visual features in S4) is as follows: S401): extracting features from each image, that is, using SIFT algorithm to extract key points and descriptors from each image, and generating a feature vector D2; S402): Similarity measurement, that is, according to the image features obtained in S401), the descriptors are matched using FLANN to calculate the similarity between the images. The formula is: In the formula, S ij Represents image I i and I j The similarity score between them, N represents the total number of matching feature points, d i,k d j,k Represents image I i and I j The descriptor of the Kth matching point in ||d i,k -d j,k || represents the Euclidean distance between matching descriptors; σ2 is a scale parameter used to control the attenuation of the distance weight, where a smaller σ2 value will emphasize the distance difference between matching points, while a larger σ2 value will weaken this difference; S403): construct an undirected graph and set a threshold, construct an undirected graph G based on the similarity matrix in S402), and set a similarity threshold τ to remove low similarity edges; S404): Maximum independent set search, based on the undirected graph G constructed in S403), a greedy algorithm is used to solve the MIS, and images with significant differences are selected.

5. The multi-stage space-ground collaborative data optimization method based on metadata and visual feature fusion according to claim 4 is characterized by: In said S403), the undirected graph G=(V, E), Among them, the node set V represents the image set, and the edge set E represents the similarity relationship between images; The weight of the edge w ij The following rules are defined as follows: The value of the similarity threshold T is set by calculating the median of all similarity scores, effectively filtering out outliers and low-relevance edges.

6. The multi-stage space-ground collaborative data optimization method based on metadata and visual feature fusion according to claim 5 is characterized by: In the process of constructing the undirected graph G in S404), the MIS algorithm is used to select the representative subset. The goal of the MIS problem is to find a node subset such that there is no edge connecting any two nodes. Its mathematical expression is as follows: Among them, x v is a binary variable used to indicate the selected state of node υ; when x v =1, indicating that node v is selected; when x v =0, it means that node v is not selected; Constraint x u +x v ≤1, which means that any adjacent nodes cannot be selected into the independent set at the same time, thus ensuring the property of the independent set.