Edge preprocessing method and system for IoT low-orbit satellite sensing data
By standardizing, fusing, and compressing low-Earth orbit satellite IoT data through edge preprocessing methods, the problem of imbalance between satellite-to-ground link resources and efficiency is solved, enabling efficient data transmission and near real-time monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAXIN ZHENGNENG GRP CO LTD
- Filing Date
- 2026-02-09
- Publication Date
- 2026-07-17
AI Technical Summary
In existing technologies for low-Earth orbit satellite IoT, there is an imbalance between resources and efficiency when transmitting multi-source sensing data from the satellite to the ground for processing. This results in delays in the transmission of critical information and redundant data occupying bandwidth in the satellite-to-ground link, making it difficult to support near real-time monitoring of dynamic targets.
An edge preprocessing method is adopted. By receiving and classifying the multi-source raw sensing data stream, standardization processing and spatiotemporal correlation analysis are performed. A lightweight ensemble learning model is used for feature-level fusion and prediction compensation to generate fused feature vectors. Data integrity judgment and quality classification are performed, feature data subsets are extracted, a spatiotemporal dynamic reinforcement grid is constructed, and finally compression and feature extraction are performed to generate optimized data packets for transmission to the ground.
It improved the consistency of spatiotemporal reference and the accuracy of data fusion, reduced redundant data, saved on-board communication resources, simplified ground processing procedures, and enabled near real-time monitoring of dynamic targets on the Earth's surface.
Smart Images

Figure CN121664286B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an edge preprocessing method and system for low-orbit satellite sensing data of the Internet of Things. Background Technology
[0002] In recent years, with the deployment of low-Earth orbit satellite constellations, satellite-based Internet of Things (IoT) has achieved wide-area, continuous sensing of the Earth's surface, demonstrating potential in fields such as environmental monitoring and urban management. However, existing technological solutions face resource and efficiency constraints in transmitting multi-source sensing data from satellites to ground processing. Specifically, it is difficult to balance the limited computing and storage resources on the satellite with the timeliness and effectiveness requirements of ground applications. Currently, the mainstream processing method is "on-board acquisition, ground processing," with the satellite primarily acting as a data relay, transmitting raw data streams from various sensors such as temperature, optics, and synthetic aperture radar through simple timestamps and... After being packaged, the data is directly transmitted via the satellite-to-ground link. However, this method may have the following shortcomings in practical applications. First, the raw sensing data is large in volume and contains a lot of redundant and invalid information (such as invalid optical images obscured by clouds). Direct transmission of this data consumes valuable satellite-to-ground link bandwidth, which may lead to delays in the transmission of critical information. Second, due to the differences in spatiotemporal references and acquisition cycles between different types of sensors (such as optical and infrared), ground systems often need to consume additional resources for complex spatiotemporal registration and interpolation when fusing these heterogeneous data for comprehensive situational analysis. This results in a long analysis process and makes it difficult to support near real-time monitoring of dynamic targets.
[0003] Taking urban flood risk assessment as an example, satellites need to simultaneously acquire visible light images of the area (to identify water bodies), rainfall radar data, and surface temperature information. Existing methods typically transmit these three types of raw data in batches. After receiving the data, the ground system first needs to perform spatial registration and interpolation on images and radar data of different time phases and resolutions before it can perform fusion analysis. This process may result in delays of several hours, and the entire set of data transmitted may include infrared data of large areas of cloudless, clear areas (which have limited value for flood assessment), resulting in a double consumption of transmission and processing resources. Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide an edge preprocessing method and system for IoT low-orbit satellite sensing data, which reduces the data transmission load of the satellite-to-ground link, saves valuable on-board communication resources, avoids critical information delays caused by redundant data occupying transmission bandwidth, and improves the timeliness of data downlink.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] Firstly, an edge preprocessing method for IoT low-Earth orbit satellite sensing data, the method comprising:
[0007] Step 1: Receive and classify the multi-source raw sensing data stream, generate a primary data frame set, and perform standardization processing to generate a standardized data frame set;
[0008] Step 2: Analyze the spatiotemporal correlation of the standardized data frame set, dynamically determine the fusion parameters and perform interpolation alignment to generate a multi-source fusion data field, and use a pre-trained lightweight ensemble learning model to perform feature-level fusion and prediction compensation on the multi-source fusion data field to generate a fusion feature vector.
[0009] Step 3: Perform data integrity assessment and quality classification on the fused feature vectors to generate an optimal feature dataset, and extract feature data subsets corresponding to urban built-up areas, vegetation coverage areas and water areas from the optimal feature dataset;
[0010] Step 4: Perform feature matching and displacement calculation on the serialized image data in the feature data subset to generate motion trajectory sequences of geographic entities. Based on the location points in the motion trajectory sequences, calculate the approximate total length of the spatial curve of each trajectory as the geometric feature parameter of the trajectory.
[0011] Step 5: Combine the geographical distribution, motion trajectory sequence and geometric feature parameters of the feature data subset to construct a spatiotemporal dynamic enhancement grid, and calculate the spatiotemporal situation confidence enhancement factor to perform weighted correction on the selected feature dataset and generate a spatiotemporal situation optimized feature dataset.
[0012] Step 6: Compress and extract features from the spatiotemporal situation optimization feature dataset to generate edge preprocessing result data packets, and transmit them to the ground data center via the satellite-to-ground link.
[0013] Secondly, the edge preprocessing system for IoT low-orbit satellite sensing data includes:
[0014] The standardization module is used to receive, classify, and buffer multi-source raw sensing data streams, generate a primary data frame set, and perform standardization processing to generate a standardized data frame set.
[0015] The feature compensation module is used to analyze the spatiotemporal correlation of the standardized data frame set, dynamically determine the fusion parameters and perform interpolation alignment, generate a multi-source fused data field, and use a pre-trained lightweight ensemble learning model to perform feature-level fusion and prediction compensation on the multi-source fused data field to generate a fused feature vector.
[0016] The feature extraction module is used to perform data integrity judgment and quality classification on the fused feature vector, generate the preferred feature dataset, and extract feature data subsets corresponding to urban built-up areas, vegetation coverage areas and water areas from the preferred feature dataset;
[0017] The geometric quantization module is used to perform feature matching and displacement calculation on the serialized image data in the feature data subset, generate the motion trajectory sequence of geographic entities, and calculate the approximate total length of the spatial curve of each trajectory based on the position points in the motion trajectory sequence, as the geometric feature parameter of the trajectory.
[0018] The grid optimization module is used to construct a spatiotemporally dynamic enhanced grid by combining the geographical distribution, motion trajectory sequence and geometric feature parameters of the feature data subset, and to calculate the spatiotemporal situation confidence enhancement factor to perform weighted correction on the preferred feature dataset and generate a spatiotemporally optimized feature dataset.
[0019] The uplink transmission module is used to compress and extract features from the spatiotemporal situation optimization feature dataset, generate edge preprocessing result data packets, and transmit them to the ground data center via the satellite-to-ground link.
[0020] Thirdly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0021] The above-described solution of the present invention has at least the following beneficial effects:
[0022] Unifying the spatiotemporal reference and format specifications of data eliminates heterogeneity issues caused by different sensors and data protocols, improving the coherence and effectiveness of the overall data processing workflow. Based on spatiotemporal correlation, fusion parameters are dynamically determined and interpolation alignment is completed. Combined with a lightweight ensemble learning model, feature-level fusion and system bias prediction compensation of multi-source data are achieved, effectively improving the fusion accuracy of multi-source sensing data, reducing data bias caused by observation conditions and sensor differences, and strengthening the feature representation capability of fused data. Accurate filtering of low-quality and invalid data is achieved through data integrity judgment and quality grading. Simultaneously, feature data subsets are extracted according to typical geographical regions, enabling refined screening and classification of sensing data and reducing the proportion of redundant data. Geographic entity motion trajectories are calculated from serialized image data, and geometric feature parameters are extracted to mine the dynamic information of geographic entities behind the sensing data, enriching the feature dimensions of the data. A spatiotemporal dynamic enhancement grid is constructed by combining geographical distribution, motion trajectory, and geometric features. Weighted correction of selected data is performed using a spatiotemporal situation confidence enhancement factor, realizing the enhancement of sensing data... The spatiotemporal situation optimization makes the spatiotemporal distribution characteristics of the data more consistent with the actual situation, effectively improving the credibility and application value of the data. After the optimized feature dataset is compressed and features are extracted, it is transmitted through the satellite-to-ground link, reducing the data transmission load of the satellite-to-ground link, saving valuable on-board communication resources, and avoiding the delay of key information caused by redundant data occupying transmission bandwidth, thus improving the timeliness of data downlink. The edge end completes the entire process of data standardization, fusion, filtering, optimization, and compression. The ground data center receives the optimized feature data packets, eliminating the need to consume a lot of resources for complex spatiotemporal registration, data cleaning, and heterogeneous fusion, simplifying the ground data processing process, improving the efficiency of ground analysis applications, and realizing near real-time monitoring of dynamic targets and spatiotemporal situation on the surface. The entire edge preprocessing process adopts lightweight algorithms and models, which are adapted to the computing and storage resource constraints of low-orbit satellite on-board processors or near-Earth edge computing nodes. While achieving efficient data processing, it will not cause excessive consumption of on-board resources, balancing data processing effect and on-board resource adaptability. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the edge preprocessing method for IoT low-orbit satellite sensing data provided in an embodiment of the present invention.
[0024] Figure 2 This is a schematic diagram of an edge preprocessing system for IoT low-orbit satellite sensing data provided in an embodiment of the present invention. Detailed Implementation
[0025] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0026] like Figure 1 As shown, embodiments of the present invention propose an edge preprocessing method for IoT low-Earth orbit satellite sensing data, the method comprising the following steps:
[0027] Step 1: Receive and classify the multi-source raw sensing data stream, generate a primary data frame set, and perform standardization processing to generate a standardized data frame set;
[0028] Step 2: Analyze the spatiotemporal correlation of the standardized data frame set, dynamically determine the fusion parameters and perform interpolation alignment to generate a multi-source fusion data field, and use a pre-trained lightweight ensemble learning model to perform feature-level fusion and prediction compensation on the multi-source fusion data field to generate a fusion feature vector.
[0029] Step 3: Perform data integrity assessment and quality classification on the fused feature vectors to generate an optimal feature dataset, and extract feature data subsets corresponding to urban built-up areas, vegetation coverage areas and water areas from the optimal feature dataset;
[0030] Step 4: Perform feature matching and displacement calculation on the serialized image data in the feature data subset to generate motion trajectory sequences of geographic entities. Based on the location points in the motion trajectory sequences, calculate the approximate total length of the spatial curve of each trajectory as the geometric feature parameter of the trajectory.
[0031] Step 5: Combine the geographical distribution, motion trajectory sequence and geometric feature parameters of the feature data subset to construct a spatiotemporal dynamic enhancement grid, and calculate the spatiotemporal situation confidence enhancement factor to perform weighted correction on the selected feature dataset and generate a spatiotemporal situation optimized feature dataset.
[0032] Step 6: Compress and extract features from the spatiotemporal situation optimization feature dataset to generate edge preprocessing result data packets, and transmit them to the ground data center via the satellite-to-ground link.
[0033] In this embodiment of the invention, a unified spatiotemporal reference and format specification for data are established to eliminate heterogeneity issues caused by different sensors and data protocols, thereby improving the coherence and effectiveness of the overall data processing flow. Based on spatiotemporal correlation, fusion parameters are dynamically determined and interpolation alignment is completed. A lightweight ensemble learning model is used to achieve feature-level fusion of multi-source data and system bias prediction compensation, effectively improving the fusion accuracy of multi-source sensing data, reducing data bias caused by observation conditions and sensor differences, and strengthening the feature expression capability of the fused data. Low-quality and invalid data are accurately filtered out through data integrity judgment and quality grading. Simultaneously, feature data subsets are extracted according to typical geographical regions to achieve refined screening and classification of sensing data, reducing the proportion of redundant data. Geographic entity motion trajectories are calculated from serialized image data, and geometric feature parameters are extracted to mine the dynamic information of geographic entities behind the sensing data, enriching the feature dimensions of the data. A spatiotemporal dynamic enhancement grid is constructed by combining geographical distribution, motion trajectory, and geometric features. The selected data is weighted and corrected using a spatiotemporal situation confidence enhancement factor, achieving… The spatiotemporal situation optimization of the sensing data makes the spatiotemporal distribution characteristics of the data more closely match the actual situation, effectively improving the data's credibility and application value. The optimized feature dataset is compressed and features extracted before being transmitted via the satellite-to-ground link, reducing the data transmission load on the satellite-to-ground link, saving valuable onboard communication resources, and avoiding critical information delays caused by redundant data occupying transmission bandwidth, thus improving the timeliness of data transmission. The edge performs a full-process preprocessing of data, including standardization, fusion, filtering, optimization, and compression. The ground data center receives the optimized feature data packets, eliminating the need for complex spatiotemporal registration, data cleaning, and heterogeneous fusion, simplifying the ground data processing flow, improving the efficiency of ground analysis applications, and enabling near real-time monitoring of dynamic targets and spatiotemporal situations on the Earth's surface. The entire edge preprocessing process uses lightweight algorithms and models, adapting to the computing and storage resource constraints of low-Earth orbit satellite onboard processors or near-Earth edge computing nodes. This achieves efficient data processing without excessive consumption of onboard resources, balancing data processing effectiveness with onboard resource compatibility.
[0034] In a preferred embodiment of the present invention, step 1 includes:
[0035] Step 100: At the onboard processor of the low-Earth orbit satellite or the near-Earth edge computing node deployed at the ground station, receive raw sensing data streams from various types of IoT sensors. Based on the source sensor type and data protocol of the raw sensing data streams, perform multi-path parallel parsing on the raw sensing data streams. Define the smallest data block with a complete protocol structure obtained from the parsing as a data unit. Based on the source identifier carried by each data unit, classify and store them in independent buffers allocated in the memory of the onboard processor or edge computing node, corresponding to each data source, forming a primary data frame set organized by data source identifier. Specifically, this includes: first, activating the multi-channel data receiving interface of the onboard embedded processor or near-Earth edge computing node. This interface has completed deep adaptation to the dedicated communication protocols of various IoT sensors, such as temperature sensors, optical sensors, synthetic aperture radar sensors, rain radar sensors, and electromagnetic signal detection sensors. It can simultaneously achieve continuous and parallel reception of raw sensing data streams uploaded in real time from various types of sensors. All raw sensing data streams transmitted by sensors are continuously... End-to-end transmission and reception are completed in the form of continuous data streams, ensuring the continuity, integrity, and real-time nature of data reception throughout the process, with no data interruptions or missed acquisitions. Then, the parallel parsing engine of the onboard embedded processor or near-ground edge computing node is invoked. Based on the source sensor type and dedicated data transmission protocol corresponding to each original sensing data stream, multi-channel parallel parsing operations are performed on the received multi-source continuous data streams. During the parsing process, the protocol format requirements corresponding to each type of sensor are followed, and the continuous data streams are precisely split bit by bit. At the same time, the protocol structure integrity of each data block after splitting is verified in all fields of the protocol header, protocol body, and check bits. Only the data blocks that pass the verification are retained. The smallest data block with a complete protocol structure confirmed by verification is defined as a data unit. Each data unit carries a unique data source identifier, which includes information such as sensor type code, sensor number, and acquisition node code. It is precisely matched one-to-one with the IoT sensor that acquired and generated the data unit, ensuring that the source of each data unit can be traced throughout the entire process.
[0036] Finally, in the physical memory space of the onboard embedded processor or near-Earth edge computing node, a dedicated independent physical buffer is pre-allocated for each independent data source. The initial storage capacity of each physical buffer is calculated and allocated based on the fixed data acquisition frequency of the corresponding sensor and the fixed byte size of a single data unit. The specific calculation formula is: Initial storage capacity = Fixed data acquisition frequency (unit: data entries / second) × Byte size of a single data unit (unit: bytes / data entry) × Preset buffer duration (uniformly set to 60 seconds to adapt to the low-Earth orbit satellite transit observation cycle). At the same time, 20% of the redundant capacity is reserved to cope with sudden data fluctuations. The final allocated initial capacity is the sum of the calculated value and the redundant capacity. The buffer supports dynamic expansion and contraction of storage capacity based on the dynamic changes in the actual acquisition frequency of the sensor and the byte fluctuations of the data unit. The specific logic is to monitor the percentage of used capacity of each buffer in real time. When the capacity exceeds 80% (expansion threshold), dynamic expansion is triggered, with an expansion rate of 50% of the current capacity. The total capacity after expansion shall not exceed the upper limit of the reserved memory on the satellite or edge node. When the used capacity percentage is below 30% for 30 consecutive seconds (shrinkage threshold), dynamic shrinkage is triggered, shrinking to 1.5 times the current used capacity, but not less than 80% of the initial storage capacity. This avoids frequent expansion and shrinkage, ensuring that the buffer storage capacity is highly matched with the actual data storage needs, without excessive or insufficient memory resources. Subsequently, each data unit is parsed one by one, and its data source identifier is read. Based on the identifier, each data unit is accurately classified and stored in the corresponding independent physical buffer. Each buffer stores the data units in order of data reception time. All the classified and organized data units stored in all independent physical buffers together form a primary data frame set organized in order of data source identifier.
[0037] Step 101 involves registering the image data and electromagnetic signal data in the primary data frame set using spatial coordinates and synchronizing the timestamps to generate registered and aligned image and signal data. It also involves unifying the units of measurement for the environmental parameter time-series data, detecting and removing outliers in the converted data, and interpolating and filling in any missing data points to generate processed environmental parameter data. Specifically, this includes extracting all image data and electromagnetic signal data from the primary data frame set in step 100, and reading the original spatial coordinate system parameters and original timestamp information for both types of data. The original spatial coordinate system for the image data is the sensor's... For example, the original spatial coordinate system for electromagnetic signal data is the sensor detection coordinate system, and the original timestamp is the local acquisition time of each sensor. A spatial coordinate system registration operation is performed to determine a unified geospatial coordinate system (using the WGS84 geodetic coordinate system) as the registration target coordinate system. Based on the parameter differences between the original and target coordinate systems, the translation, rotation, and scaling matrices are calculated. The specific calculation process is as follows: the translation matrix is obtained by solving for the three-dimensional coordinate offset of the origin of the original coordinate system in the target coordinate system. Let the rectangular coordinates of the origin of the original coordinate system corresponding to the target coordinate system be (X0, Y0, Z0), then the translation matrix is a 3×1 matrix. The matrix elements directly correspond to the origin offset; the rotation matrix is calculated based on the azimuth α, pitch β, and roll γ of the original coordinate system relative to the target coordinate system. It is derived using the Euler angle transformation formula, first converting the azimuth, pitch, and roll angles into rotation matrices for their respective axes, and then obtaining the final 3×3 rotation matrix through matrix multiplication. The elements of each axis rotation matrix are composed of trigonometric function values (e.g., the elements of the rotation matrix around the Z-axis are cosα, -sinα, sinα, and cosα); the scaling matrix is calculated based on the conversion relationship between the units of the original and target coordinate systems. The target coordinate system unit is meters. The scaling matrix for image data (sensor imaging coordinate system, unit is pixels) is a 3×3 diagonal matrix, where the diagonal elements are equal to the actual ground distance corresponding to the sensor pixel resolution (i.e., the scaling factor). For optical sensors, the scaling factor is 0.1 meters / pixel. For synthetic aperture radar... The sensor scaling factor is set to 0.5 meters per pixel, and the scaling factor for electromagnetic signal data (sensor detection coordinate system, unit: signal sampling point) is set to 1 meter per sampling point. The diagonal elements of the scaling matrix correspond to the scaling factors for the X, Y, and Z axes, respectively, while the off-diagonal elements are all 0. Coordinate transformation is performed using the formula: target spatial coordinates = original spatial coordinates × rotation matrix + translation vector. Then, a scaling operation is performed on the transformed coordinates: target coordinate value = transformed coordinate value × corresponding scaling factor (0.1 or 0.5 for image data depending on sensor type, 1 for electromagnetic signal data). The spatial coordinates of all image and electromagnetic signal data are unified to the target geospatial coordinate system, completing spatial coordinate system registration. A timestamp synchronization alignment operation is performed, selecting the UTC (Coordinated Universal Time) time provided by a high-precision rubidium atomic clock on a low-Earth orbit satellite as the unified time reference. This time reference has an accuracy of [insert accuracy here]. At the second (nanosecond) level, time drift can be calibrated using satellite ephemeris data. The time difference between the original timestamp of each data point and the unified time reference is read, and the time correction amount is calculated. The time correction amount = unified time reference value - original timestamp value. The original timestamps of each data point are superimposed with the corresponding time correction amount to obtain the calibrated unified timestamp. The timestamp synchronization and alignment of all image data and electromagnetic signal data are completed, and finally, the registered and aligned image and signal data are generated.
[0038] All environmental parameter time-series data were extracted from the initial data frame set. This data includes different types of environmental monitoring parameters such as temperature, humidity, air pressure, and rainfall. Each parameter has different units of measurement. The humidity parameter was initially collected as relative humidity, with values derived from real-time data collected by humidity sensors. The value range is 0 to 100 (unitless, representing the percentage of actual water vapor content in the air compared to saturated water vapor content at the same temperature). Following internationally accepted metrological standards, a unified conversion operation was performed on environmental parameters with different units of measurement. The conversion process followed fixed metrological conversion formulas, such as Kelvin temperature = Celsius temperature + 273.15, and absolute humidity = relative humidity × saturated water vapor pressure at the same temperature ÷ air pressure. This completed the unified conversion of the units of measurement for all environmental parameter time-series data. Outlier detection and removal were then performed on the converted environmental parameter time-series data. The 3σ principle was used to determine outliers. First, the unit of measurement for each type of environmental parameter dataset was calculated. The mean μ is calculated, and then the sample standard deviation σ of the dataset is calculated. The normal range of this type of environmental parameter is determined to be μ-3σ to μ+3σ. All data exceeding this range are identified as outliers and removed directly from the dataset. For the environmental parameter time series data after outlier removal, a data missing point interpolation operation is performed. Linear interpolation is used to calculate missing values. First, the missing point in the dataset is located. Two adjacent valid data points are read. Let the timestamp of the previous valid data point be x1 and the parameter value be y1, the timestamp of the subsequent valid data point be x2 and the parameter value be y2, and the timestamp of the missing point be x. The interpolation slope is calculated as k=(y2-y1) / (x2-x1). Then, the parameter value of the missing point is calculated based on the slope: the parameter value of the missing point is y=y1+k×(x-x1). The calculated missing point parameter value is then filled into the corresponding position. This process completes the interpolation of all missing data points, ultimately generating the processed environmental parameter data.
[0039] Step 102: Integrate the registered and aligned image and signal data, along with the processed environmental parameter data, to generate a standardized data frame set with a unified spatiotemporal reference. Specifically, this includes: First, extracting unified geospatial coordinate system parameters and unified timestamp reference parameters from the registered and aligned image and signal data and the processed environmental parameter data after standardization. The unified geospatial coordinate system parameters include core information such as coordinate system type, longitude range, latitude range, and elevation reference. The unified timestamp reference parameters include core information such as time reference type, time accuracy, and time calibration basis. This set of spatiotemporal parameters serves as the sole core reference for this dataset integration. Next, the two types of data are finely organized according to a spatiotemporal dual dimension. Using the unified timestamp as the vertical organization dimension, all data are hierarchically divided according to the chronological order of the timestamps. Using the unified geospatial coordinates as the horizontal organization dimension, all data are regionally divided according to the latitude and longitude intervals of the geospatial coordinates. Finally, using a precise spatiotemporal coordinate matching algorithm, data with the same timestamp and spatial coordinates falling within the same geospatial latitude and longitude range are matched. Image data, electromagnetic signal data, and environmental parameter data within the interval are precisely correlated and matched, enabling different types of monitoring data to form an organic whole that corresponds and is interconnected in the spatiotemporal dimension. Each data packet that has completed spatiotemporal correlation matching forms an independent spatiotemporal data frame unit containing multiple types of monitoring data. Each data frame unit carries a unique spatiotemporal identifier, including the corresponding timestamp and geospatial coordinate interval information. Finally, all data frame units that have completed spatiotemporal correlation matching are sorted in the first order according to the order of timestamps. Within the same timestamp, they are then sorted in the second order according to the distribution pattern of geospatial coordinates from longitude to latitude from small to large. All data frame units after the two-level sorting are integrated and constructed in an integrated manner, and all data frame units are combined into a complete dataset in sequence, ultimately forming a standardized data frame set with a unified spatiotemporal reference. The spatiotemporal reference of all data units in this standardized data frame set is highly consistent, the data format is standardized and unified, and all data has completed preprocessing and quality correction, eliminating the need for additional spatiotemporal registration, format conversion, and quality processing operations.
[0040] This embodiment relies on onboard processors or near-Earth edge computing nodes to receive and parse multi-source raw data, moving the data processing flow to the edge to avoid bandwidth consumption on the space-to-ground link caused by direct downlink of raw data. By classifying and storing raw data into independent buffers according to data source identifiers, it achieves orderly organization of multi-source heterogeneous data and avoids mixed storage of data from different sensors. It completes spatial coordinate system registration and timestamp synchronization alignment for image and electromagnetic signal data, eliminating spatiotemporal heterogeneity problems caused by different acquisition bases of different sensors, and improving the accuracy and effectiveness of subsequent data fusion. It performs unified conversion of measurement units for environmental parameter time-series data, eliminating differences in measurement expressions of different parameters. At the same time, it detects and removes outliers and fills in missing points using linear interpolation, improving the integrity and accuracy of environmental parameter data and reducing the interference of low-quality data on subsequent analysis results.
[0041] In a preferred embodiment of the present invention, step 2 includes:
[0042] Step 200: For multi-source data points located within the same geographical area in the standardized data frame set, analyze the statistical continuity of their spatial distribution and calculate the dynamic radius parameter used for spatial interpolation; for time series data belonging to the same observation target in the standardized data frame set, analyze the rate of change and periodicity characteristics of the time series data over time and calculate the dynamic window size parameter used for time alignment. Specifically, this includes: firstly, dividing the geographical area covered by the standardized data frame set into several regular grids at fixed latitude and longitude intervals (0.01 degrees × 0.01 degrees), with each grid being an independent location within the same geographical area. The data is analyzed by associating data source identifiers with spatial coordinates to identify all time-series data for the same fixed geographic entity and geographic area (such as building entities, vegetation cover areas, water bodies, and road sections), defining them as time-series data for the same observation target. For multi-source data points within each geographic region, the statistical continuity of their spatial distribution is analyzed. First, the total number of data points within the region is counted, and the three-dimensional spatial distance between every two adjacent data points is calculated. The sum of all adjacent distances is then divided by the total number of adjacent distances to obtain the average adjacent spacing of data points within the region. The initial value of the dynamic radius parameter is set to 1 / 3 of the average adjacent spacing. The 0.5x average distance is determined based on the optimal coverage criterion for local statistical interpolation. This 1.5x average distance can cover over 90% of the correlation range of adjacent data points within the region, balancing the sufficiency of interpolation reference samples with computational efficiency. It avoids situations where the radius is too small, resulting in no effective reference samples for interpolation points, or too large, introducing irrelevant data and reducing interpolation accuracy. Furthermore, data point density is adjusted. The preset threshold for data point density is 10 points / square kilometer. When the data point density (number of data points in the region divided by the region area) is lower than this preset threshold, the initial value is multiplied by 1.2x to enlarge the radius. This value is based on the correlation between data points in low-density areas. The 1.2x magnification ratio ensures that the interpolation range covers the entire geographic area without significantly reducing interpolation accuracy, without generating any blank interpolation points, thus guaranteeing the integrity of spatial interpolation. When the data point density exceeds the preset threshold, the initial value is multiplied by 0.8 to reduce the radius. This value is based on the fact that the data points in high-density areas have small intervals. The 0.8x reduction ratio allows the interpolation points to refer only to nearby highly correlated sample points, improving the spatial accuracy of interpolation and meeting the accuracy requirements of local kriging interpolation. The final value obtained is the dynamic radius parameter for the spatial interpolation of multi-source data in this region, with each geographic region corresponding to an independent dynamic radius.
[0043] For time series data of the same observation target, the rate of change and periodicity characteristics over time are analyzed. The rate of change is calculated by taking the difference between the observed data values of two adjacent timestamps, dividing it by the time interval between the two timestamps to obtain the instantaneous rate of change for each time interval, and then averaging all instantaneous rates of change to obtain the average rate of change. The preset threshold for the average rate of change is 0.05 units / second, where the unit is the basic unit of measurement for various environmental monitoring parameters and image feature parameters. Periodicity characteristics are statistically analyzed using the sliding window method. An initial sliding window is set (the duration is 3 times the data sampling interval), and the fluctuation amplitude of the data within the window is statistically analyzed. The window is gradually expanded until the fluctuation amplitude tends to stabilize. The duration of the window at this point is the data periodicity duration. The dynamic window size parameter is calculated according to the following rules, using the data sampling interval as the basic unit. When the average rate of change is greater than the preset threshold of 0.05 units / second, the window size is 2 times the sampling interval to adapt to scenarios with rapid data changes. When the average rate of change is less than or equal to the preset threshold of 0.05 units / second, the window size is the result of dividing the periodicity duration by the sampling interval to ensure coverage of a complete cycle and achieve stable alignment of the time series.
[0044] Step 201: Based on the spatiotemporal range defined by the dynamic radius parameter and the dynamic window size parameter, perform spatial interpolation and temporal series alignment operations on the standardized data frames from different sensors to generate a spatiotemporally continuous and consistent multi-source fusion data field. Specifically, this includes: using the dynamic radius of each geographic region as the range, performing spatial interpolation on the multi-source discrete data points within that region, employing the Kriging interpolation method, with the dynamic radius as the interpolation influence range, using the discrete data points from different sensors within the region as known samples, calculating the spatial covariance of the sample points, and solving for the interpolation weights through the covariance matrix. The interpolation weights range from 0.05 to 0.9, with the weight of a single sample point not exceeding 0.9 to avoid excessive influence of a single sample point on the interpolation result. The weight calculation method is to multiply the inverse matrix of the covariance matrix by the covariance vector from the sample point to the interpolation point, then multiply the known sample data by the corresponding weight and sum them to obtain the value of the interpolation point; traversing the entire coverage area at a fixed grid spacing (consistent with the geographic region grid in step 200), supplementing the values of all blank grid points, thus transforming the discrete data into a spatially continuous data layer. Using the dynamic window size as a benchmark, alignment processing is performed on the time series data from different sensors. The multi-source time series data of the same observation target is divided into several continuous time windows according to the dynamic window size, and each window contains several sampling points. For each time window, the average timestamp of the data from different sensors within the window is calculated. This average is used as the unified timestamp of the window. Then, the arithmetic mean of all data values of each sensor within the window is taken to obtain the aligned data of that sensor under the unified timestamp. If there is missing data for a certain sensor within the window, linear interpolation is used to fill it (refer to the missing value filling logic in step 101) to ensure that all sensors have corresponding data within each time window. The data layers of each sensor that have completed spatial interpolation are associated with the time-aligned time series data according to the spatiotemporal dimension. Using the unified timestamp as the vertical dimension and the spatial grid as the horizontal dimension, the interpolated data and aligned data of different sensors at the same spatiotemporal location are superimposed and integrated. Each spatiotemporal grid cell contains the fused data of all sensors, and finally a spatiotemporally continuous multi-source fused data field with consistent data benchmarks is formed.
[0045] Step 202 involves inputting the multi-source fusion data field into a pre-trained lightweight ensemble learning model based on a gradient boosting decision tree architecture. This model performs nonlinear mapping and feature-level fusion of observation data from different sensor sources, and predicts and compensates for system biases caused by differences in observation conditions or sensors, generating fused feature vectors. Specifically, the lightweight ensemble learning model uses a gradient boosting decision tree (GBDT) architecture as its core. It is customized to meet the computational and memory constraints of onboard embedded processors or near-Earth edge computing nodes. The specific construction process involves setting the model to consist of 50 weak decision trees, with each tree's depth controlled within 8 layers (limiting the depth reduces the number of model parameters and onboard computing resources), and each decision tree having no more than 32 leaf nodes (reducing node splitting computation and improving inference speed). The mean squared error loss function is used to adapt to numerical feature fusion scenarios of multi-source sensor data. Data is collected covering different observation conditions (clear / sunny). Historical multi-source sensor data from various regions (urban, suburban, water, mountainous areas, etc.), including clear skies, cloudy skies, rainfall, atmospheric disturbances, etc., and different sensor types (optical, synthetic aperture radar, electromagnetic signal detection, environmental parameter sensors), were divided into training, validation, and test sets in a 7:2:1 ratio. All data underwent standardization processing in steps 100 and 101 to ensure a consistent data baseline. Mini-batch gradient descent was used for training, with a batch size of 64 (adapting to the memory capacity of edge nodes). The initial learning rate was 0.1, which dynamically decreased with each training epoch (multiplying by 0.9 every 10 epochs), for a total of 80 training epochs. During training, the fusion error of the validation set (mean squared error between the fusion result and the true value) was used as the monitoring indicator. Training was stopped when the validation set error no longer decreased after 5 consecutive epochs to avoid overfitting. The resulting lightweight model had fewer than 10MB of parameters and a single inference time of no more than 50ms, making it fully compatible with the hardware computing environment of satellites or edge nodes.
[0046] A lightweight model is pre-trained using multi-source fusion data fields. The model assigns independent feature branches to data from different sensor sources. Each branch designs a unique nonlinear mapping rule based on the data characteristics of its corresponding sensor. Specifically, the optical sensor feature branch focuses on features such as image texture and grayscale values, transforming the original image data into 128-dimensional features through decision tree node splitting rules (with information gain maximization as the splitting criterion). The synthetic aperture radar sensor feature branch focuses on features such as echo intensity and scattering coefficient, transforming them into 64-dimensional features. The electromagnetic signal detection sensor feature branch focuses on features such as signal frequency and amplitude, transforming them into 64-dimensional features. The environmental parameter sensor feature branch focuses on time-series features such as temperature and humidity, transforming them into 32-dimensional high-dimensional features. The dimensions of all high-dimensional features from all branches are uniformly standardized to 128 dimensions to ensure consistent fusion dimensions. A weighted summation method is used to fuse the high-dimensional features from multiple branches. Weighted summation is chosen because different sensors contribute differently to the multi-source fusion result; it assigns higher weights to sensor features with higher contributions, improving the accuracy of the fusion result. The weights of various sensor features are determined based on the actual needs of low-orbit satellite IoT multi-source sensing scenarios. The weight of optical sensor features is set to 0.35, as optical sensors can provide high-resolution... Surface visual features are the core basis for scene recognition and target localization, with the highest feature recognition and contribution, thus assigned the highest weight. The synthetic aperture radar (SAR) sensor feature weight is 0.3. SAR sensors possess all-weather, all-time observation capabilities, compensating for the weather-dependent limitations of optical sensors. They exhibit strong feature stability and complementarity, making their contribution second only to optical sensors. The electromagnetic signal detection sensor feature weight is 0.15. SAR sensors focus on collecting electromagnetic features and are mainly used for identifying specific targets (such as communication base stations and power facilities). Their applicable scenarios are relatively specific, and their contribution is moderate. Environmental parameter sensors (temperature, humidity...) The feature weights (air pressure, rainfall, etc.) are set to 0.2. Environmental parameter sensors provide basic environmental background features of the scene, which can help improve the interpretation accuracy of optical and radar features. They are an important supplement to the fusion results and their contribution is higher than that of electromagnetic signal sensors. The feature weights of all sensors are summed to 1 to meet the normalization requirements of weight allocation. The feature importance is obtained by summing the loss function reduction caused by each feature in all decision trees. The weight of a certain type of sensor feature is equal to its feature importance divided by the sum of the importance of all features. Then, the 128-dimensional high-dimensional features mapped from each branch are multiplied by the corresponding weights, and the summation is performed dimension by dimension to obtain the 128-dimensional preliminary fusion features.
[0047] The model calls upon a pre-trained bias prediction module to accurately predict and compensate for system biases caused by observation conditions (such as cloud cover and atmospheric disturbance) or sensor differences (such as hardware accuracy and calibration deviation). Specifically, the bias prediction module, based on the correlation patterns learned during training between different observation conditions, sensor types, and system biases, first identifies the observation condition labels (such as cloud cover degree and atmospheric disturbance level) and sensor type labels corresponding to the current data. It then calls upon historical datasets of similar scenarios (same observation conditions + same sensor type) and calculates the difference between the actual values in these datasets and the model's initial fusion results. The system bias prediction value for the current scene is obtained by summing all the differences and dividing by the total number of differences. Then, the system bias prediction value is calculated by adding the value of each dimension of the 128-dimensional preliminary fused features to the system bias prediction value. The correction logic is as follows: if the bias prediction value is positive, it indicates that the overall value of the preliminary fused features is too low, and the value is increased after adding; if the bias prediction value is negative, it indicates that the overall value of the preliminary fused features is too high, and the value is decreased after adding. Finally, a unified (128-dimensional) fused feature vector with optimized accuracy is output, which eliminates the system error caused by differences in observation conditions and sensors.
[0048] In this embodiment, the dynamic radius and dynamic window size parameters are calculated based on the spatial continuity and temporal variation of the data itself. This allows for adaptation to different geographical data densities and temporal characteristics of different observation targets, avoiding interpolation distortion or time alignment deviations caused by fixed parameters and improving the adaptability of multi-source data fusion. By spatially interpolating discrete data and aligning the time window with the time series of different sensors, the problem of inconsistent spatiotemporal references for multi-source data is solved. Data regularization is completed in advance at the edge, eliminating the need for the ground system to consume additional resources to process heterogeneous data and shortening the subsequent analysis process. The optimized gradient boosting decision tree model is small in size and has low computational load, making it suitable for the limited computing and storage resources of onboard or edge nodes. At the same time, through feature-level fusion and bias compensation, errors caused by observation conditions and sensor differences are effectively reduced, improving the accuracy and reliability of fused data. Deep processing of multi-source data is completed at the edge, generating a high-quality fused feature vector. The ground data center can directly perform analysis based on this vector without repeating the fusion and bias correction operations, improving the timeliness of data processing and supporting the near real-time monitoring needs of dynamic targets.
[0049] In a preferred embodiment of the present invention, step 3 includes:
[0050] Step 300: For each data unit in the fused feature vector, sequentially perform checks on the continuity of data frame counts, the integrity of observation time segment coverage, and the reliability of sensor-reported status to obtain the results of the data frame count continuity check, the integrity of observation time segment coverage, and the reliability of sensor-reported status assessment. Specifically, this includes: firstly, extracting the original sensing data frame sequence corresponding to a single data unit, determining the preset total observation duration and fixed data acquisition interval for the observation task to which the data unit belongs, where the preset total observation duration is 600 seconds, which is the typical duration for a low-orbit satellite to complete a complete observation of the target area in a single pass, and the fixed data acquisition interval is preset according to each sensor type (consistent with step 100), and the data frame count is calculated. The theoretical data frame count of the unit is calculated using the formula: Theoretical data frame count = Preset total observation time ÷ Fixed data acquisition interval. Then, the actual valid data frame count corresponding to the data unit is counted. Actual valid data frames are those that have undergone standardization and fusion processing and are free of format errors and missing values. Based on the theoretical and actual counts, a continuity coefficient is calculated using the formula: Continuity coefficient = Actual valid data frame count ÷ Theoretical data frame count. The continuity coefficient ranges from 0 to 1; the closer the value is to 1, the better the continuity of the data frame count. This continuity coefficient is the verification result of the data frame count continuity. After completing this verification for a single data unit, all data units in the fused feature vector are traversed sequentially according to the same logic, and the verification is completed one by one, with the results recorded.
[0051] For a single data unit that has completed the frame count continuity verification, firstly, its corresponding actual observation time segment is extracted, and the preset observation time range of the observation task to which the data unit belongs is determined (corresponding to the preset total observation duration of 600 seconds, which is a continuous time interval). The initial time coverage coefficient is calculated using the formula: Initial time coverage coefficient = Total duration of actual observation time segment ÷ Total duration of preset observation time range. Subsequently, fragment detection is performed on the actual observation time segment. If there is an interruption in the actual observation time segment within the same observation time range with a time interval greater than twice the acquisition interval, it is determined as a time fragment. For each time a time fragment is detected, a deduction operation of 0.1 is performed on the initial time coverage coefficient. The minimum value of the time coverage coefficient after deduction is 0. The final corrected time coverage coefficient is the judgment result of the observation time segment coverage integrity. The coefficient ranges from 0 to 1. The closer the value is to 1, the better the time coverage integrity. The same logic is used to traverse all data units in the fused feature vector to complete the full coverage integrity judgment.
[0052] For a single data unit that has completed the time coverage integrity assessment, the full reporting status information of its data source sensor is extracted. A reliability score is calculated from three dimensions: data packet loss rate, data value fluctuation amplitude, and sensor calibration status. The overall reliability score of the sensor reporting status is then obtained through a weighted summation. The data packet loss rate score is calculated as 1 - the actual data packet loss rate of the sensor, where the packet loss rate is the ratio of the number of lost data frames during sensor reporting to the total number of reported data frames, with a score ranging from 0 to 1. The data value fluctuation amplitude score is calculated by setting a preset fluctuation threshold based on the sensor type. Values (optical sensor grayscale value fluctuation threshold is 5, synthetic aperture radar sensor echo intensity fluctuation threshold is 8, electromagnetic signal detection sensor amplitude fluctuation threshold is 3, environmental parameter sensors are assigned values according to parameter type, i.e., temperature 2℃, humidity 5%RH, air pressure 10hPa, rainfall 2mm / h). These thresholds are the maximum allowable fluctuation range calibrated by the manufacturer for each sensor. If the actual fluctuation range of the data value is ≤ the corresponding preset fluctuation threshold, the score is 1; if the actual fluctuation range is > the preset fluctuation threshold, the score is 1 - (actual fluctuation range - preset fluctuation threshold) ÷ preset fluctuation threshold, with the highest score being... The lowest value is 0. The sensor calibration status score is calculated as follows: if the sensor completes formal calibration during the data acquisition period and the calibration status is qualified, the score is 1; if calibration is not completed or the calibration status is unqualified, the score is 0.5. Weights are assigned based on the degree of influence of each dimension on the sensor's reported status. The data packet loss rate is weighted at 0.4 because packet loss directly leads to data loss and has the most significant impact on the integrity of the reported data; the data value fluctuation amplitude is weighted at 0.3 because fluctuation amplitude exceeding the threshold will reduce data accuracy, but its impact is less than that of packet loss; the sensor calibration status is weighted at 0.3. Since the calibration status determines the basic reliability of the data, and its impact on credibility is comparable to that of fluctuation amplitude, the overall credibility score is calculated as follows: Overall Credibility Score = Data Loss Rate Score × 0.4 + Data Value Fluctuation Amplitude Score × 0.3 + Sensor Calibration Status Score × 0.3. The score ranges from 0 to 1. The closer the value is to 1, the higher the credibility of the sensor's reported status. This score is the evaluation result of the credibility of the sensor's reported status. Following the same logic, all data units in the fusion feature vector are traversed and merged. Finally, each data unit is matched with the corresponding three quantitative results: verification, judgment, and evaluation.
[0053] Step 301: For each data unit in the fused feature vector, based on the data frame count continuity verification result, observation time segment coverage integrity judgment result, and sensor reporting status reliability assessment result corresponding to the data unit, calculate the comprehensive usability score of the data unit, and filter out data units with a comprehensive usability score lower than the preset quality threshold according to the preset quality threshold; add a corresponding quality level label to each data unit retained after filtering, the quality level label is obtained by mapping the comprehensive usability score of the data unit, forming a preferred feature dataset with quality labels, specifically including: firstly, based on the logarithm of the three quantization results... The weights are assigned according to the degree of impact on quality. The continuity coefficient of data frame counts and the coverage coefficient of observation time segments are both weighted at 0.35. This is because both jointly determine the spatiotemporal integrity of the data, which is a core dimension for assessing data usability, and their impact is comparable. The reliability score of sensor reported status is weighted at 0.3. This is because the reliability of the data source is a fundamental guarantee, but its impact is less than that of spatiotemporal integrity. This weight allocation takes into account the continuity, integrity, and source reliability of the data, achieving a comprehensive assessment of multi-dimensional quality. Subsequently, a comprehensive usability score is calculated for each individual data unit using the formula: Comprehensive Usability Score = Continuity Coefficient × 0.35 +Time Coverage Factor × 0.35 + Sensor Reported Status Reliability Score × 0.3, where the score ranges from 0 to 1. A value closer to 1 indicates higher overall usability of the data unit. The same logic is used to calculate the comprehensive score for all data units in the fused feature vector. A preset data quality threshold of 0.6 is set; this threshold is the lowest quality score verified by a large amount of historical observation data to ensure the effectiveness of subsequent data analysis. All data units in the fused feature vector are iterated through, and low-quality data units with a comprehensive usability score below 0.6 are directly filtered out from the dataset, retaining only valid data units with a score ≥ 0.6. Based on the comprehensive usability score... The scope is defined by establishing a three-level quality level mapping rule to achieve accurate conversion of scores to quality level labels. Specifically, the mapping rule is as follows: a comprehensive usability score in the range of 0.8 to 1 is mapped to a quality level label of "Excellent"; a score in the range of 0.6 to 0.8 is mapped to a quality level label of "Good"; and a score equal to 0.6 is mapped to a quality level label of "Pass". Corresponding quality level labels are attached to all valid data units that have been filtered out and retained. The labels are associated with and stored with the core information of the data units, such as spatial coordinates, timestamps, and feature values. All valid data units with quality level labels are organized in an orderly manner according to the spatiotemporal dimension to form a preferred feature dataset with quality labels.
[0054] Step 302: Based on the pre-set spatial extent index of geographic elements, retrieve and extract all feature data whose spatial coordinates fall within the boundaries of urban built-up areas, vegetation cover areas, and water areas from the preferred feature dataset with quality labels. This constitutes a subset of feature data corresponding to the urban built-up areas, vegetation cover areas, and water areas. Specifically, the pre-set spatial extent index of geographic elements is a vector spatial index registered with the unified geographic spatial coordinate system (WGS84 geodetic coordinate system). This index pre-inputs the complete three-dimensional coordinate set of geographic boundaries for the three types of geographic elements: urban built-up areas, vegetation cover areas, and water areas. Each type of geographic element corresponds to... For each unique geographic feature identifier, boundary coordinates are standardized and stored according to the dimensions of longitude, latitude, and elevation. All coordinates use the same geospatial coordinate system as the previously standardized data frame set and the fused feature vector, eliminating the need for additional coordinate transformation. Furthermore, this index constructs a gridded index for the boundary coordinates of various geographic features, improving the efficiency of subsequent spatial retrieval and adapting to the limited computing resources at the edge. Each data unit in the optimized feature dataset with quality labels is processed individually. First, the standardized spatial coordinates (longitude, latitude, and elevation) carried by the data unit are extracted. These spatial coordinates are then indexed against the pre-defined spatial range of geographic features. The spatial inclusion relationships of the boundary coordinates of the three types of geographic elements are determined sequentially. A closed geographic area is defined as the two-dimensional planar region enclosed by the boundary coordinates of each type of geographic element, superimposed with the corresponding elevation intervals (0 to 500 meters for urban built-up areas, 0 to 3000 meters for vegetated areas, and -50 to 100 meters for water areas), forming a three-dimensional closed space. The determination rule is that if the spatial coordinates of a data unit fall within both the two-dimensional planar boundary of a certain type of geographic element and the corresponding elevation interval, then the data unit is considered to match that type of geographic element; if the coordinates of a data unit do not fall within the three-dimensional closed space of any type of geographic element, it is temporarily excluded from this determination. Extraction scope; classify and aggregate the successfully matched data units according to geographic feature type, organize all data units matching urban built-up areas in an orderly manner to form a subset of urban built-up area feature data; organize all data units matching vegetation cover areas in an orderly manner to form a subset of vegetation cover area feature data; organize all data units matching water area in an orderly manner to form a subset of water area feature data. Each feature data subset retains all information such as the quality level label, spatiotemporal coordinates, and feature values of the original data unit, and is arranged in an orderly manner according to the time stamp order and spatial coordinate distribution pattern, thus completing the generation of three types of geographic feature data subsets.
[0055] This embodiment completes the quantitative verification and evaluation of fused feature vectors from three dimensions: frame count, time coverage, and sensor reporting status. This multi-dimensional and comprehensive approach reflects the integrity and reliability of data units, avoiding quality judgment bias caused by single-dimensional evaluation. Weighted calculation achieves a comprehensive usability score for data units, with weight allocation considering data continuity, integrity, and source reliability. Combined with preset quality thresholds to filter low-quality data, it effectively improves the overall quality of the selected feature dataset, reduces interference from low-quality and invalid data on subsequent data analysis results, and enhances the effectiveness of data application. Data retrieval and extraction are completed based on a pre-built standardized geographic element spatial range index, eliminating the need for additional coordinate transformation and improving the efficiency of spatial retrieval and data extraction, adapting to the limited computing resource constraints of satellite or edge nodes. Dedicated feature data subsets are extracted according to urban built-up areas, vegetation coverage areas, and water areas, achieving scenario-based and refined classification of selected feature data, meeting the differentiated monitoring and analysis needs of different geographical regions. The feature data subsets completely retain the quality markers and core information of the original data and are organized in an orderly manner according to the spatiotemporal dimension, eliminating the need for subsequent data processing and quality judgment, simplifying the subsequent data processing flow, and improving the efficiency of overall edge preprocessing.
[0056] In a preferred embodiment of the present invention, step 4 includes:
[0057] Step 400: For each set of temporally continuous sequential image data corresponding to the feature data subsets of urban built-up areas, vegetation coverage areas, and water areas, extract apparent feature points between consecutive image frames; compare and associate the extracted apparent feature points of two consecutive image frames to obtain a set of successfully matched feature point pairs. Specifically, this includes: extracting temporally continuous sequential image data from the three types of feature data subsets according to geographic feature type, requiring that the time interval between adjacent image frames in the sequence is consistent with the fixed acquisition interval of the sensor, and that all image data have excellent or good quality labels to ensure the reliability of feature point extraction; for the selected consecutive image frames, use the SIFT algorithm to extract apparent feature points, setting during the extraction process... The response threshold for fixed feature points is set to 0.03. This value has been verified through extensive image testing and can effectively filter low-contrast noise points while retaining real and effective weak texture feature points, balancing the number and reliability of feature points. The edge threshold is set to 10. This threshold can accurately distinguish between edge and non-edge regions, excluding feature points that are easily disturbed and have poor stability at the edges, thus avoiding such points from affecting subsequent matching accuracy. Only feature points with response values higher than the threshold and not located in edge regions are retained, excluding noise interference points. The extracted apparent feature points must contain three core pieces of information: pixel coordinates, grayscale gradient, and feature descriptor. The feature descriptor is generated by performing gradient statistics on the 16×16 pixel region around the feature point, ensuring the uniqueness and recognizability of the feature point.
[0058] The apparent feature points extracted from the two consecutive frames are compared pairwise. The Euclidean distance of the feature descriptors is used as the matching criterion. The Euclidean distance of the descriptors of a single feature point in the previous frame image and all feature points in the subsequent frame image is calculated. The two subsequent frame feature points with the smallest distance are selected as candidate matching points. If the ratio of the smallest distance to the second smallest distance is less than 0.6, it is determined that the feature point is successfully matched with the subsequent frame feature point corresponding to the smallest distance, forming a set of feature point pairs. Here, 0.6 is the confidence threshold for feature point matching. When the ratio is less than 0.6, it can be ensured that the mismatch rate of the successfully matched feature point pairs is less than 5%. This avoids the introduction of a large number of mismatched points due to an excessively large threshold, and also prevents the loss of effective matching points due to an excessively small threshold. Following this logic, all feature points in the previous frame are traversed. After completing the pairwise comparison, duplicate matching and mismatched feature point pairs are removed. Finally, a set of successfully matched and highly reliable feature point pairs is obtained.
[0059] Step 401: Based on the pixel coordinate difference between the successfully matched feature point pairs in the two image frames, calculate the displacement vector of the geographic entity on the image plane; combining the sensor pose parameters corresponding to the acquisition of the serialized image data, map the displacement vector from the image plane coordinate system to the geographic space coordinate system to obtain a location point of the geographic entity in geographic space. Specifically, this includes: for each successfully matched feature point pair, reading the pixel coordinates of the feature point in the previous frame and the pixel coordinates of the corresponding feature point in the subsequent frame, calculating the pixel coordinate difference between the two in the image plane coordinate system, i.e., the difference in the horizontal coordinate is the difference between the horizontal coordinate of the subsequent frame and the horizontal coordinate of the previous frame, and the difference in the vertical coordinate is... The difference is calculated by subtracting the ordinate of the previous frame's pixels from the ordinate of the subsequent frame's pixels. Using the difference in horizontal coordinates as the x-component and the difference in vertical coordinates as the y-component, a two-dimensional displacement vector of the geographic entity on the image plane is formed. The vector direction represents the orientation of the displacement, and the vector magnitude represents the pixel distance of the displacement. The sensor pose parameters corresponding to the acquisition of this set of serialized image data are obtained, including the sensor's imaging focal length, principal point coordinates, attitude angles (azimuth, pitch, roll), and position coordinates. All parameters are extracted from the sensor acquisition log and have been registered with a unified geospatial coordinate system. Based on the sensor's imaging focal length and principal point coordinates, the image plane is... The surface displacement vector is converted into an image space angular offset. The angular offset is calculated by dividing the offset pixel value by the imaging focal length to obtain the corresponding angular change values in the horizontal and vertical directions. Combining the sensor attitude angle and position coordinates, the image space angular offset is mapped to the geospatial coordinate system. The specific implementation process is as follows: First, a rotation matrix is calculated based on the sensor's attitude angles (azimuth α, pitch β, roll γ). The rotation matrix is obtained by multiplying the rotation matrices around the three coordinate axes of the geospatial coordinate system. That is, the rotation matrix for rotating the azimuth α around the Z-axis (sky axis) is: the first row [cosα, -sinα, 0], the second row... The first row is [sinα,cosα,0], the second row is [0,0,1]; the rotation matrix for rotating the pitch angle β around the Y-axis (north axis) is: the first row is [cosβ,0,sinβ], the second row is [0,1,0], the third row is [-sinβ,0,cosβ]; the rotation matrix for rotating the roll angle γ around the X-axis (east axis) is: the first row is [1,0,0], the second row is [0,cosγ,-sinγ], the third row is [0,sinγ,cosγ]. Multiplying these three basic rotation matrices in the order of roll, then pitch, and finally azimuth gives the overall rotation matrix corresponding to the sensor attitude.
[0060] The image space angular offsets (horizontal angular offset Δθ, vertical angular offset Δφ) are converted into image space unit direction vectors [Δθ, Δφ, 1]. This vector is then multiplied by the overall rotation matrix to complete the coordinate system transformation from image space to geographic space, yielding the geographic space direction vector. Starting from the sensor's position coordinates (longitude L0, latitude B0, elevation H0), and combined with the distance D from the sensor to the geographic entity (obtained from the sensor ranging parameters), the sensor position coordinate correction is calculated. The correction is equal to the change in the geographic space direction vector per unit distance. Each component is multiplied by the ranging distance D, i.e., the longitude component of the correction amount = longitude component of the geospatial direction vector × D, the latitude component of the correction amount = latitude component of the geospatial direction vector × D, and the elevation component of the correction amount = elevation component of the geospatial direction vector × D; finally, the geospatial coordinates of the geographic entity at the time of image acquisition in the next frame are: longitude of the location point = L0 + longitude component of the correction amount, latitude of the location point = B0 + latitude component of the correction amount, and elevation of the location point = H0 + elevation component of the correction amount, thus completing the mapping from the image plane coordinate system to the geospatial coordinate system.
[0061] Step 402 involves repeatedly performing feature point extraction, matching, displacement calculation, and coordinate mapping for each pair of consecutive image frames in the serialized image data. This accumulates over time to obtain a series of geospatial location points, which together form the motion trajectory sequence of the geographic entity. Specifically, for each set of serialized image data, each pair of consecutive image frames is selected sequentially over time, and the apparent feature point extraction and comparison operations of step 400, as well as the displacement vector calculation and geospatial coordinate mapping operations of step 401, are repeatedly performed. Each pair of consecutive frames processed yields a corresponding geospatial location point for the geographic entity. All obtained geospatial location points are arranged in order according to the chronological order of the image frames. Each location point is associated with a corresponding acquisition timestamp and quality level marker. All location points of the same geographic entity within the entire serialized image data acquisition cycle are accumulated over time to form a series of geospatial location points. This series of location points is then combined in an orderly manner to form the motion trajectory sequence of the geographic entity. Each motion trajectory sequence corresponds to a unique geographic entity identifier and geographic feature type.
[0062] Step 403: For each motion trajectory in the motion trajectory sequence, sequentially read all the geospatial location points contained in that trajectory; for each read motion trajectory, calculate the three-dimensional Euclidean distance between every two adjacent geospatial location points in that trajectory. Specifically, for motion trajectory sequences corresponding to subsets of urban built-up areas, vegetation cover areas, and water area feature data, classify them according to geographic element type, and sequentially read all the geospatial location points contained in each motion trajectory. Each location point contains three coordinate parameters: longitude, latitude, and elevation, and the coordinate parameters are based on a unified geospatial coordinate system; for a single motion trajectory, starting from the first location point, sequentially select every two adjacent geospatial location points and calculate the three-dimensional Euclidean distance between them. The calculation process first converts the longitude and latitude coordinates into plane rectangular coordinates. The conversion formula is: the plane horizontal coordinate equals the longitude value multiplied by 111319.9, and the plane vertical coordinate equals the latitude value. Multiply by 111319.9 by the cosine of latitude, where 111319.9 is the conversion factor between longitude, latitude, and meters. Physically, it represents the actual ground distance (in meters) corresponding to 1 degree of longitude or latitude on the Earth's equator. This value is calculated by dividing the Earth's equatorial circumference (approximately 40076 kilometers) by 360 degrees. It is used to convert latitude and longitude coordinates into Cartesian coordinates in meters, eliminating the difference between latitude and longitude angle units and distance units. Multiplying by the cosine of latitude corrects the problem that the ground distance corresponding to the same longitude difference is shorter at higher latitudes, ensuring the accuracy of the plane coordinate conversion. Then, substitute it into the three-dimensional Euclidean distance formula to calculate, that is, the distance equals the square root of (the square of the difference in the plane abscissa of two adjacent points + the square of the difference in the plane ordinate of two adjacent points + the square of the difference in elevation of two adjacent points). Following this logic, traverse all adjacent positions of a single trajectory, calculate the corresponding three-dimensional Euclidean distance one by one, and simultaneously record the time interval corresponding to each distance segment, and associate it with the trajectory sequence.
[0063] Step 404 involves summing the 3D Euclidean distances between all adjacent geographic locations along a trajectory. The summation result is used as the approximate total length of the spatial curve of the trajectory, and this approximate total length is recorded as a geometric feature parameter of the trajectory. Specifically, for a single trajectory, the 3D Euclidean distances between all adjacent geographic locations calculated in step 403 are collected, and these distance values are summed sequentially. The summation result is the approximate total length of the spatial curve of the trajectory, measured in meters and rounded to two decimal places. The calculated approximate total length of the spatial curve is used as the core geometric feature parameter of the trajectory and stored in association with information such as the geographic entity identifier, geographic feature type, time span, and number of locations in the trajectory sequence. Following the same logic, all trajectory sequences are traversed, and the total length calculation and parameter recording are completed one by one, ultimately forming a trajectory feature dataset containing the geometric feature parameters of each trajectory, providing data support for subsequent trajectory analysis and target recognition.
[0064] This embodiment employs the SIFT algorithm to extract apparent feature points and combines it with descriptor comparison and matching, along with strict matching rules, to effectively reduce false and duplicate matching, improve feature point pair matching accuracy, and provide a reliable foundation for displacement vector calculation and trajectory construction. It fuses sensor pose parameters to complete coordinate system mapping, accurately converting image planar displacement to geospatial coordinates, solving the heterogeneity problem between image coordinates and geographic coordinates, ensuring the spatial accuracy of geographic entity location points and motion trajectories, and adapting to the spatial positioning requirements of low-Earth orbit satellite multi-source sensing. Finally, it accumulates location points in chronological order to form a motion trajectory sequence, completely reconstructing the movement of geographic entities. The system tracks the motion process and associates it with quality level markers and timestamps, providing multi-dimensional references for determining the validity of the trajectory and analyzing its movement patterns. Based on the calculation and summation of Euclidean distance in three-dimensional space, it accurately obtains the approximate total length of the trajectory spatial curve, quantifies the geometric characteristics of the trajectory, and provides a quantitative basis for assessing the movement amplitude and activity range of geographic entities within different geographic element areas. The system categorizes and processes trajectory data according to geographic element types, adapting to differentiated scenarios such as urban built-up areas, vegetation-covered areas, and water areas. The trajectory feature parameters can directly support target motion analysis in various scenarios, improving the targeting of data processing and application, and adapting to the needs of edge-end special monitoring.
[0065] In a preferred embodiment of the present invention, step 5 includes:
[0066] Step 500: Determine the geographic spatial range covered by the feature data subsets corresponding to urban built-up areas, vegetation-covered areas, and water areas, as well as the geographic spatial range traversed by the movement trajectory sequence. Union these two geographic spatial ranges to obtain the target spatial domain. Specifically, this includes: First, determining two types of geographic spatial ranges: one type is the geographic spatial range covered by the three feature data subsets of urban built-up areas, vegetation-covered areas, and water areas. Then, by extracting the spatial coordinates of all data units in each feature data subset, statistically obtain the maximum and minimum longitude, maximum and minimum latitude, maximum and minimum elevation. The first category is the coverage area of the feature data subset, which is the three-dimensional spatial region enclosed by these six extreme values. The three categories of regions are merged to form the overall feature data coverage area. The second category is the geographic spatial range traversed by all motion trajectory sequences. The coordinates of all geographic spatial locations contained in each motion trajectory are extracted, and the six extreme values of longitude, latitude, and elevation are also statistically obtained. The three-dimensional spatial region enclosed by these extreme values is the trajectory traversal range. The above two categories of spatial ranges are subjected to a union operation, that is, all spatial regions of the two regions are merged, and the overlapping parts and their unique parts are retained. The final three-dimensional spatial region is the preliminary range of the target spatial domain.
[0067] Step 501: Based on preset longitude and latitude intervals, the target spatial domain is divided into multiple regular spatial grid units in the horizontal direction. The set of spatial grid units constitutes the initial spatial grid. Specifically, this includes: defining a circular structural element centered on each geographic location point in the motion trajectory sequence and using a preset morphological operation radius as a reference; performing a morphological dilation operation on the point set composed of all geographic location points using the circular structural element to generate a continuous spatial range covering all location points and their adjacent buffer areas; performing a geometric union operation between the continuous spatial range and the geographic spatial range covered by the feature data subset, and defining the result as the target spatial domain for refined grid division; and dividing the target spatial domain horizontally based on the outer rectangle of the target spatial domain, using preset longitude and latitude intervals as steps to generate multiple regular spatial grid units that completely cover the outer rectangle. The set of spatial grid units constitutes the initial spatial grid. Specifically, this includes:
[0068] The morphological operation radius is set to 50 meters. This value is suitable for the spatial resolution of low-Earth orbit satellite sensing data and can effectively cover the adjacent buffer area of the trajectory location point, avoiding discontinuous grid coverage due to single-point discreteness. The preset longitude and latitude intervals are both 0.001 degrees, corresponding to an actual ground distance of approximately 111 meters. This interval matches the accuracy of the previous trajectory location points, balancing the degree of grid refinement and the computational efficiency at the edge. Taking each geospatial location point in the motion trajectory sequence as the center and the 50-meter morphological operation radius as the benchmark, a circular structural element is defined. The scope of the circular structural element includes all spatial points within a radius of 50 meters centered on the location point. A circular structuring element is used to perform a morphological dilation operation on the point set consisting of all geographic location points. That is, the circular structuring element regions of each location point are superimposed and merged to generate a continuous spatial range covering all location points and a 50-meter buffer zone around them, filling the spatial gaps between discrete trajectory points. The continuous spatial range generated by morphological dilation is then geometrically unioned with the geographic spatial range covered by the feature data subset in step 500 to merge the entire space of the two regions, eliminating redundancy in the overlapping parts of the regions. The result of the operation is defined as the final target spatial domain of the refined grid division, ensuring that the grid covers all feature data and includes the trajectory and the surrounding buffer zone.
[0069] Based on the final target spatial domain, its enclosing rectangle is first determined. This enclosing rectangle is the smallest horizontal rectangular area that can completely encompass the entire target spatial domain, covering only the two-dimensional plane of latitude and longitude, and simultaneously covering the entire elevation range of the target spatial domain in the vertical direction (longitude 116.1900 to 116.5100 degrees, latitude 39.7900 to 40.0900 degrees, elevation -60 to 3050 meters). The boundary parameters of the enclosing rectangle strictly correspond to the extreme values of the target spatial domain, where the longitude range is taken from the minimum longitude of the target spatial domain (116.1900 degrees) to the maximum longitude (116.5100 degrees), and the latitude range is taken from the minimum latitude of the target spatial domain. From 39.7900 degrees to a maximum of 40.0900 degrees, the rectangle is ensured to have no redundant space and completely cover the target area. Then, using preset longitude and latitude intervals of 0.001 degrees as step sizes, the outer rectangle is uniformly meshed horizontally. Before meshing, the number of mesh units is calculated: Longitude mesh count = (maximum longitude of outer rectangle - minimum longitude of outer rectangle) ÷ longitude interval; Latitude mesh count = (maximum latitude of outer rectangle - minimum latitude of outer rectangle) ÷ latitude interval. A total of several regular square spatial mesh units are generated, resulting in a total of (number of longitude meshes × number of latitude meshes). Each grid cell has fixed and uniform spatial parameters, with a horizontal longitude and latitude span of 0.001 degrees, corresponding to an actual ground distance of approximately 111 meters × 111 meters (derived based on the conversion factor of latitude / longitude to meters of 111319.9). Vertically, it covers the entire elevation range of the target spatial domain (-60 meters to 3050 meters), ensuring that each grid cell can fully incorporate data from different elevations within its corresponding horizontal range. All spatial grid cells are sequentially coded according to the rule of longitude-latitude sequence number, generating a unique grid identifier. For example, the first grid cell in the longitude direction and the first grid cell in the latitude direction are identified as 001- 001, the 320th grid in the longitude direction and the 300th grid in the latitude direction are labeled 320-300, which facilitates quick retrieval and positioning. At the same time, each grid unit is associated with its own clearly defined longitude, latitude and elevation range. The longitude and latitude ranges are sequentially extended according to the division order. For example, the grid labeled 001-001 has a longitude range of 116.1900-116.1910 degrees and a latitude range of 39.7900-39.7910 degrees. Subsequent grids follow the same pattern. The elevation range is uniformly from -60 to 3050 meters. All grid units are combined in an orderly manner according to the labeling order to form an initial spatial grid that covers the entire target spatial domain.
[0070] Step 502: For each spatial grid cell in the initial spatial grid, count the number of all data points whose spatial coordinates fall within that spatial grid cell, using this as the data point density of that spatial grid cell; count the number of all motion trajectories that pass through that spatial grid cell in the motion trajectory sequence; calculate the average value of the geometric feature parameters of all motion trajectories that pass through that spatial grid cell. Specifically, this includes: traversing each spatial grid cell, extracting its corresponding latitude and longitude range and elevation range, filtering all data points in the feature data subset whose spatial coordinates simultaneously fall within both the latitude and longitude range and elevation range of that grid cell, counting the total number of these data points, and calculating the area of that grid cell. The area is calculated by multiplying the ground distance corresponding to a longitude interval (0.001 degrees) by the ground distance corresponding to a latitude interval (0.001 degrees), and then dividing the total number of data points by the area of the grid cell to obtain the spatial grid. The data point density of the unit is expressed in units per square meter, with values rounded to four decimal places. For each motion trajectory sequence, it is determined whether it passes through the current spatial grid unit. The criteria for determination are that at least one location point among the geographic location points contained in the trajectory falls within the spatial range of the grid unit, or the trajectory line segment (the line connecting two adjacent location points) passes through the spatial range of the grid unit. The total number of motion trajectories that meet the conditions is counted and taken as the number of motion trajectories in the spatial grid unit. The geometric feature parameters (i.e., the approximate total length of the spatial curve calculated in step 404) corresponding to all motion trajectories that pass through the current grid unit are collected. These geometric feature parameter values are summed sequentially to obtain the total length sum. If the number of trajectories passing through the grid is greater than 0, the average value of the geometric feature parameters is the sum of the total lengths divided by the number of trajectories. If the number of trajectories is 0, the average value is 0, and the average value is rounded to two decimal places and taken as the average value of the geometric feature parameters of the grid unit.
[0071] Step 503: For each spatial grid cell in the initial spatial grid, a value is calculated using a preset weighting function based on the data point density, the number of motion trajectories in the spatial grid cell, and the average geometric feature parameters of the spatial grid cell. This value serves as the spatiotemporal situation confidence enhancement factor for the spatial grid cell. Specifically, for each spatial grid cell in the initial spatial grid, three types of statistical indicators are first normalized to ensure that the indicator values are uniformly within the range of 0 to 1, adapting to the weighted calculation requirements. The data point density is normalized by dividing the current grid density by the maximum density value among all grid cells; the number of motion trajectories is normalized by dividing the current number of trajectories by the maximum number of trajectories among all grid cells; the average geometric feature parameters are... The value normalization method is to divide the current grid average value by the maximum average value among all grid cells. If the maximum average value is 0, the normalized value is 0. Weights are assigned based on the influence of three types of indicators on the spatiotemporal situation: data point density (0.35), number of motion trajectories (0.35), and average geometric feature parameter (0.3). Data point density and the number of trajectories directly reflect the spatial activity and data richness within the grid, with roughly equal influence. The average geometric feature parameter reflects the motion amplitude, with a slightly lower influence. A spatiotemporal situation confidence enhancement factor is calculated using a preset weighting function. The formula is: Enhancement Factor = (Normalized Data Point Density × 0.35) + (Normalized Number of Motion Trajectories × 0.35) + (Normalized Average Geometric Feature Parameter × 0.3). The enhancement factor ranges from 0 to 1; the closer the value is to 1, the richer and more reliable the spatiotemporal situation data within the grid cell. The enhancement factor is stored in association with the grid identifier.
[0072] Step 504: For each data unit in the preferred feature dataset with quality labels, determine the spatial grid cell to which the data unit belongs based on its spatial coordinates, and read the spatiotemporal situation confidence enhancement factor corresponding to the determined spatial grid cell. Using this spatiotemporal situation confidence enhancement factor as a weight, perform a weighted multiplication operation on the original feature values carried by the data unit to obtain the corrected feature values. Specifically, this includes: firstly, traversing the preferred feature dataset according to geographic feature type (urban built-up area, vegetation cover area, water area), and performing the correction operation on each data unit one by one. The process of categorizing and traversing data improves operational efficiency and avoids data confusion across geographical elements. For a single data unit, its standardized spatial coordinates (longitude, latitude, and elevation) are first extracted. Then, spatial range matching is performed with the grid units in the initial spatial grid. The matching adopts a precise three-dimensional range determination rule: first, it is determined whether the longitude of the data unit falls within the longitude range of a certain grid unit; then, it is determined whether the latitude falls within the latitude range of the corresponding grid; finally, it is verified whether the elevation is within the elevation coverage range of the grid (-60 meters to 3050 meters). If all three conditions are met, the data unit is determined to belong to that grid unit, and the data is then verified through the grid. A unique grid identifier ensures a unique attribution relationship, preventing a single data unit from corresponding to multiple grids. After matching, the spatiotemporal situation confidence enhancement factor corresponding to that grid unit is retrieved from the database using the unique grid identifier. This enhancement factor is then used as a weight to perform a weighted multiplication operation on the original feature values carried by the data unit. The corrected feature value = original feature value × spatiotemporal situation confidence enhancement factor. Weighted multiplication is used because the enhancement factor is a comprehensive quantitative result of data richness, trajectory activity, and situation confidence within the grid unit; a higher value indicates greater reliability and effectiveness of the data in that area. The multiplication method can directly bind feature values to regional confidence levels. Feature values of high-confidence grid cells are preserved or even enhanced after multiplication (feature values remain basically unchanged when the factor is close to 1; feature values are highlighted proportionally when the factor is slightly higher than 0.5), while feature values of low-confidence grid cells are weakened proportionally (feature values approach 0 when the factor is close to 0). This achieves a differentiated optimization effect of strengthening effective features and suppressing ineffective features. At the same time, the multiplication operation logic is simple and the amount of computation is small, which is suitable for the lightweight computing needs of satellite or edge nodes. It can quickly complete the correction of the entire data without complex iterations.
[0073] During the calibration process, the feature values before and after calibration, the enhancement factor used, and the grid identifier of each unit are recorded. At the same time, the original quality level label, spatiotemporal coordinates, data source identifier, and other core information of the data unit are fully preserved to ensure full traceability of the data. If the enhancement factor of the grid to which the data unit belongs is 0, it means that the data points in the grid are sparse and there is no effective motion trajectory support, and the spatiotemporal situation is extremely unreliable. In this case, the calibrated feature value is directly set to 0 to completely weaken the interference of such data on subsequent analysis and avoid low-quality data from affecting the overall judgment results. Following the above logic, the feature value calibration of all data units under all geographic feature types is completed, and the calibration results are generated and temporarily stored one by one.
[0074] Step 505 involves reorganizing the feature values obtained after correction of all data units in the preferred feature dataset with quality labels to form a spatiotemporal situation optimized feature dataset. Specifically, this includes: classifying and grouping the feature values obtained after correction of all data units in the preferred feature dataset with quality labels according to geographic feature type (urban built-up area, vegetation cover area, water area); arranging data units of the same geographic feature type in an orderly manner according to timestamp order and spatial grid identifier; ensuring that each data unit retains complete information such as the corrected feature value, original quality level label, spatiotemporal coordinates, grid identifier, and corresponding spatiotemporal situation confidence enhancement factor to ensure data traceability and subsequent analysis needs; and integrating all the classified and organized data units to form a spatiotemporal situation optimized feature dataset. This dataset eliminates the interference of low-confidence grid data and strengthens the feature expression of high-confidence areas.
[0075] In this embodiment, the target spatial domain is fused with dual ranges and optimized through morphological dilation. This not only covers all feature data and motion trajectories but also fills in the gaps between discrete trajectories, ensuring that the grid division is comprehensive and adaptable to the situational analysis needs of multi-geographical feature regions. Fine-grained grid division, combined with reasonable parameter settings, balances spatial resolution and edge computing efficiency. Simultaneously, three types of indicators are used to statistically quantify the grid situation, providing a scientific weighting basis for feature correction and avoiding subjective judgment bias. The spatiotemporal situation confidence enhancement factor is generated through weighted calculation, integrating data density, the number of trajectories, and motion characteristics to accurately quantify the confidence of grid units, achieving differentiated correction of feature values and improving the quality and reliability of the overall feature dataset. The corrected data is organized according to geographic feature type, retaining complete core information and quality markers, eliminating the need for subsequent additional processing. This allows for direct adaptation to specialized analysis scenarios, simplifying the process and improving data processing efficiency, meeting the near-real-time monitoring needs at the edge. The impact of low-confidence grid data is weakened, while high-confidence regional features are strengthened, effectively filtering out invalid information interference and improving the accuracy of subsequent target identification and situational assessment, adapting to the refined application needs of low-orbit satellite multi-source sensing data.
[0076] In a preferred embodiment of the present invention, step 6 includes:
[0077] Step 600 involves performing lossy compression based on discrete cosine transform on the spatiotemporal situation optimization feature dataset to generate a compressed feature data matrix. Specifically, this includes: first, splitting the spatiotemporal situation optimization feature dataset according to geographic feature types; arranging each type of feature data in an ordered manner according to timestamp-grid identifier, converting it into a two-dimensional numerical matrix, where rows correspond to data units and columns correspond to feature dimensions (including corrected feature values, quality level mapping values, etc.), ensuring a unified data format compatible with the compression algorithm; dividing the two-dimensional numerical matrix into several non-overlapping sub-data blocks of 8×8 size, padding edge sub-blocks smaller than 8×8 with zeros, the zero-padding value being the mean of the corresponding feature dimension to avoid edge effects affecting compression accuracy; and performing a two-dimensional discrete cosine transform on each 8×8 sub-data block. The core of the transform is... The data is transformed from the spatial domain to the frequency domain, retaining low-frequency components (core features) and discarding high-frequency components (redundant information). The cosine basis function formula for the two-dimensional discrete cosine transform is as follows: For an 8×8 sub-block, let the position coordinates of the original spatial domain element within the sub-block be (i, j), and the position coordinates of the transformed frequency domain coefficients be (u, v), where i and j represent the row and column indices of the elements within the original 8×8 spatial domain sub-block, respectively, taking integer values from 0 to 7, corresponding to the position of each original data element from the top left to the bottom right corner of the sub-block; u and v represent the row and column indices of the coefficients within the 8×8 frequency domain coefficient matrix, also taking integer values from 0 to 7, corresponding to the position of each frequency coefficient after the transformation; π is the mathematical constant pi, used to calculate the angle value in the cosine function, and the basis function is cosine. ×cos The transformation coefficients are calculated by multiplying the original element values at all positions in the sub-data block by the corresponding cosine basis function values, summing all the product results, and finally multiplying by a scaling factor. For the coefficient (DC component) at the top left corner of the sub-block (0, 0), the scaling factor is 1 / 8; for the frequency coefficients at other positions, the scaling factor is 1 / 4, thus ensuring that the numerical range of the frequency coefficients after transformation is controllable. The rule for determining the frequency level is as follows: in the 8×8 frequency coefficient matrix, the frequency level is determined by the sum of the row index u and column index v of the coefficient position. The smaller the sum of the row index and column index, the lower the frequency of the corresponding coefficient. The DC component at coordinate (0, 0) is zero frequency, which is the lowest frequency of the entire sub-block. The larger the sum of the row index and column index, the higher the frequency of the corresponding coefficient. The coefficient at coordinate (7, 7) is the highest frequency of the entire sub-block. The low-frequency region and the high-frequency region can be clearly divided according to the size of this sum.
[0078] A custom 8×8 quantization matrix is used to quantize the transformed frequency coefficients. The quantization matrix is perfectly matched to the frequency coefficient matrix in size. The element values of the quantization matrix are set according to the frequency coefficient at the corresponding position: low-frequency regions (small row and column indices and values) have small element values to ensure high quantization accuracy and complete preservation of core feature information in these regions; high-frequency regions (large row and column indices and values) have large element values, resulting in lower quantization accuracy and allowing for reasonable loss of redundant details in these regions. The quantization process involves dividing each frequency coefficient by its corresponding quantization matrix element, and rounding the result to the nearest integer. Lossy data compression is achieved by discarding subtle numerical differences in high-frequency coefficients. After quantization of all sub-blocks, the quantized frequency coefficient matrices of each sub-block are reorganized into the overall frequency coefficient matrix according to the original block order. Then, the two-dimensional frequency coefficient matrix is converted into a one-dimensional sequence by zig-zag sorting (a serpentine traversal from the low-frequency region in the upper left corner to the high-frequency region in the lower right corner). The blank and redundant space in the high-frequency region of the two-dimensional matrix is removed to further compress the data volume. Finally, a compressed feature data matrix is generated with a compression ratio controlled between 10:1 and 15:1 to ensure that there is no significant loss of core features after data compression.
[0079] Step 601: The compressed feature data matrix is subjected to feature dimension filtering. The variance and information entropy of each feature dimension in the spatiotemporal dimension are calculated. Feature dimensions with both variance and information entropy higher than a preset threshold are selected to form a core feature subset. Specifically, this includes: first, performing inverse quantization on the compressed feature data matrix to restore it to a frequency domain coefficient matrix; then, using inverse discrete cosine transform to restore it to a spatial domain data matrix, which is only used for feature dimension filtering calculations and does not retain the complete restored data; for each feature dimension, its variance in the spatiotemporal dimension is calculated to quantify the dispersion of the feature values. The calculation process is as follows: first, the mean of all data units in that dimension is calculated; then, the square of the difference between each data unit and the mean is calculated; the sum of all squares is divided by the total number of data units to obtain the variance of that dimension. The larger the variance, the more prominent the feature region. The higher the gradation, the more effective the information entropy. For each feature dimension, the probability of all data values in that dimension is calculated, and the effective information contained in the feature is quantified by calculating the information entropy. First, the data values of that dimension are divided into equal segments (16 segments to match the feature value range). The number of data units in each segment is counted, and the probability of the corresponding segment is obtained by dividing by the total number. Then, the information entropy is calculated. The larger the information entropy, the richer the effective information carried by the feature. The variance preset threshold is set to 0.02 and the information entropy preset threshold is set to 1.5. Both thresholds have been verified by historical data and can effectively filter constant values and low-information dimensions with noise. The selection rule is to retain only feature dimensions with variance higher than 0.02 and information entropy higher than 1.5. The feature data corresponding to these dimensions are extracted in the original order to form the core feature subset. After removing invalid dimensions, the data volume can be reduced by about 30%.
[0080] Step 602 encapsulates the core feature subset, the metadata header describing the data source and processing process, and the cyclic redundancy check (CRC) code used to verify data integrity, generating a uniformly formatted edge preprocessing result data packet. Specifically, this includes: constructing the metadata header, generating a metadata header describing the data source and processing process, containing eight core pieces of information: geographic feature type, data time span, grid partitioning parameters (interval, range), compression parameters (transformation block size, quantization matrix version), feature dimension description, preprocessing step identifier (including a summary of parameters for each step), data size, and generation timestamp. All information uses ASCII encoding and has a fixed length of 512 bytes for rapid parsing; calculating the CRC code on the core feature subset and metadata header, using the CRC-32 algorithm to ensure data integrity. The calculation process involves concatenating the core feature subset and metadata header into a binary data stream, using this data stream as the divisor, and using a fixed generator polynomial (…). The data is divided by a modulo-2 division, and the 32-bit remainder is used as a cyclic redundancy check code to verify whether the data is corrupted at the receiving end. The data is encapsulated in the following order: metadata header, core feature subset, and cyclic redundancy check code. The header occupies 512 bytes, the core feature subset is stored in the middle as a binary stream, and the check code occupies 4 bytes. The whole data is encapsulated into a binary edge preprocessing result data packet with a uniform format. A 4-byte end marker is added to the end of the data packet, with the marker being a fixed value of 0xFFFFFFFF, which makes it easy for the receiving end to identify the data packet boundary.
[0081] Step 603: Transmit the edge preprocessing result data packets to the ground data center via the satellite-to-ground wireless communication link established between the low-Earth orbit satellite and the ground data center. Specifically, this includes: activating the satellite-to-ground communication module onboard the low-Earth orbit satellite, using the L-band (1.5GHz-1.6GHz) as the transmission frequency band. This band has strong anti-interference capabilities, low propagation loss, and is suitable for the motion characteristics of low-Earth orbit satellites. The link transmission rate is dynamically adjusted according to real-time bandwidth, ranging from 1Mbps to 5Mbps, with a default rate of 2Mbps to balance transmission efficiency and stability. Since the maximum transmission unit (MTU) of the satellite-to-ground communication link is 1024 bytes, edge preprocessing result data packets exceeding this size are fragmented, with a fragment size set to 1020 bytes (with 4 bytes reserved for fragmentation). Each fragment is labeled with the total number of fragments minus the current fragment number to ensure reassembly at the receiving end. Data packets are sent sequentially via the satellite-to-ground wireless communication link in fragment order. During transmission, after each fragment is transmitted, an acknowledgment signal is received from the ground data center. If no acknowledgment is received within a timeout period, the data is retransmitted, with a maximum of three retransmissions to avoid data loss. Link-layer error correction coding (convolutional code) is enabled with a coding rate of 1 / 2, and redundant check bits are added to reduce the impact of channel noise on the data. After all fragments have been transmitted, the ground data center completes data packet reassembly and cyclic redundancy check. If the check passes, a transmission completion signal is sent back, and the satellite terminates the transmission process upon receiving the signal. If the check fails, a retransmission instruction is sent back, and the satellite retransmits the corresponding fragment until the data is completely received.
[0082] In this embodiment, Discrete Cosine Transform (DCT) lossy compression precisely controls redundant information, reducing data volume while ensuring the accuracy of core features. The compression ratio is adapted to the bandwidth constraints of low-Earth orbit satellite-to-ground communication links, reducing transmission time and avoiding link congestion. Combining variance and information entropy as dual dimensions for feature selection effectively eliminates dimensions with low discriminative power and low information content, further simplifying data volume while retaining core features, thus reducing computational burden for subsequent analysis in the ground data center. The metadata header fully records the data source and processing process, achieving end-to-end data traceability. Cyclic Redundancy Check (CRC) codes and link-layer error correction coding provide dual protection to ensure data packet transmission. The integrity and reliability of transmission and parsing are ensured, avoiding data corruption caused by noise and link interference; standardized data packet format and dynamically adapted transmission parameters are used to adapt to the motion characteristics of low-orbit satellites and the instability of satellite-to-ground communication links; fragmented transmission and retransmission mechanisms further improve the data transmission success rate, ensuring that edge preprocessing results are delivered to the ground efficiently and securely; compression, filtering and encapsulation are completed at the edge end throughout the process, without the need for ground center participation in preprocessing, maximizing the advantages of edge computing, reducing the amount of satellite-to-ground communication data, and laying the foundation for ground center to quickly carry out situation analysis and target identification, thereby improving the overall data processing link efficiency.
[0083] like Figure 2As shown, embodiments of the present invention also provide an edge preprocessing system for IoT low-Earth orbit satellite sensing data, including:
[0084] The standardization module is used to receive, classify, and buffer multi-source raw sensing data streams, generate a primary data frame set, and perform standardization processing to generate a standardized data frame set.
[0085] The feature compensation module is used to analyze the spatiotemporal correlation of the standardized data frame set, dynamically determine the fusion parameters and perform interpolation alignment, generate a multi-source fused data field, and use a pre-trained lightweight ensemble learning model to perform feature-level fusion and prediction compensation on the multi-source fused data field to generate a fused feature vector.
[0086] The feature extraction module is used to perform data integrity judgment and quality classification on the fused feature vector, generate the preferred feature dataset, and extract feature data subsets corresponding to urban built-up areas, vegetation coverage areas and water areas from the preferred feature dataset;
[0087] The geometric quantization module is used to perform feature matching and displacement calculation on the serialized image data in the feature data subset, generate the motion trajectory sequence of geographic entities, and calculate the approximate total length of the spatial curve of each trajectory based on the position points in the motion trajectory sequence, as the geometric feature parameter of the trajectory.
[0088] The grid optimization module is used to construct a spatiotemporally dynamic enhanced grid by combining the geographical distribution, motion trajectory sequence and geometric feature parameters of the feature data subset, and to calculate the spatiotemporal situation confidence enhancement factor to perform weighted correction on the preferred feature dataset and generate a spatiotemporally optimized feature dataset.
[0089] The uplink transmission module is used to compress and extract features from the spatiotemporal situation optimization feature dataset, generate edge preprocessing result data packets, and transmit them to the ground data center via the satellite-to-ground link.
[0090] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0091] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0092] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. An edge preprocessing method for low-orbit satellite sensing data of the Internet of Things, characterized in that, The method includes: Step 1: Receive and classify the multi-source raw sensing data stream, generate a primary data frame set, and perform standardization processing to generate a standardized data frame set; Step 2: Analyze the spatiotemporal correlation of the standardized data frame set, dynamically determine the fusion parameters and perform interpolation alignment to generate a multi-source fusion data field, and use a pre-trained lightweight ensemble learning model to perform feature-level fusion and prediction compensation on the multi-source fusion data field to generate a fusion feature vector. Step 3: Perform data integrity assessment and quality classification on the fused feature vectors to generate an optimal feature dataset, and extract feature data subsets corresponding to urban built-up areas, vegetation coverage areas and water areas from the optimal feature dataset; Step 4: Perform feature matching and displacement calculation on the serialized image data in the feature data subset to generate motion trajectory sequences of geographic entities. Based on the location points in the motion trajectory sequences, calculate the approximate total length of the spatial curve of each trajectory as the geometric feature parameter of the trajectory. Step 5: Combining the geographical distribution, motion trajectory sequences, and geometric feature parameters of the feature data subset, a spatiotemporal dynamic enhancement grid is constructed, and a spatiotemporal situation confidence enhancement factor is calculated to perform weighted correction on the selected feature dataset, generating a spatiotemporal situation optimized feature dataset, including: Determine the geographic spatial range covered by the feature data subsets corresponding to the urban built-up area, vegetation coverage area and water area, as well as the geographic spatial range traversed by the movement trajectory sequence, and take the union of the two geographic spatial ranges to obtain the target spatial domain. Based on preset longitude and latitude intervals, the target spatial domain is divided into multiple regular spatial grid units in the horizontal direction. The collection of spatial grid units constitutes the initial spatial grid, including: Centered on each geospatial location point in the motion trajectory sequence, and based on a preset morphological operation radius, a circular structural element is defined; the circular structural element is used to perform a morphological dilation operation on the point set composed of all geospatial locations to generate a continuous spatial range covering all locations and their adjacent buffer areas. Perform a geometric union operation on the continuous spatial range and the geographic spatial range covered by the feature data subset, and define the result as the target spatial domain for fine-grained grid division. Based on the bounding rectangle of the target spatial domain, the grid is divided horizontally with preset longitude and latitude intervals as steps, generating multiple regular spatial grid cells that completely cover the bounding rectangle. The collection of spatial grid cells constitutes the initial spatial grid. For each spatial grid cell in the initial spatial grid, the number of all data points whose spatial coordinates fall within that spatial grid cell in the statistical feature data subset is used as the data point density of that spatial grid cell; the number of all motion trajectories that pass through that spatial grid cell in the motion trajectory sequence is counted; and the average value of the geometric feature parameters of all motion trajectories that pass through that spatial grid cell is calculated. For each spatial grid cell in the initial spatial grid, a value is calculated using a preset weighting function based on the data point density of the spatial grid cell, the number of motion trajectories of the spatial grid cell, and the average value of the geometric characteristic parameters of the spatial grid cell. This value serves as the spatiotemporal situation confidence enhancement factor for the spatial grid cell. For each data cell in the preferred feature dataset with quality labels, the spatial grid cell to which the data cell belongs is determined based on the spatial coordinates of the data cell, and the spatiotemporal situation confidence enhancement factor corresponding to the determined spatial grid cell is read. Using the spatiotemporal situation confidence enhancement factor as a weight, a weighted multiplication operation is performed on the original feature values carried by the data cell to obtain the corrected feature values. The feature values obtained after correcting all data units in the preferred feature dataset with quality labels are reorganized to form a spatiotemporal situation optimized feature dataset. Step 6: Compress and extract features from the spatiotemporal situation optimization feature dataset to generate edge preprocessing result data packets, and transmit them to the ground data center via the satellite-to-ground link.
2. The edge preprocessing method for IoT low-orbit satellite sensing data according to claim 1, characterized in that, Step 1 includes: Onboard processors of low-Earth orbit satellites or near-Earth edge computing nodes deployed at ground stations receive raw sensing data streams from various types of IoT sensors. Based on the source sensor type and data protocol of the raw sensing data streams, the raw sensing data streams are parsed in multiple parallel streams. The smallest data block with a complete protocol structure obtained from the parsing is defined as a data unit. Based on the source identifier carried by each data unit, it is classified and stored in an independent buffer allocated in the memory of the onboard processor or edge computing node, corresponding to each data source, forming a primary data frame set organized according to the data source identifier. The image data and electromagnetic signal data in the primary data frame set are registered in spatial coordinate system and synchronized with timestamps to generate registered and aligned image and signal data. The environmental parameter time series data are converted to a unified unit of measurement. The converted data are then subjected to outlier detection and removal, and the detected missing data points are interpolated and filled to generate processed environmental parameter data. By integrating the registered and aligned image and signal data, as well as the processed environmental parameter data, a standardized set of data frames with a unified spatiotemporal reference is generated.
3. The edge preprocessing method for IoT low-orbit satellite sensing data according to claim 2, characterized in that, Step 2 includes: For multi-source data points located within the same geographic area in a standardized data frame set, the statistical continuity of the spatial distribution of the multi-source data points is analyzed, and the dynamic radius parameter used for spatial interpolation is calculated; for time series data belonging to the same observation target in a standardized data frame set, the rate of change and periodicity characteristics of the time series data over time are analyzed, and the dynamic window size parameter used for time alignment is calculated. Based on the spatiotemporal range defined by the dynamic radius parameter and the dynamic window size parameter, spatial interpolation and temporal series alignment operations are performed on the standardized data frames of different sensors to generate a spatiotemporally continuous and consistent multi-source fusion data field. The multi-source fusion data field is input into a pre-trained lightweight ensemble learning model based on gradient boosting decision tree architecture to perform nonlinear mapping and feature-level fusion of observation data from different sensor sources, and to predict and compensate for system biases caused by differences in observation conditions or sensors, thereby generating fusion feature vectors.
4. The edge preprocessing method for IoT low-orbit satellite sensing data according to claim 3, characterized in that, Step 3 includes: For the data units in the fused feature vector, the continuity of data frame counts is checked, the integrity of observation time segment coverage is judged, and the reliability of sensor reported status is evaluated in sequence to obtain the results of the data frame count continuity check, the results of the observation time segment coverage integrity judgment, and the results of the reliability of sensor reported status evaluation. For each data unit in the fused feature vector, based on the data frame count continuity verification result, observation time segment coverage integrity judgment result, and sensor reporting status reliability assessment result corresponding to the data unit, the comprehensive usability score of the data unit is calculated. According to the preset quality threshold, data units with a comprehensive usability score lower than the quality threshold are filtered out. A corresponding quality level label is attached to each data unit that is retained after filtering. The quality level label is mapped from the comprehensive usability score of the data unit, forming a preferred feature dataset with quality labels. Based on the pre-set spatial extent index of geographic elements, all feature data whose spatial coordinates fall within the boundaries of urban built-up areas, vegetation cover areas, and water areas are retrieved and extracted from the preferred feature dataset with quality labels, forming a subset of feature data corresponding to urban built-up areas, vegetation cover areas, and water areas.
5. The edge preprocessing method for IoT low-orbit satellite sensing data according to claim 4, characterized in that, Step 4 includes: For each set of temporally continuous sequential image data in the feature data subset corresponding to urban built-up areas, vegetation coverage areas, and water areas, apparent feature points are extracted between consecutive image frames; the extracted apparent feature points of two consecutive frames are compared and associated to obtain a set of successfully matched feature point pairs. Based on the pixel coordinate difference between the successfully matched feature point pairs in two frames of images, the displacement vector of the geographic entity on the image plane is calculated; combined with the sensor pose parameters corresponding to the acquisition of serialized image data, the displacement vector is mapped from the image plane coordinate system to the geographic space coordinate system to obtain a location point of the geographic entity in geographic space. For each pair of consecutive image frames in the serialized image data, the process of feature point extraction, matching, displacement calculation and coordinate mapping is repeatedly performed. A series of geospatial location points are accumulated in time order, and a series of geospatial location points constitute the motion trajectory sequence of the geographic entity. For each motion trajectory in the motion trajectory sequence, all geospatial location points contained in the motion trajectory are read sequentially; for each read motion trajectory, the three-dimensional Euclidean distance between every two adjacent geospatial location points in the motion trajectory is calculated sequentially. The three-dimensional Euclidean distances between all adjacent geographic locations in a motion trajectory are summed, and the summation result is used as the approximate total length of the spatial curve of the motion trajectory. The approximate total length of the spatial curve is recorded as the geometric characteristic parameter of the motion trajectory.
6. The edge preprocessing method for IoT low-orbit satellite sensing data according to claim 5, characterized in that, Step 6 includes: The feature dataset for spatiotemporal situation optimization is subjected to lossy compression based on discrete cosine transform to generate a compressed feature data matrix. The compressed feature data matrix is filtered by feature dimension. The variance and information entropy of each feature dimension in the spatiotemporal dimension are calculated. Feature dimensions whose variance and information entropy are both higher than a preset threshold are selected to form the core feature subset. The core feature subset, metadata header describing the data source and processing process, and cyclic redundancy check code used to verify data integrity are encapsulated to generate edge preprocessing result data packets with a uniform format. The edge preprocessing result data packets are transmitted to the ground data center through a satellite-to-ground wireless communication link established between the low-orbit satellite and the ground data center.
7. An edge preprocessing system for low-Earth orbit satellite sensing data of the Internet of Things, wherein the system implements the method as described in any one of claims 1 to 6, characterized in that, include: The standardization module is used to receive, classify, and buffer multi-source raw sensing data streams, generate a primary data frame set, and perform standardization processing to generate a standardized data frame set. The feature compensation module is used to analyze the spatiotemporal correlation of the standardized data frame set, dynamically determine the fusion parameters and perform interpolation alignment, generate a multi-source fused data field, and use a pre-trained lightweight ensemble learning model to perform feature-level fusion and prediction compensation on the multi-source fused data field to generate a fused feature vector. The feature extraction module is used to perform data integrity judgment and quality classification on the fused feature vector, generate the preferred feature dataset, and extract feature data subsets corresponding to urban built-up areas, vegetation coverage areas and water areas from the preferred feature dataset; The geometric quantization module is used to perform feature matching and displacement calculation on the serialized image data in the feature data subset, generate the motion trajectory sequence of geographic entities, and calculate the approximate total length of the spatial curve of each trajectory based on the position points in the motion trajectory sequence, as the geometric feature parameter of the trajectory. The grid optimization module is used to construct a spatiotemporally dynamic enhanced grid by combining the geographical distribution, motion trajectory sequence and geometric feature parameters of the feature data subset, and to calculate the spatiotemporal situation confidence enhancement factor to perform weighted correction on the preferred feature dataset and generate a spatiotemporally optimized feature dataset. The uplink transmission module is used to compress and extract features from the spatiotemporal situation optimization feature dataset, generate edge preprocessing result data packets, and transmit them to the ground data center via the satellite-to-ground link.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 6.