Data bus based on multi-source remote sensing data processing
By introducing a data bus into multi-source remote sensing data processing, real-time data perception, ETL processing, hierarchical storage and data state tree construction are realized, and the problems of performance bottlenecks, data quality optimization and real-time response in the existing technology are solved, and efficient, accurate and real-time remote sensing data processing is achieved.
Patent Information
- Application Number
- CN202510472818.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing technology has performance bottlenecks in multi-source remote sensing data processing, insufficient optimization of data quality and acquisition methods, and lacks real-time adjustment and optimization mechanisms, so it is impossible to quickly respond to different application needs.
It provides a data bus based on multi-source remote sensing data processing. Through the unified data organization management module, association and integration components, hierarchical storage components and state information synchronization components, real-time data perception, ETL processing, hierarchical storage and data state tree construction, ensuring efficient data management and application.
Through refined hierarchical storage system and dynamic ETL processing, data integrity and high traceability are improved, data matching and fusion accuracy are improved, data update cycle is shortened, and the accurate, real-time and efficient processing needs of remote sensing data are met.
Smart Images

Figure CN119988478A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of data processing, and in particular to a data bus based on multi-source remote sensing data processing. Background Art
[0002] With the rapid development of earth observation technology, the scale of multi-source remote sensing data processing business is constantly expanding, and its business system is becoming increasingly complex and diverse. Due to historical reasons and technical limitations, the construction of remote sensing satellite ground systems was previously divided according to business, and different departments carried out "chimney-style" independent construction. Data production standards were not unified, data storage and management were independent, and data cross-system interconnection and interoperability were difficult, resulting in many problems in data management and application. Facing the demand for data standardization in multi-source remote sensing data processing, a more efficient and flexible data integration and interaction solution is urgently needed.
[0003] The patent document with publication number CN116737854A discloses a spatiotemporal data lake management system based on multi-source remote sensing data and its security protection method. The system includes: a data identification layer, which performs data identification on multi-source remote sensing data, monitors abnormal data during identification, and extracts and disinfects virus data separately; after disinfection, the data is monitored again to ensure that the data is normal, and then the data is restored to its original position; a data collection layer, which organizes the identified data, eliminates redundant data, and obtains digital remote sensing data from different historical periods according to the time series relationship; and collects The data is analyzed and converted to become analytical data that can be directly used; the data storage layer builds a data lake architecture, first caching the spatiotemporal data in the Kafka queue and classifying and storing it; then the spatiotemporal data is stored from the Kafka queue to the data lake through the storage component, and the data lake is partitioned; the data fusion layer converts the digital remote sensing data into spatial resolution images based on multi-source digital remote sensing data; then the medium spatial resolution images and high spatial resolution images are fused to generate composite medium and high resolution images, thereby obtaining high-quality remote sensing images.
[0004] It can be seen that the spatiotemporal data lake management system based on multi-source remote sensing data has the following problems: although the spatiotemporal data lake architecture stores data through Kafka queue caching and partition management, storage management and data migration face performance bottlenecks when processing large-scale remote sensing data; the system fuses images of different resolutions in the data fusion layer, with the goal of generating high-quality remote sensing images, but the details of data quality and collection methods are not fully optimized during the data collection process; the system relies on static data for data identification, storage and processing. When faced with changes in different data quality and storage requirements, it lacks a real-time adjustment and optimization mechanism and cannot quickly respond to different application requirements. Summary of the invention
[0005] To this end, the present invention provides a data bus based on multi-source remote sensing data processing, which is used to overcome the problems in the prior art of small application scenarios of multi-source remote sensing data processing and low fusion of multi-source remote sensing data due to over-reliance on static data processing and low data sharing through a data management mechanism of unified organization management and data status tree.
[0006] To achieve the above object, the present invention provides a data bus based on multi-source remote sensing data processing, comprising: The unified data organization and management module includes an archiving and aggregation component, which is used to perceive the product data updated by each remote sensing service plug-in set in real time, and archive and aggregate the product data to the remote sensing data global resource pool to form an archive data set; an association integration component, which is connected to the archive aggregation component and is used to perform ETL processing on the archive data set to form a standardized data set; A hierarchical storage component connected to the associated integration component for storing the standardized data set in a detailed data layer, a data intermediate layer, a data service layer, and a data application layer according to preset business requirements to form a hierarchical storage data set; A state information synchronization component connected to the hierarchical storage component to construct a data state tree based on the hierarchical storage data set, and to mark the data state in the data state tree and assign a version number to form a versioned data set; A data unified service module, which is connected to the data unified organization and management module, includes a service construction component, which is used to construct a unified service interface according to preset multi-source data processing requirements and the original data, standardized data and version control information in the versioned data set to form a data service interface; A data service component, which is connected to the service construction component, is used to provide query, browsing, statistics, download, subscription and customization services of remote sensing data through the data service interface to form a unified data service; The hierarchical storage data set includes a detailed data set, an intermediate processing data set, a data service set and an application data set.
[0007] Furthermore, the association integration component includes: A quality verification unit, used to verify the integrity, spatiotemporal consistency and data accuracy of the archived data set to form a quality verification result; A preprocessing unit, configured to perform format conversion, outlier removal and data completion on the archived data set according to the quality verification result to form a preprocessed data set; A feature matching unit connected to the preprocessing unit and configured to perform feature matching based on a feature vector of the preprocessing data set and a preset threshold group to form a matching data set; A threshold adjustment unit connected to the feature matching unit, configured to adjust the preset threshold group according to the number of matching data sets formed within a preset time period to form an adjusted threshold group; An ETL conversion unit is connected to the feature matching unit and is used to perform ETL conversion on the matching data set formed based on the adjustment threshold group to form a standardized data set.
[0008] Furthermore, the feature matching unit includes: A similarity calculation subunit, used to calculate the similarity between each feature according to the feature vector and a preset similarity measurement method to form a plurality of similarities; A similar data grabbing subunit, connected to the similarity calculating subunit, for grabbing corresponding pre-processed data in the pre-processed data set when the similarity is greater than a preset similarity threshold in the preset threshold group, to form a similar data set; A time series matching subunit, which is connected to the similar data grabbing subunit, and is used to perform time data matching based on the similar data set, and determine that the match is successful when the time offset of the data match is less than the preset time offset threshold in the preset threshold group, so as to form a time series data set; A spatial matching subunit, connected to the temporal matching subunit, for performing spatial data matching based on the temporal data set, and determining that the matching is successful when the spatial overlap of the data matching is greater than a preset overlap threshold in the preset threshold group, to form a spatial data set; A feature fusion subunit is connected to the spatial matching subunit and is used to fuse the spatial features in the spatial data set to form the matching data set.
[0009] Furthermore, the threshold adjustment unit includes: A quantity fluctuation calculation subunit, used to calculate the standard deviation of the formed quantity to form a quantity fluctuation value; An adjustment subunit is connected to the quantity fluctuation calculation subunit, and is used to reduce the preset similarity threshold according to the relative deviation between the quantity fluctuation value and the maximum value of the preset quantity fluctuation range when the quantity fluctuation value is greater than the maximum value of the preset quantity fluctuation range, so as to form an adjustment threshold group; or to increase the preset time offset threshold when the quantity fluctuation value is greater than the midpoint value within the preset quantity fluctuation range, so as to form an adjustment threshold group; or to reduce the preset overlap threshold when the quantity fluctuation value is between the minimum value and the midpoint value of the preset quantity fluctuation range, so as to form an adjustment threshold group.
[0010] Furthermore, the hierarchical storage component includes: A detailed data storage unit, used for storing the standardized data that has not been aggregated and processed in the standardized data set to form the detailed data set; A data intermediate layer storage unit, which is connected to the detailed data storage unit and is used to clean, convert, multi-dimensionally aggregate and store the detailed data set to form the intermediate processing data set; A data service layer storage unit connected to the data intermediate layer storage unit for performing index optimization, partition storage, query acceleration and storage on the intermediate processing data set to form the data service set; A data application layer storage unit is connected to the data service layer storage unit and is used to extract and store visual analysis data and decision support data from the data service set based on the preset business requirements to form the application data set.
[0011] Furthermore, the state information synchronization component includes: A state tree construction unit, configured to initialize each node in a preset state tree based on the data update time and storage level in the hierarchical storage data set to form the data state tree; A status marking unit, used for classifying and marking the data according to the data update time to form a data status marking set; A version control unit connected to the state marking unit, used to assign a unique version number to each data state in the state marking data set, and maintain version change records to form version management information; A state synchronization unit is connected to the version control unit and is used to synchronize a number of version management information into the data state tree according to the version management information and a preset time window to form a versioned data set.
[0012] Furthermore, the state synchronization unit includes: a change value calculation subunit, used to calculate the difference between the timestamps in the data state of the current version and the previous adjacent version in the version management information to form a data state change value; The state synchronization subunit is connected to the change value calculation subunit, and is used to synchronize the version management information to the data state tree to form a versioned data set when the data state change value is greater than the preset time window.
[0013] Furthermore, the service construction component includes: A requirement analysis unit, used to analyze the preset multi-source data processing requirements to obtain functional requirements and interface requirements; A data mapping unit, used to uniformly map and format-convert the original data, standardized data and the version control information in the versioned data set to form a data mapping table; An interface design unit, which is connected to the demand analysis unit and the data mapping unit respectively, and is used to design the unified service interface according to the functional requirements, the interface requirements, the data mapping table and the preset business requirements to form a service interface design document; An interface encapsulation unit is connected to the interface design unit and is used to perform encapsulation according to the service interface design document to construct the unified service interface and form the data service interface.
[0014] Furthermore, the data service component includes: Demand identification unit, used to identify the user's query, browsing, statistics, downloading, subscription and customization needs, and form a user demand set based on the data request input by the user; A demand processing unit, connected to the demand identification unit, for processing and screening the original data, the standardized data and the version control information to form a processed data set; A data interaction unit is connected to the demand processing unit to provide the processed data set through a constructed service interface to form a unified data service.
[0015] Furthermore, the unified data service includes a general data service and a dedicated data service: The general data services include metadata services, remote sensing data publishing services, remote sensing data subscription services and remote sensing data retrieval services; The dedicated data services include auxiliary data services, calibration data services, optical remote sensing image data services and multi-spectral remote sensing image services.
[0016] Compared with the prior art, the beneficial effect of the present invention is that, through a refined hierarchical storage system and dynamic ETL processing, the whole process from the original collection of remote sensing data to the fine processing is seamlessly connected, ensuring that the data is effectively preserved and independently managed in the detail layer, intermediate processing layer, service layer and application layer, thereby greatly improving the integrity and high traceability of the data. The data state tree constructed by the state information synchronization component can monitor the update time and state changes of each layer of data in real time, and dynamically adjust the data state by using version control and negative feedback mechanism, accurately calculate the change value between adjacent versions, and realize the adaptive correction of the threshold. This not only improves the accuracy of data matching and fusion, but also significantly shortens the data update cycle, ensuring that the processing process meets the actual business needs. At the same time, a variety of data services based on a unified service interface make data calls in different application scenarios efficient and convenient, meet the requirements for accurate, real-time and efficient processing of remote sensing data, and effectively solve the problems of small application scenarios of multi-source remote sensing data processing and low fusion of multi-source remote sensing data due to over-reliance on static data processing and low data sharing.
[0017] Furthermore, the basic quality of the data was ensured through quality verification and preprocessing. Combined with feature matching and dynamic threshold adjustment mechanism, multi-level processing from archived data to standardized data was achieved, which significantly improved the accuracy and consistency of data matching and effectively reduced the error rate. Finally, the standardized data set generated by ETL transformation provided a high-quality, standardized data foundation for subsequent data storage and application, greatly improving the operational efficiency of data processing and data utilization.
[0018] Furthermore, through multi-level similarity, time series, spatial matching and fusion processing, the matching accuracy and consistency of remote sensing data from different sources are improved. Similarity calculation and similar data capture ensure the similarity of data features, time series matching ensures the temporal correlation of data, and spatial matching enhances the spatial consistency of data. In addition, feature fusion further improves the integrity and availability of data, so that the final generated matching data set can better support subsequent data analysis and application, and improve the service value and reliability of remote sensing data.
[0019] Furthermore, by dynamically adjusting the matching threshold, the adaptability of the matching data set is enhanced, so that the matching rules can be optimized with data fluctuations, avoiding data matching failure or over-screening problems caused by fixed parameters. At the same time, this method can improve the flexibility and applicability of data integration while ensuring data quality, making data matching more stable and accurate, and suitable for a variety of complex application scenarios.
[0020] Furthermore, through the layered storage method, not only the efficiency of data storage is optimized, but also the flexibility and speed of data processing and access are improved. The detailed data storage unit ensures the integrity of the data, the middle-layer storage unit enhances the availability of the data, the data service layer storage unit ensures efficient query and fast response, and the application layer storage unit meets the fast data extraction and analysis of specific business needs, effectively improving the scalability, flexibility and efficiency, so that data at different levels can be reasonably managed and used according to specific needs.
[0021] Furthermore, through real-time tracking and version control of data status, the consistency, accuracy and traceability of data are effectively guaranteed. Through precise version management and synchronization mechanisms, the consistency and stability of data at all stages can be ensured, and the reliability and transparency of data processing can be improved.
[0022] Furthermore, by calculating the timestamp difference between the current version and the previous version, the change in data status can be effectively judged, ensuring that the data status tree is synchronized only when there is a significant change in the data status. This not only avoids unnecessary frequent updates, but also ensures efficient operation of the system. When the data status change value is greater than the preset time window, the system immediately synchronizes the version management information to the data status tree to ensure that the data status reflects the actual changes in a timely and accurate manner, thereby improving data consistency and accuracy. In this way, data management becomes more flexible and intelligent, and can dynamically adapt to different data update requirements, reduce unnecessary storage and computing overhead, and improve the response speed and stability of the system.
[0023] Furthermore, demand analysis and data mapping ensure the accuracy and compatibility of interface design, avoiding the problem of inconsistent data formats. Through standardized interface design, the system can provide users with consistent and flexible data access interfaces, improving the scalability and maintainability of data services. At the same time, interface encapsulation operations make service integration easier, enabling rapid response to changes in business requirements, and helping to improve overall system performance and user experience.
[0024] Furthermore, by providing high-quality and high-availability data services through clear screening and processing processes, the flexibility and response efficiency of the system are improved. The unified data service interface simplifies the user's operation process, reduces dependence on different data sources and service interfaces, and makes data access more efficient and convenient.
[0025] Furthermore, general data services can meet the basic needs of most users for remote sensing data, ensuring the efficiency and wide applicability of data access and management, while dedicated data services provide deeply customized support for specific fields and needs, improving the application value and accuracy of remote sensing data. This service classification structure enhances the flexibility and adaptability of the system and can provide optimized service solutions under different user needs and application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a schematic diagram of the architecture of the data bus based on multi-source remote sensing data processing in this embodiment; Figure 2 A flow chart of forming a matching data set by the feature matching unit of this embodiment; Figure 3 A flowchart of forming a hierarchical storage data set for the hierarchical storage component of this embodiment; Figure 4 A decision logic diagram for forming a versioned data set for the state synchronization unit of this embodiment. DETAILED DESCRIPTION
[0027] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0028] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0029] See also Figure 1 As shown, it is a schematic diagram of the data bus based on multi-source remote sensing data processing in this embodiment; This embodiment provides a data bus based on multi-source remote sensing data processing, including: The unified data organization and management module includes an archiving and aggregation component, which is used to perceive the product data updated by each remote sensing service plug-in set in real time, and archive and aggregate the product data to the remote sensing data global resource pool to form an archive data set; an association integration component, which is connected to the archive aggregation component and is used to perform ETL processing on the archive data set to form a standardized data set; A hierarchical storage component connected to the associated integration component for storing the standardized data set in a detailed data layer, a data intermediate layer, a data service layer, and a data application layer according to preset business requirements to form a hierarchical storage data set; A state information synchronization component connected to the hierarchical storage component to construct a data state tree based on the hierarchical storage data set, and to mark the data state in the data state tree and assign a version number to form a versioned data set; A data unified service module, which is connected to the data unified organization and management module, includes a service construction component, which is used to construct a unified service interface according to preset multi-source data processing requirements and the original data, standardized data and version control information in the versioned data set to form a data service interface; A data service component, which is connected to the service construction component, is used to provide query, browsing, statistics, download, subscription and customization services of remote sensing data through the data service interface to form a unified data service; The hierarchical storage data set includes a detailed data set, an intermediate processing data set, a data service set and an application data set.
[0030] Product data refers to the remote sensing data results obtained and updated in real time by different remote sensing business plug-in sets (i.e. business modules designed for different remote sensing data types or application scenarios) during data collection, preprocessing, analysis or product generation. It usually includes remote sensing images that have been initially processed or corrected, feature extraction results, data indicators or other data information related to remote sensing information, providing a basic basis for subsequent data integration, storage and application services.
[0031] Preset business requirements are a series of goals and performance indicators that are pre-established at the beginning of system design based on specific business scenarios and user needs to guide the construction of data storage and service interfaces. Their formulation depends on the actual scenarios of the application field, user expectations, and the quality and quantity of existing data resources.
[0032] First, the archive aggregation component in the data unified organization and management module perceives the product data updated by each remote sensing business plug-in set in real time, and archives and aggregates it into the remote sensing data global resource pool to form an archive data set; then, the association integration component performs ETL processing on the archive data set to generate a standardized data set; then, the hierarchical storage component stores the standardized data set in the detailed data layer, intermediate processing data layer, data service layer and data application layer according to the preset business requirements, forming a complete hierarchical storage data set; then, the status information synchronization component builds a data status tree based on the hierarchical storage data set, marks the status of each node and assigns a version number to form a versioned data set; finally, the service construction component in the data unified service module builds a unified service interface based on the preset multi-source data processing requirements and the original data, standardized data and version control information in the versioned data set, and the data service component uses this interface to provide remote sensing data query, browsing, statistics, download, subscription and customization services to form a unified data service.
[0033] Through the refined hierarchical storage system and dynamic ETL processing, the whole process from the original collection of remote sensing data to the fine processing is seamlessly connected, ensuring that the data is effectively preserved and independently managed in the detail layer, intermediate processing layer, service layer and application layer, thereby greatly improving the integrity and high traceability of the data. The data state tree constructed by the state information synchronization component can monitor the update time and state changes of each layer of data in real time, and dynamically adjust the data state using version control and negative feedback mechanisms, accurately calculate the change value between adjacent versions, and realize adaptive correction of thresholds. This not only improves the accuracy of data matching and fusion, but also significantly shortens the data update cycle, ensuring that the processing process meets actual business needs. At the same time, a variety of data services based on a unified service interface make data calls in different application scenarios efficient and convenient, meet the requirements for accurate, real-time and efficient processing of remote sensing data, and effectively solve the problems of small application scenarios and low fusion of multi-source remote sensing data processing due to over-reliance on static data processing and low data sharing.
[0034] Specifically, the association integration component includes: A quality verification unit, used to verify the integrity, spatiotemporal consistency and data accuracy of the archived data set to form a quality verification result; A preprocessing unit, configured to perform format conversion, outlier removal and data completion on the archived data set according to the quality verification result to form a preprocessed data set; A feature matching unit connected to the preprocessing unit and configured to perform feature matching based on a feature vector of the preprocessing data set and a preset threshold group to form a matching data set; A threshold adjustment unit connected to the feature matching unit, configured to adjust the preset threshold group according to the number of matching data sets formed within a preset time period to form an adjusted threshold group; An ETL conversion unit is connected to the feature matching unit and is used to perform ETL conversion on the matching data set formed based on the adjustment threshold group to form a standardized data set.
[0035] The preset threshold group refers to a series of threshold parameters set in the process of multi-source remote sensing data processing to ensure the accuracy and consistency of data matching, including the preset similarity threshold, the preset time offset threshold and the preset overlap threshold. These thresholds are used to determine the degree of matching, time synchronization accuracy and spatial overlap ratio between data, so as to improve the quality of data fusion, reduce mismatching and information loss, and ensure the reliability and stability of data processing results.
[0036] First, the integrity, spatiotemporal consistency and data accuracy of the archived data set are verified by the quality verification unit to form a quality verification result; then, the preprocessing unit performs format conversion, outlier elimination and data completion on the archived data set based on the verification result to form a preprocessed data set; subsequently, the feature matching unit uses the feature vectors in the preprocessed data set and the preset threshold group to perform feature matching, extract similar data and form a matching data set; next, the threshold adjustment unit dynamically adjusts the preset threshold group according to the number of matching data sets formed within a preset time length to form an adjusted threshold group; finally, the ETL conversion unit performs ETL conversion on the matching data set based on the adjusted threshold group to form a standardized data set.
[0037] The basic quality of the data is ensured through quality verification and preprocessing. Combined with feature matching and dynamic threshold adjustment mechanism, multi-level processing from archived data to standardized data is realized, which significantly improves the accuracy and consistency of data matching and effectively reduces the error rate. Finally, the standardized data set generated by ETL transformation provides a high-quality and standardized data foundation for subsequent data storage and application, greatly improving the operational efficiency of data processing and data utilization effect.
[0038] Please continue reading Figure 2 As shown, it is a flow chart of the feature matching unit forming a matching data set in this embodiment; The feature matching unit comprises: A similarity calculation subunit, used to calculate the similarity between each feature according to the feature vector and a preset similarity measurement method to form a plurality of similarities; A similar data grabbing subunit, connected to the similarity calculating subunit, for grabbing corresponding pre-processed data in the pre-processed data set when the similarity is greater than a preset similarity threshold in the preset threshold group, to form a similar data set; A time series matching subunit, which is connected to the similar data grabbing subunit, and is used to perform time data matching based on the similar data set, and determine that the match is successful when the time offset of the data match is less than the preset time offset threshold in the preset threshold group, so as to form a time series data set; A spatial matching subunit, connected to the temporal matching subunit, for performing spatial data matching based on the temporal data set, and determining that the matching is successful when the spatial overlap of the data matching is greater than a preset overlap threshold in the preset threshold group, to form a spatial data set; A feature fusion subunit is connected to the spatial matching subunit and is used to fuse the spatial features in the spatial data set to form the matching data set.
[0039] The preset similarity threshold is the minimum standard for measuring the similarity between data feature vectors, which depends on the data type, the calculation method of the feature vector and the business needs. It is usually set between 0.7 and 0.95. In this embodiment, it is set to 0.85 to ensure high accuracy of matching data while taking into account data coverage.
[0040] The preset time offset threshold is the maximum allowable deviation for determining whether the time attributes of two data items match. It depends on the acquisition frequency of remote sensing data, the time sensitivity of the application scenario, and the data alignment requirements. It is usually set between 5 seconds and 30 minutes to meet the time matching requirements of different business needs. In this embodiment, it is set to 10 minutes to ensure data timeliness while providing sufficient time redundancy to match asynchronously collected data.
[0041] The preset overlap threshold is the minimum standard for measuring the degree of data overlap during spatial matching. It depends on the resolution of the remote sensing image, the business requirements of data overlap, and the spatial distribution characteristics of the target area. It is usually set between 50% and 90%. In this embodiment, it is set to 75% to ensure a high degree of spatial matching of the data while taking into account the availability of the data.
[0042] The preset similarity measurement method is an algorithm for calculating the similarity between data feature vectors. In this embodiment, cosine similarity is used to effectively measure the directional similarity between feature vectors while avoiding the influence of numerical scale on the calculation results.
[0043] The feature matching unit first calculates the similarity between each feature vector in the preprocessed data set through the similarity calculation subunit, and filters out the data with similarity higher than the preset similarity threshold to form a similar data set. Subsequently, the time series matching subunit performs time matching on the similar data set to ensure that the time offset is within the preset time offset threshold to form a time series data set. Next, the spatial matching subunit performs spatial matching on the time series data set, filters out the data with spatial overlap higher than the preset overlap threshold to form a spatial data set. Finally, the feature fusion subunit performs feature fusion on the spatial data set to generate a matching data set to ensure the feature consistency and integrity of multi-source data.
[0044] Through multi-level similarity, time series, spatial matching and fusion processing, the matching accuracy and consistency of remote sensing data from different sources are improved. Similarity calculation and similar data capture ensure the similarity of data features, time series matching ensures the temporal relevance of data, and spatial matching enhances the spatial consistency of data. In addition, feature fusion further improves the integrity and availability of data, so that the final generated matching data set can better support subsequent data analysis and application, and improve the service value and reliability of remote sensing data.
[0045] Specifically, the threshold adjustment unit includes: A quantity fluctuation calculation subunit, used to calculate the standard deviation of the formed quantity to form a quantity fluctuation value; An adjustment subunit is connected to the quantity fluctuation calculation subunit, and is used to reduce the preset similarity threshold according to the relative deviation between the quantity fluctuation value and the maximum value of the preset quantity fluctuation range when the quantity fluctuation value is greater than the maximum value of the preset quantity fluctuation range, so as to form an adjustment threshold group; or to increase the preset time offset threshold when the quantity fluctuation value is greater than the midpoint value within the preset quantity fluctuation range, so as to form an adjustment threshold group; or to reduce the preset overlap threshold when the quantity fluctuation value is between the minimum value and the midpoint value of the preset quantity fluctuation range, so as to form an adjustment threshold group.
[0046] The preset quantity fluctuation range refers to a predetermined range used to measure the fluctuation of the quantity of the matching data set, which is set according to the historical performance of the data set and the expected data fluctuation, and is usually set between 0.05 and 0.2. In this embodiment, it is set to 0.1 to 0.15, which can balance the tolerance for normal fluctuations while avoiding excessive threshold adjustment caused by excessive fluctuations, ensuring the stability and adaptability of the data matching process.
[0047] First, the standard deviation of the number of matching data sets is calculated by the quantity fluctuation calculation subunit to obtain the quantity fluctuation value. Then, the adjustment subunit makes dynamic adjustments based on the range of the quantity fluctuation value: when the quantity fluctuation value exceeds the maximum value of the preset quantity fluctuation range, the preset similarity threshold is lowered to expand the matching range; when the quantity fluctuation value is higher than the midpoint value of the range but does not exceed the maximum value, the preset time offset threshold is increased to allow for a larger time matching error; when the quantity fluctuation value is between the minimum value and the midpoint value, the preset overlap threshold is lowered to relax the spatial matching requirements, and finally a new adjustment threshold group is formed.
[0048] By dynamically adjusting the matching threshold, the adaptability of the matching data set is enhanced, so that the matching rules can be optimized as the data fluctuates, avoiding data matching failure or over-screening problems caused by fixed parameters. At the same time, this method can improve the flexibility and applicability of data integration while ensuring data quality, making data matching more stable and accurate, and suitable for a variety of complex application scenarios.
[0049] Please continue reading Figure 3 As shown, it is a flow chart of the hierarchical storage component of this embodiment forming a hierarchical storage data set; The hierarchical storage component includes: A detailed data storage unit, used for storing the standardized data that has not been aggregated and processed in the standardized data set to form the detailed data set; A data intermediate layer storage unit, which is connected to the detailed data storage unit and is used to clean, convert, multi-dimensionally aggregate and store the detailed data set to form the intermediate processing data set; A data service layer storage unit connected to the data intermediate layer storage unit for performing index optimization, partition storage, query acceleration and storage on the intermediate processing data set to form the data service set; A data application layer storage unit is connected to the data service layer storage unit and is used to extract and store visual analysis data and decision support data from the data service set based on the preset business requirements to form the application data set.
[0050] First, the detailed data storage unit stores the original standardized data that has not been aggregated to form a detailed data set; then, the data intermediate layer storage unit cleans, converts and multi-dimensionally aggregates these data to form an intermediate processing data set; then, the data service layer storage unit optimizes the index and accelerates the query of the intermediate processing data to form a data service set; finally, the data application layer storage unit extracts visual analysis data and decision support data from the data service set according to preset business needs to form an application data set.
[0051] The hierarchical storage method not only optimizes the efficiency of data storage, but also improves the flexibility and speed of data processing and access. The detailed data storage unit ensures the integrity of the data, the middle-layer storage unit enhances the availability of the data, the data service layer storage unit ensures efficient query and fast response, and the application layer storage unit meets the fast data extraction and analysis of specific business needs, effectively improving the scalability, flexibility and efficiency, so that data at different levels can be reasonably managed and used according to specific needs.
[0052] Specifically, the state information synchronization component includes: A state tree construction unit, configured to initialize each node in a preset state tree based on the data update time and storage level in the hierarchical storage data set to form the data state tree; A status marking unit, used for classifying and marking the data according to the data update time to form a data status marking set; A version control unit connected to the state marking unit, used to assign a unique version number to each data state in the state marking data set, and maintain version change records to form version management information; A state synchronization unit is connected to the version control unit and is used to synchronize a number of version management information into the data state tree according to the version management information and a preset time window to form a versioned data set.
[0053] Data status refers to the relevant information used to describe data at different time points and processing stages, including timestamps, version numbers, processing status, quality tags, storage levels, and update times. Timestamps mark the time when data is generated or updated, version numbers are used to distinguish different data versions, processing status reflects whether the data has been cleaned, converted, and other steps, quality tags indicate the quality of the data, storage levels show the location of data in different storage layers, and update time records the latest update time of the data. Together, these elements help manage and track changes in data throughout its life cycle, ensuring data validity and consistency.
[0054] The state information synchronization component coordinates work through multiple units to ensure that the state and version information of the data are effectively managed. First, the state tree construction unit initializes each node based on the update time and storage level of the hierarchical storage data set to form a data state tree. Then, the state marking unit classifies and marks the data according to the data update time to form a data state marking set. The version control unit assigns a version number to the state marking data set and maintains the version change record to form version management information. Finally, the state synchronization unit synchronizes the version management information to the data state tree to form a versioned data set.
[0055] Through real-time tracking and version control of data status, the consistency, accuracy and traceability of data are effectively guaranteed. Through precise version management and synchronization mechanisms, the consistency and stability of data at all stages can be ensured, and the reliability and transparency of data processing can be improved.
[0056] Please continue reading Figure 4 As shown, it is a decision logic diagram of the state synchronization unit in this embodiment to form a versioned data set; Specifically, the state synchronization unit includes: a change value calculation subunit, used to calculate the difference between the timestamps in the data state of the current version and the previous adjacent version in the version management information to form a data state change value; The state synchronization subunit is connected to the change value calculation subunit, and is used to synchronize the version management information to the data state tree to form a versioned data set when the data state change value is greater than the preset time window.
[0057] The preset time window refers to a pre-set time range or period during the data processing or synchronization process, which is used to determine and control the frequency of data status changes, synchronization or updates. It depends on the real-time requirements of the system, the frequency of data updates, and the needs of specific business scenarios. It is set between 10 seconds and 2 hours. In this embodiment, it is set to 1 hour, which can balance the real-time performance of data synchronization and the system load, avoid too frequent updates or synchronization operations, and ensure that the system can process data efficiently and stably.
[0058] The state synchronization unit calculates the timestamp difference between the current version and the previous version in the version management information through the change value calculation subunit to form a data state change value. When the data state change value exceeds the preset time window, the state synchronization subunit synchronizes the version management information to the data state tree to ensure the update of the versioned data set.
[0059] By calculating the timestamp difference between the current version and the previous version, the change in data status can be effectively judged, ensuring that the data status is synchronized to the data status tree only when there is a significant change in the data status. This not only avoids unnecessary frequent updates, but also ensures efficient operation of the system. When the data status change value is greater than the preset time window, the system immediately synchronizes the version management information to the data status tree to ensure that the data status reflects the actual changes in a timely and accurate manner, thereby improving data consistency and accuracy. In this way, data management becomes more flexible and intelligent, and can dynamically adapt to different data update requirements, reduce unnecessary storage and computing overhead, and improve the response speed and stability of the system.
[0060] Specifically, the service building components include: A requirement analysis unit, used to analyze the preset multi-source data processing requirements to obtain functional requirements and interface requirements; A data mapping unit, used to uniformly map and format-convert the original data, standardized data and the version control information in the versioned data set to form a data mapping table; An interface design unit, which is connected to the demand analysis unit and the data mapping unit respectively, and is used to design the unified service interface according to the functional requirements, the interface requirements, the data mapping table and the preset business requirements to form a service interface design document; An interface encapsulation unit is connected to the interface design unit and is used to perform encapsulation according to the service interface design document to construct the unified service interface and form the data service interface.
[0061] Preset multi-source data processing requirements refer to pre-set data processing requirements based on the specific needs and goals of multiple data sources. These requirements cover the rules and standards for how to acquire, process, fuse and present data from multiple different data sources (such as sensors, remote sensing equipment, databases, etc.) to ensure the quality, availability and consistency of the data. They depend on factors such as the business objectives of the system, the type of data source, the complexity of data processing and real-time requirements.
[0062] First, the requirements analysis unit analyzes the preset multi-source data processing requirements to determine the functional requirements and interface requirements; then, the data mapping unit uniformly maps and formats the original data, standardized data, and version control information in the versioned data set according to the requirements to form a data mapping table; the interface design unit designs a unified service interface based on the functional requirements, interface requirements, data mapping table, and preset business requirements, and generates a service interface design document; finally, the interface encapsulation unit encapsulates the interface according to the service interface design document to construct a data service interface.
[0063] Demand analysis and data mapping ensure the accuracy and compatibility of interface design, avoiding the problem of inconsistent data formats. Through standardized interface design, the system can provide users with a consistent and flexible data access interface, improving the scalability and maintainability of data services. At the same time, the interface encapsulation operation makes service integration easier, can quickly respond to changes in business requirements, and help improve the overall system performance and user experience.
[0064] Specifically, the data service components include: Demand identification unit, used to identify the user's query, browsing, statistics, downloading, subscription and customization needs, and form a user demand set based on the data request input by the user; A demand processing unit, connected to the demand identification unit, for processing and screening the original data, the standardized data and the version control information to form a processed data set; A data interaction unit is connected to the demand processing unit to provide the processed data set through a constructed service interface to form a unified data service.
[0065] First, the demand identification unit identifies the user's data needs such as query, browsing, statistics, downloading, subscription and customization, and generates a user demand set. Then, the demand processing unit screens and processes the original data, standardized data and version control information to form a processed data set. Finally, the data interaction unit provides the processed data set to the user through the pre-built service interface to form a unified data service.
[0066] Providing high-quality and high-availability data services through clear screening and processing processes improves the flexibility and response efficiency of the system. The unified data service interface simplifies the user's operation process, reduces dependence on different data sources and service interfaces, and makes data access more efficient and convenient.
[0067] Specifically, the unified data service includes general data services and dedicated data services: The general data services include metadata services, remote sensing data publishing services, remote sensing data subscription services and remote sensing data retrieval services; The dedicated data services include auxiliary data services, calibration data services, optical remote sensing image data services and multi-spectral remote sensing image services.
[0068] General data services provide metadata services, remote sensing data publishing services, remote sensing data subscription services and remote sensing data retrieval services to meet users' needs for basic data query, publishing, subscription and retrieval. Special data services provide more professional services, including auxiliary data services, calibration data services, optical remote sensing image data services and multispectral remote sensing image services, focusing on data processing and application needs in specific fields.
[0069] General data services can meet the basic needs of most users for remote sensing data, ensuring the efficiency and wide applicability of data access and management, while dedicated data services provide deeply customized support for specific fields and needs, improving the application value and accuracy of remote sensing data. This service classification structure enhances the flexibility and adaptability of the system, and can provide optimized service solutions under different user needs and application scenarios.
[0070] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A data bus based on multi-source remote sensing data processing, characterized in that: include: The unified data organization and management module includes an archiving and aggregation component, which is used to perceive the product data updated by each remote sensing service plug-in set in real time, and archive and aggregate the product data to the remote sensing data global resource pool to form an archive data set; an association integration component, which is connected to the archive aggregation component and is used to perform ETL processing on the archive data set to form a standardized data set; A hierarchical storage component connected to the associated integration component for storing the standardized data set in a detailed data layer, a data intermediate layer, a data service layer, and a data application layer according to preset business requirements to form a hierarchical storage data set; A state information synchronization component connected to the hierarchical storage component to construct a data state tree based on the hierarchical storage data set, and to mark the data state in the data state tree and assign a version number to form a versioned data set; A data unified service module, which is connected to the data unified organization and management module, includes a service construction component, which is used to construct a unified service interface according to preset multi-source data processing requirements and the original data, standardized data and version control information in the versioned data set to form a data service interface; A data service component, which is connected to the service construction component, is used to provide query, browsing, statistics, download, subscription and customization services of remote sensing data through the data service interface to form a unified data service; The hierarchical storage data set includes a detailed data set, an intermediate processing data set, a data service set and an application data set.
2. The data bus based on multi-source remote sensing data processing according to claim 1, characterized in that: The association integration component includes: A quality verification unit, used to verify the integrity, spatiotemporal consistency and data accuracy of the archived data set to form a quality verification result; A preprocessing unit, configured to perform format conversion, outlier removal and data completion on the archived data set according to the quality verification result to form a preprocessed data set; A feature matching unit connected to the preprocessing unit and configured to perform feature matching based on a feature vector of the preprocessing data set and a preset threshold group to form a matching data set; A threshold adjustment unit connected to the feature matching unit, configured to adjust the preset threshold group according to the number of matching data sets formed within a preset time period to form an adjusted threshold group; An ETL conversion unit is connected to the feature matching unit and is used to perform ETL conversion on the matching data set formed based on the adjustment threshold group to form a standardized data set.
3. The data bus based on multi-source remote sensing data processing according to claim 2, characterized in that: The feature matching unit comprises: A similarity calculation subunit, used to calculate the similarity between each feature according to the feature vector and a preset similarity measurement method to form a plurality of similarities; A similar data grabbing subunit, connected to the similarity calculating subunit, for grabbing corresponding pre-processed data in the pre-processed data set when the similarity is greater than a preset similarity threshold in the preset threshold group, to form a similar data set; A time series matching subunit, which is connected to the similar data grabbing subunit, and is used to perform time data matching based on the similar data set, and determine that the match is successful when the time offset of the data match is less than the preset time offset threshold in the preset threshold group, so as to form a time series data set; A spatial matching subunit, connected to the temporal matching subunit, for performing spatial data matching based on the temporal data set, and determining that the matching is successful when the spatial overlap of the data matching is greater than a preset overlap threshold in the preset threshold group, to form a spatial data set; A feature fusion subunit is connected to the spatial matching subunit and is used to fuse the spatial features in the spatial data set to form the matching data set.
4. The data bus based on multi-source remote sensing data processing according to claim 3, characterized in that: The threshold adjustment unit comprises: A quantity fluctuation calculation subunit, used to calculate the standard deviation of the formed quantity to form a quantity fluctuation value; An adjustment subunit is connected to the quantity fluctuation calculation subunit, and is used to reduce the preset similarity threshold according to the relative deviation between the quantity fluctuation value and the maximum value of the preset quantity fluctuation range when the quantity fluctuation value is greater than the maximum value of the preset quantity fluctuation range, so as to form an adjustment threshold group; or to increase the preset time offset threshold when the quantity fluctuation value is greater than the midpoint value within the preset quantity fluctuation range, so as to form an adjustment threshold group; or to reduce the preset overlap threshold when the quantity fluctuation value is between the minimum value and the midpoint value of the preset quantity fluctuation range, so as to form an adjustment threshold group.
5. The data bus based on multi-source remote sensing data processing according to claim 4, characterized in that: The hierarchical storage component includes: A detailed data storage unit, used for storing the standardized data that has not been aggregated and processed in the standardized data set to form the detailed data set; A data intermediate layer storage unit, which is connected to the detailed data storage unit and is used to clean, convert, multi-dimensionally aggregate and store the detailed data set to form the intermediate processing data set; A data service layer storage unit connected to the data intermediate layer storage unit for performing index optimization, partition storage, query acceleration and storage on the intermediate processing data set to form the data service set; A data application layer storage unit is connected to the data service layer storage unit and is used to extract and store visual analysis data and decision support data from the data service set based on the preset business requirements to form the application data set.
6. The data bus based on multi-source remote sensing data processing according to claim 5, characterized in that: The state information synchronization component includes: A state tree construction unit, configured to initialize each node in a preset state tree based on the data update time and storage level in the hierarchical storage data set to form the data state tree; A status marking unit, used for classifying and marking the data according to the data update time to form a data status marking set; A version control unit connected to the state marking unit, used to assign a unique version number to each data state in the state marking data set, and maintain version change records to form version management information; A state synchronization unit is connected to the version control unit and is used to synchronize a number of version management information into the data state tree according to the version management information and a preset time window to form a versioned data set.
7. The data bus based on multi-source remote sensing data processing according to claim 6, characterized in that: The state synchronization unit comprises: a change value calculation subunit, used to calculate the difference between the timestamps in the data state of the current version and the previous adjacent version in the version management information to form a data state change value; The state synchronization subunit is connected to the change value calculation subunit, and is used to synchronize the version management information to the data state tree to form a versioned data set when the data state change value is greater than the preset time window.
8. The data bus based on multi-source remote sensing data processing according to claim 7, characterized in that: The service building components include: A requirement analysis unit, used to analyze the preset multi-source data processing requirements to obtain functional requirements and interface requirements; A data mapping unit, used to uniformly map and format-convert the original data, standardized data and the version control information in the versioned data set to form a data mapping table; An interface design unit, which is connected to the demand analysis unit and the data mapping unit respectively, and is used to design the unified service interface according to the functional requirements, the interface requirements, the data mapping table and the preset business requirements to form a service interface design document; An interface encapsulation unit is connected to the interface design unit and is used to perform encapsulation according to the service interface design document to construct the unified service interface and form the data service interface.
9. The data bus based on multi-source remote sensing data processing according to claim 8, characterized in that: The data service components include: Demand identification unit, used to identify the user's query, browsing, statistics, downloading, subscription and customization needs, and form a user demand set based on the data request input by the user; A demand processing unit, connected to the demand identification unit, for processing and screening the original data, the standardized data and the version control information to form a processed data set; A data interaction unit is connected to the demand processing unit to provide the processed data set through a constructed service interface to form a unified data service.
10. The data bus based on multi-source remote sensing data processing according to claim 9, characterized in that: The unified data service includes general data service and dedicated data service: The general data services include metadata services, remote sensing data publishing services, remote sensing data subscription services and remote sensing data retrieval services; The dedicated data services include auxiliary data services, calibration data services, optical remote sensing image data services and multi-spectral remote sensing image services.
Citation Information
Patent Citations
Space-time data lake management system based on multi-source remote sensing data and safety protection method thereof
CN116737854A
GIS (Geographical Information System) interface platform as well as network GIS management system and management method
CN101763347A
Spatial data storage management system based on big data storage architecture
CN110019089A
Method for designing data warehouse in earth observation field
CN118626469A
Multi-source heterogeneous remote sensing data organization and management method and system, medium and computer program product
CN118885544A