Aerospace big data intelligent data governance system and method based on four-database collaboration and knowledge generation
By constructing a space-air big data governance system that integrates four databases and generates knowledge, the dynamic adaptability and intelligence of the space-air big data governance system have been solved. This system enables the automated transformation of data into knowledge and the reverse empowerment of knowledge, thereby improving the efficiency and adaptability of data governance and forming an intelligent closed loop of self-learning and optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-04-10
AI Technical Summary
Existing aerospace big data governance systems lack the ability to dynamically adapt to new data characteristics and business needs, have insufficient knowledge assetization, limited retrieval intelligence, and lack multimodal feature fusion and semantic understanding capabilities, resulting in high data governance costs, poor real-time response, and the failure to form a complete data-to-knowledge reverse empowerment closed loop.
By employing a four-repository collaboration and knowledge generation approach, a temporary repository, a controlled repository, a product repository, and a knowledge repository are constructed. Combined with data governance, productization, and a knowledge generation engine, automated data cleaning, transformation, and quality checks are achieved, generating and accumulating knowledge assets. The intelligent empowerment engine enables intelligent and adaptive optimization of the data governance process.
It enhances the adaptability and intelligence of data governance, realizes the automated transformation of data from raw data to knowledge, reduces manual intervention, improves the response speed and accuracy of multi-dimensional queries, reduces the configuration time for new data processing, and lowers maintenance costs.
Smart Images

Figure CN121834012A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of aerospace big data governance technology, and in particular to an intelligent data governance system and method for aerospace big data based on four-database collaboration and knowledge generation. Background Technology
[0002] Current aerospace big data governance mainly adopts architectures such as data lakes and data middleware, which have achieved certain results in data aggregation and basic management. However, the following technical bottlenecks still exist in practical applications:
[0003] First, there is the issue of rigid data governance processes. Existing data governance rules are mostly pre-defined static rules, employing a hard-coded "if-then" model, lacking the ability to dynamically adapt to new data characteristics and business needs. For example, when new satellite sensor data is added to the database, the system cannot automatically adapt to the corresponding processing flow, requiring manual reconfiguration of data governance rules. Second, the degree of knowledge assetization is insufficient. Valuable experience gained during data processing (such as optimal processing parameters, effective quality inspection rules, and successful business models) exists mostly in the form of documents or tacit expert knowledge, failing to be transformed into structured knowledge assets that are understandable and callable by machines. This results in the inability to scalably reuse data governance knowledge, leading to high data governance costs. Third, the level of intelligent retrieval is limited. Traditional remote sensing data retrieval methods mainly rely on keyword matching or single visual features, lacking multimodal feature fusion and semantic understanding capabilities. When dealing with the multi-dimensional query needs of aerospace big data, recall and precision are insufficient to meet practical application requirements. Fourth, the business response time is poor. The transformation from raw data to usable business products requires multiple independent steps. There is a lack of intelligent recommendations and automated pipelines based on historical knowledge, which cannot support real-time or near-real-time decision-making needs.
[0004] While some existing methods attempt to combine convolutional neural networks with image retrieval techniques to improve remote sensing image classification or utilize attention mechanisms to enhance multi-label retrieval accuracy, these techniques fail to form a complete closed loop from data to knowledge, and then use knowledge to empower data governance. They also lack dedicated indexes and retrieval mechanisms for the multi-dimensional characteristics of aerospace big data and fail to address the issues of structured accumulation and adaptive reuse of data governance knowledge. Summary of the Invention
[0005] To address the aforementioned problems, the present invention aims to provide an intelligent data governance system and method for aerospace big data based on four-database collaboration and knowledge generation. This system enables the transformation and reuse of implicit knowledge in data governance into explicit knowledge, and enhances the adaptability and intelligence of the data governance process through knowledge reverse empowerment. Ultimately, it forms an intelligent closed-loop data governance system capable of self-learning, self-optimization, and continuous adaptation to business needs.
[0006] This invention provides an intelligent data governance system and method for aerospace big data based on four-database collaboration and knowledge generation.
[0007] The first aspect: A smart data governance system for aerospace big data based on the collaboration of four databases and knowledge generation, comprising a data layer, a core layer, a service layer, and an application layer, wherein:
[0008] The data layer is used to import multi-source heterogeneous aerospace data, including satellite remote sensing, navigation and positioning, meteorological data, and geographic information.
[0009] The core layer employs four storage modules: a temporary storage library, a controlled storage library, a product library, and a knowledge base, forming a four-library architecture.
[0010] The service layer employs four engines—data governance engine, productization engine, knowledge generation engine, and intelligent empowerment engine—for data intelligence and data governance.
[0011] The application layer is used for intelligent search portals, product service gateways, and monitoring and analysis platforms.
[0012] In one embodiment of the present invention, in the four-library storage module:
[0013] A temporary storage area is used for raw data storage, metadata parsing, and data caching.
[0014] A controlled library used for standardized data, process data traceability, and version control;
[0015] The product library is used to generate standard products, thematic products, and analytical products.
[0016] The knowledge base includes processing model knowledge, retrieval model knowledge, quality inspection rule knowledge, and business template knowledge.
[0017] In one embodiment of the present invention, the following is stated:
[0018] A data governance engine is used to automate the cleaning, transformation, standardization, and quality checks of data.
[0019] The productization engine provides a configurable product generation pipeline, supporting the rapid production of various data products.
[0020] The knowledge generation engine is responsible for extracting, structuring, and storing knowledge assets from the data governance process.
[0021] The intelligent empowerment engine enables intelligent retrieval, matching, and application of knowledge assets.
[0022] In one embodiment of the present invention, the knowledge asset includes:
[0023] Processing model knowledge to optimize the set of image processing parameters for a specific satellite sensor;
[0024] Retrieve model knowledge, based on a hierarchical retrieval index built from multi-dimensional metadata;
[0025] Knowledge of quality inspection rules; based on a deep learning-based automatic data quality assessment model, intelligent identification of common data quality problems;
[0026] Business template knowledge, including standardized processing procedures and product templates for typical application scenarios.
[0027] The second aspect: A smart data governance method for aerospace big data based on four-database collaboration and knowledge generation, including:
[0028] S1. Based on a four-repository architecture consisting of a temporary repository, a controlled repository, a product repository, and a knowledge base, combined with a data governance engine and a productization engine, data governance is carried out in a collaborative manner.
[0029] S2. Relying on the knowledge generation engine, knowledge assets are automatically generated and accumulated in the data governance process and stored in the knowledge base;
[0030] S3. Leveraging the intelligent empowerment engine, in new data governance tasks, the optimal knowledge assets are matched from the knowledge base, and intelligent empowerment applications are carried out in each link of the data governance process under the drive of knowledge assets.
[0031] S4 provides feedback on the effectiveness of intelligent empowerment applications, continuously verifying and optimizing knowledge assets in the knowledge base.
[0032] In one embodiment of the present invention, the automatic generation and accumulation of knowledge assets in the data governance process in S2 includes: multimodal feature extraction, knowledge structuring processing, and intelligent retrieval construction, wherein:
[0033] The multimodal feature extraction includes: visual feature extraction, metadata feature extraction, and spatiotemporal feature extraction.
[0034] The knowledge structuring process includes: multi-source feature fusion using discriminant correlation analysis; constructing a mapping relationship between the feature layer and the semantic layer based on association rule mining technology; and establishing semantic associations between knowledge assets using knowledge graph technology.
[0035] The intelligent retrieval system includes: building a multi-level hybrid index based on Elasticsearch; employing multiple retrieval modes such as exact matching, range query, wildcard matching, and semantic search; and dynamically adjusting the index strategy according to the query mode to perform adaptive optimization of the index.
[0036] In one embodiment of the present invention, the intelligent empowerment application of each stage of the data governance process driven by knowledge assets in S3 includes:
[0037] Intelligent lake entry recommendation, automated quality inspection, intelligent product generation, and adaptive search optimization.
[0038] In one embodiment of the present invention, the following is stated:
[0039] Intelligent data entry recommendation includes: real-time parsing of metadata features of newly entered data; matching the optimal preprocessing process from the knowledge base based on similarity calculation; and automatically recommending personalized quality inspection solutions.
[0040] The automated quality inspection includes: calling the AI quality inspection model in the knowledge base to automatically assess data quality; intelligently identifying potential data defects based on historical problem patterns; and generating detailed quality assessment reports and improvement suggestions.
[0041] The intelligent product generation includes: parsing business requirements, automatically matching relevant processing models and business templates; dynamically combining and generating the optimal product production line; and automating and monitoring the product production process in real time.
[0042] The adaptive retrieval optimization includes: intelligently selecting retrieval strategies based on query context; improving retrieval efficiency by using query sharding and index pruning techniques; and supporting multi-dimensional, multi-granular combined queries and semantic queries.
[0043] Third aspect: An electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, performs the steps of the method provided in the second aspect.
[0044] Fourth aspect: A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method provided in the second aspect.
[0045] The beneficial effects of this invention are:
[0046] 1. The system and method of this invention not only realize the automated and intelligent transformation of aerospace big data from raw data to knowledge products, but also significantly improve the efficiency and adaptability of data governance through the continuous accumulation and reverse empowerment of the knowledge base, and are especially suitable for aerospace remote sensing application scenarios with multiple sources, multiple time phases and multiple modes.
[0047] 2. Based on the system and method of this invention, data governance efficiency is improved, the automation level of data processing and product generation is increased, and manual intervention is reduced; retrieval performance is optimized, the response time of multi-dimensional queries is reduced from minutes to seconds, and the recall and precision rates are significantly improved; knowledge reuse is good, the reuse rate of data governance knowledge is high, and the configuration time for new data processing is reduced; the system has strong adaptability, the system can automatically adapt to new data sources and business needs, and maintenance costs are reduced. Attached Figure Description
[0048] Figure 1 This is a framework diagram of the aerospace big data intelligent data governance system based on four-database collaboration and knowledge generation, as described in this invention.
[0049] Figure 2 This is a flowchart illustrating the intelligent data governance method for aerospace big data based on four-database collaboration and knowledge generation, as described in this invention. Detailed Implementation
[0050] Embodiments of the present invention are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar symbols denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0051] Current methods for governing aerospace big data suffer from several shortcomings: failure to form a complete closed loop from data to knowledge, and then using knowledge to empower data governance; lack of dedicated indexes and retrieval mechanisms for the multi-dimensional characteristics of aerospace big data; and unresolved issues regarding the structured accumulation and adaptive reuse of data governance knowledge.
[0052] To address the aforementioned problems, this invention provides an intelligent data governance method for aerospace big data based on four-database collaboration and knowledge generation.
[0053] Example 1:
[0054] like Figure 1 This embodiment discloses an intelligent data governance system for aerospace big data that integrates four databases and knowledge generation, comprising a data layer, a core layer, a service layer, and an application layer.
[0055] The data layer is used to acquire heterogeneous aerospace data from multiple sources, such as satellite remote sensing, navigation and positioning, meteorological data, and geographic information data.
[0056] The core layer employs a four-database storage module, comprising a temporary storage library for receiving raw data, a controlled storage library for processing data, a product library for storing standardized products, and a knowledge base for accumulating and reusing knowledge assets.
[0057] The temporary storage adopts a distributed storage architecture, supports parallel access and buffer management of multi-source heterogeneous aerospace data, and integrates automatic metadata parsing and data deduplication functions for raw data reception and storage, metadata parsing and data caching.
[0058] The controlled library uses dynamic data management based on quality thresholds to achieve version control and traceability of data, and is used for standardized data, process data tracking and version control.
[0059] The product library provides standardized product packaging and service interfaces, supporting rapid product release and business calls, and is used to generate standard products, thematic products, and analytical products.
[0060] The knowledge base adopts a hybrid storage mode of graph database and relational database, which supports multi-dimensional association and complex query of knowledge assets, including processing model knowledge, retrieval model knowledge, quality inspection rule knowledge and business template knowledge.
[0061] The service layer employs four engines: a data governance engine, a productization engine, a knowledge generation engine, and an intelligent empowerment engine.
[0062] The data governance engine enables automated cleaning, transformation, standardization, and quality inspection of multi-source heterogeneous aerospace data.
[0063] The productization engine provides a configurable product generation pipeline, supporting the rapid production of various data products.
[0064] The knowledge generation engine is responsible for extracting, structuring, and storing knowledge assets from the entire data governance process.
[0065] The intelligent empowerment engine enables intelligent retrieval, matching, and application of knowledge assets.
[0066] The application layer is used for intelligent search portals, product service gateways, and monitoring and analysis platforms.
[0067] The intelligent retrieval portal is used for multidimensional queries and semantic searches, the product service gateway is used for product release and API interfaces, and the monitoring and analysis platform is used for data governance dashboards and quality monitoring.
[0068] Using this system architecture, such as Figure 1 As shown, the data flow starts from the underlying multi-source aerospace data, passes through the temporary storage and controlled storage in sequence, and finally forms usable products in the product library, thus forming a standard data governance process.
[0069] Knowledge flow includes knowledge asset generation and knowledge empowerment, among which:
[0070] Knowledge asset generation involves using a knowledge generation engine to extract knowledge assets from the entire data governance process across the temporary storage, controlled storage, and product storage, and then structuring and storing them in the knowledge base. This process... Figure 1 The arrows pointing from each library to the "knowledge generation engine" are used to represent this.
[0071] Knowledge empowerment utilizes an intelligent empowerment engine to retrieve knowledge from the knowledge base, which then acts in reverse on the temporary storage (intelligent lake entry), the controlled storage (AI quality inspection), and the product storage (intelligent production). This process... Figure 1The arrows pointing from the "Intelligent Empowerment Engine" to the various libraries represent the Chinese and Israeli systems.
[0072] Example 2:
[0073] Based on the system in Example 1, this example discloses a smart data governance method for aerospace big data that integrates four-database collaboration and knowledge generation. The workflow diagram is as follows: Figure 2 As shown, taking the "Remote Sensing Active Survey and Three-Dimensional Monitoring Technology for Typical Land Uses of Natural Resources" project as an example, this project focuses on monitoring changes in typical land use types such as forest land, cultivated land, and water bodies in the Qinling Mountains and Guanzhong Plain regions, demonstrating the application of the "Four-Repository Collaboration and Knowledge Generation" mechanism of this invention in a practical scenario. The steps include:
[0074] S1. Based on a four-repository architecture consisting of a temporary repository, a controlled repository, a product repository, and a knowledge base, and combined with a data governance engine and a productization engine, data governance is carried out in a collaborative manner.
[0075] The data governance engine is used to automate the cleaning, transformation, standardization, and quality checks of data, while the productization engine provides a configurable product generation pipeline to support the rapid production of various data products.
[0076] Starting with the input of multi-source aerospace data, the first knowledge empowerment is implemented during the temporary storage stage. The intelligent empowerment engine intervenes for the first time, making intelligent recommendations based on the historical knowledge of the knowledge base and guiding the data into the correct preprocessing process.
[0077] Initial data is stored in a temporary storage library. Multi-source satellite data (including Gaofen series, Landsat and SAR, etc.) and their metadata are stored in the temporary storage library. The data covers multi-modal resources such as optical images and radar data, covering data assets of different time phases and resolutions in the study area.
[0078] Then, the intelligent empowerment-lake entry recommendation process is initiated. The intelligent empowerment engine is triggered, reading the data's metadata (such as data source, sensor type, spatiotemporal range, etc.) and querying the knowledge base. The knowledge base returns the previously verified "multi-source satellite data standard preprocessing package," including steps such as radiometric correction, geometric correction, and cloud detection, and recommends the optimal processing flow applicable to the region. After confirmation by the data administrator, the system automatically executes the preprocessing flow and stores the processed data in a controlled database.
[0079] In the data governance and knowledge generation process, data undergoes further cleaning, fusion, and quality checks within a controlled repository. For example, experts constructed an "AI model for forest change detection (Model_Forest_Change)" and a "farmland extraction template" through manual annotation and model training.
[0080] Simultaneously, the knowledge generation engine automatically extracts key parameters, model structure, applicable conditions (such as data source, region, and time phase), and performance indicators (such as accuracy and recall) from these successful tasks, and structures them into knowledge assets, storing them in the knowledge base. At the same time, the system automatically updates the "multimodal sample library," which includes optical and SAR feature samples of typical land types.
[0081] Then, intelligent empowerment and product production are carried out. When business users put forward the "2024 Forest Land Change Monitoring Report in Qinling Area", the intelligent empowerment engine intelligently retrieves and matches existing "Forest Land Change Detection AI Model", "Multi-source Data Fusion Rules" and "Natural Resource Change Report Template" from the knowledge base according to the requirement description (target land type: forest land, region: Qinling, data source: GF-2, SAR).
[0082] The system automatically assembles the product production workflow: acquire the latest images → call the change detection model → fuse multi-source features → generate change patches → fill in the report template, and finally form a standardized data product and store it in the product library.
[0083] S2. Relying on the knowledge generation engine, knowledge assets are automatically generated and accumulated in the data governance process and stored in the knowledge base.
[0084] Through a knowledge generation engine, automatic knowledge extraction and structuring are achieved throughout the entire data governance process. The multimodal feature extraction includes:
[0085] Visual feature extraction: The improved VGGNet-16 model is used to extract depth features from remote sensing images, and combined with traditional visual features (color, texture, shape) to form a comprehensive feature representation; Metadata feature extraction: Key metadata features are extracted from the header file and auxiliary information of the data file; Spatiotemporal feature extraction: Spatiotemporal feature vectors are constructed based on geocoding and temporal information.
[0086] Knowledge structuring includes: using discriminant relevance analysis (DCA) to fuse multi-source features and improve the discriminative ability of features; constructing a mapping relationship between the feature layer and the semantic layer based on association rule mining technology to narrow the semantic gap; and using knowledge graph technology to establish semantic associations between knowledge assets to support complex reasoning and querying.
[0087] Intelligent retrieval construction includes: building multi-level hybrid indexes based on Elasticsearch, including inverted indexes, FST (Finite State Transformer), and BKD-Tree; supporting multiple retrieval modes such as exact matching, range queries, wildcard matching, and semantic search; and achieving adaptive optimization of the index, dynamically adjusting the index strategy according to the query mode.
[0088] The process begins with the input of multi-source aerospace data into the database. During the temporary storage stage, the intelligent empowerment engine intervenes for the first time, making intelligent recommendations based on historical knowledge in the knowledge base to guide the data into the correct preprocessing flow. In the controlled database and product database stages, the knowledge generation engine continues to work, extracting knowledge from successful data processing tasks and products and storing it in the knowledge base.
[0089] S3. Leveraging the intelligent empowerment engine, in new data governance tasks, the optimal knowledge assets are matched from the knowledge base, and intelligent empowerment applications are carried out in each link of the data governance process under the drive of knowledge assets.
[0090] Through the intelligent empowerment engine, knowledge can be accurately applied in all aspects of data governance, including intelligent lake entry recommendation, automated quality inspection, intelligent product generation, and adaptive retrieval optimization.
[0091] The intelligent data entry recommendation includes: real-time parsing of metadata features of newly entered data, matching the optimal preprocessing process from the knowledge base based on similarity calculation, and automatically recommending personalized quality inspection solutions.
[0092] Automated quality checks include: calling AI quality inspection models in the knowledge base to automatically assess data quality, intelligently identifying potential data defects based on historical problem patterns, and generating detailed quality assessment reports and improvement suggestions.
[0093] Intelligent product generation includes: parsing business requirements, automatically matching relevant processing models and business templates; dynamically combining and generating the optimal product production line; and realizing automated execution and real-time monitoring of the product production process.
[0094] Adaptive retrieval optimization: Intelligently selects retrieval strategies based on query context, and improves retrieval efficiency by using query sharding and index pruning techniques. It supports multi-dimensional and multi-granular combined queries and semantic queries.
[0095] S4 provides feedback on the effectiveness of intelligent empowerment applications, continuously verifying and optimizing knowledge assets in the knowledge base.
[0096] A closed loop is formed. Figure 2 The dashed line clearly indicates the core closed loop: knowledge from the knowledge base is invoked by the intelligent empowerment engine to guide data governance of new data. Business applications and feedback also serve as knowledge sources, feeding back into the knowledge generation engine, enabling the knowledge system to continuously optimize based on actual results, forming a closed loop of learning and evolution.
[0097] Closed-loop optimization involves the knowledge generation engine recording the entire product development process, including model execution results and user feedback. Based on this feedback, the system automatically updates the applicability score and usage frequency of knowledge assets, optimizing subsequent recommendation strategies. When similar monitoring tasks are received again, the system can respond more quickly and accurately, forming a self-evolving closed loop of "data-driven knowledge accumulation and knowledge-enabled data governance."
[0098] Invention solution:
[0099] First, a collaborative data governance architecture of four repositories is constructed: the four repositories include a temporary storage repository for receiving raw data, a controlled repository for processing process data, a product repository for storing standardized products, and a knowledge repository for accumulating and reusing knowledge assets.
[0100] Secondly, a knowledge generation and accumulation mechanism is implemented: In the data processing flow from the controlled database to the product database, the system automatically extracts and structures the successful experiences to form processing models, quality inspection rules, business templates, etc., and stores them in the knowledge base.
[0101] Then, a knowledge reverse empowerment mechanism is implemented. When a new data processing task is initiated, the system uses an intelligent recommendation engine to match and call the best knowledge assets from the knowledge base, and applies them automatically or semi-automatically to data entry inspection, quality assessment and product production.
[0102] Finally, closed-loop optimization, through continuous application and feedback, ensures that knowledge assets are constantly verified, optimized, and enriched, thereby continuously improving the intelligence level and efficiency of the entire data governance system.
[0103] The provided intelligent governance methods and systems for aerospace big data enable the transformation and reuse of tacit knowledge in data governance into explicit knowledge. Through knowledge-based reverse empowerment, the adaptability and intelligence of the governance process are improved, thereby forming an intelligent closed-loop data governance system that can learn, optimize, and continuously adapt to business needs.
[0104] This invention also provides an electronic device, which may include: a processor, a communications interface, a memory, and a communication bus, wherein the processor, the communications interface, and the memory communicate with each other via the communication bus. The processor can invoke logical instructions stored in the memory, for example, to execute the following method:
[0105] S1. Based on a four-repository architecture consisting of a temporary repository, a controlled repository, a product repository, and a knowledge base, combined with a data governance engine and a productization engine, data governance is carried out in a collaborative manner.
[0106] S2. Relying on the knowledge generation engine, knowledge assets are automatically generated and accumulated in the data governance process and stored in the knowledge base;
[0107] S3. Leveraging the intelligent empowerment engine, in new data governance tasks, the optimal knowledge assets are matched from the knowledge base, and intelligent empowerment applications are carried out in each link of the data governance process under the drive of knowledge assets.
[0108] S4 provides feedback on the effectiveness of intelligent empowerment applications, continuously verifying and optimizing knowledge assets in the knowledge base.
[0109] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0110] This invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments, including, for example:
[0111] S1. Based on a four-repository architecture consisting of a temporary repository, a controlled repository, a product repository, and a knowledge base, combined with a data governance engine and a productization engine, data governance is carried out in a collaborative manner.
[0112] S2. Relying on the knowledge generation engine, knowledge assets are automatically generated and accumulated in the data governance process and stored in the knowledge base;
[0113] S3. Leveraging the intelligent empowerment engine, in new data governance tasks, the optimal knowledge assets are matched from the knowledge base, and intelligent empowerment applications are carried out in each link of the data governance process under the drive of knowledge assets.
[0114] S4 provides feedback on the effectiveness of intelligent empowerment applications, continuously verifying and optimizing knowledge assets in the knowledge base.
[0115] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0116] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A smart data governance system for aerospace big data based on four-database collaboration and knowledge generation, characterized in that, It includes a data layer, a core layer, a service layer, and an application layer, among which: The data layer is used to import multi-source heterogeneous aerospace data, including satellite remote sensing, navigation and positioning, meteorological data, and geographic information. The core layer employs four storage modules: a temporary storage library, a controlled storage library, a product library, and a knowledge base, forming a four-library architecture. The service layer employs four engines—data governance engine, productization engine, knowledge generation engine, and intelligent empowerment engine—for data intelligence and data governance. The application layer is used for intelligent search portals, product service gateways, and monitoring and analysis platforms.
2. The system according to claim 1, characterized in that, In the four-database storage module: A temporary storage area is used for raw data storage, metadata parsing, and data caching. A controlled library used for standardized data, process data traceability, and version control; The product library is used to generate standard products, thematic products, and analytical products. The knowledge base includes processing model knowledge, retrieval model knowledge, quality inspection rule knowledge, and business template knowledge.
3. The system according to claim 1, characterized in that, The following is stated: A data governance engine is used to automate the cleaning, transformation, standardization, and quality checks of data. The productization engine provides a configurable product generation pipeline, supporting the rapid production of various data products. The knowledge generation engine is responsible for extracting, structuring, and storing knowledge assets from the data governance process; The intelligent empowerment engine enables intelligent retrieval, matching, and application of knowledge assets.
4. The system according to claim 3, characterized in that, The knowledge assets include: Processing model knowledge to optimize the set of image processing parameters for a specific satellite sensor; Retrieve model knowledge, based on a hierarchical retrieval index built from multi-dimensional metadata; Knowledge of quality inspection rules; based on a deep learning-based automatic data quality assessment model, intelligent identification of common data quality problems; Business template knowledge, including standardized processing procedures and product templates for typical application scenarios.
5. A smart data governance method for aerospace big data based on four-database collaboration and knowledge generation, characterized in that, The system according to any one of claims 1 to 4 comprises: S1. Based on a four-repository architecture consisting of a temporary repository, a controlled repository, a product repository, and a knowledge base, combined with a data governance engine and a productization engine, data governance is carried out in a collaborative manner. S2. Relying on the knowledge generation engine, knowledge assets are automatically generated and accumulated in the data governance process and stored in the knowledge base; S3. Leveraging the intelligent empowerment engine, in new data governance tasks, the optimal knowledge assets are matched from the knowledge base, and intelligent empowerment applications are carried out in each link of the data governance process under the drive of knowledge assets. S4 provides feedback on the effectiveness of intelligent empowerment applications, continuously verifying and optimizing knowledge assets in the knowledge base.
6. The method according to claim 5, characterized in that, In S2, knowledge assets are automatically generated and accumulated during the data governance process, including: multimodal feature extraction, knowledge structuring processing, and intelligent retrieval construction, wherein: The multimodal feature extraction includes: visual feature extraction, metadata feature extraction, and spatiotemporal feature extraction; The knowledge structuring process includes: multi-source feature fusion using discriminant correlation analysis; constructing a mapping relationship between the feature layer and the semantic layer based on association rule mining technology; and establishing semantic associations between knowledge assets using knowledge graph technology. The intelligent retrieval system includes: building a multi-level hybrid index based on Elasticsearch; employing multiple retrieval modes such as exact matching, range query, wildcard matching, and semantic search; and dynamically adjusting the index strategy according to the query mode to perform adaptive optimization of the index.
7. The method according to claim 5, characterized in that, The intelligent empowerment applications in each stage of the data governance process driven by knowledge assets in S3 include: Intelligent lake entry recommendation, automated quality inspection, intelligent product generation, and adaptive search optimization.
8. The method according to claim 7, characterized in that, The following is stated: Intelligent data entry recommendation includes: real-time parsing of metadata features of newly entered data; matching the optimal preprocessing process from the knowledge base based on similarity calculation; and automatically recommending personalized quality inspection solutions. The automated quality inspection includes: calling the AI quality inspection model in the knowledge base to automatically assess data quality; intelligently identifying potential data defects based on historical problem patterns; and generating detailed quality assessment reports and improvement suggestions. The intelligent product generation includes: parsing business requirements, automatically matching relevant processing models and business templates; dynamically combining and generating the optimal product production line; and automating and monitoring the product production process in real time. The adaptive retrieval optimization includes: intelligently selecting retrieval strategies based on query context; improving retrieval efficiency by using query sharding and index pruning techniques; and supporting multi-dimensional, multi-granular combined queries and semantic queries.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in claim 5.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method as described in claim 5.
Citation Information
Patent Citations
Large-model-based manufacturing enterprise data medium system and method
CN120296095A
Intelligent knowledge base system based on multi-domain model collaboration and construction method
CN120317345A
Convergence management system and method for multi-source remote sensing big data
CN120448445A
Application system integrated management method based on artificial intelligence
CN120763171A
A computer-implemented system and method for managing product development projects
WO2025079008A1