A system for constructing a multi-constraint aerospace data element circulation network based on a large model
Patent Information
- Application Number
- CN202610826060.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-06-09
AI Technical Summary
[0014]为解决以上技术问题,本发明提出一种基于大模型的多约束空天数据要素流通网络构建系统,通过引入真实证据数据约束、结构化网络Schema与拓扑约束,并结合智能化生成与自动校验修复机制,以解决现有技术路线的问题,从而获得连通、细粒度、可追溯且可用于质量传播分析的全生命周期数据要素流通网络
[0021] 1. Enhance the authenticity and auditability of the network:
Smart Images

Figure CN122366526B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology and data elements, specifically involving a system for constructing a multi-constraint aerospace data element circulation network based on a large model. Background Technology
[0002] With the rapid development of multi-source aerospace observation systems, including Earth observation satellites, airborne remote sensing, UAVs, and ground sensors, the daily volume of remote sensing data is growing exponentially. Simultaneously, market-oriented reforms in data element allocation are driving the transformation of data from a "resource" to a "factor." Data is no longer merely a research input but circulates among multiple entities as a factor that can be identified, priced, traded, reused, and regulated. For aerospace data, this factor circulation spans the entire lifecycle of "data production - processing - derivative products - application services - reproduction - storage management," and is highly coupled with algorithm models, research teams, and service institutions, leading to increasingly complex and automated processing flows. To meet the aforementioned need for constructing a data element circulation network, existing technologies are primarily exploring two technical routes, but both have significant technical limitations.
[0003] Building a roadmap based on information extraction from system logs:
[0004] In recent years, a large number of scientific workflow systems have emerged in the field of geographic information processing, such as GEE and OpenEO. These systems encapsulate data and routine processing flows within a closed software system. Algorithm functions are encapsulated as independent computation function modules, and users construct processing flows by defining "data flows" between modules. When users perform data processing operations within the system, the system automatically captures the execution logs of each computation module in the background management module, including its input data, output data, executed module or algorithm, running parameters, and timestamps. These logs are stored in the system's backend database, and users can download and query them through the system's provided interface. After obtaining the log information, domain experts write data processing scripts to map the complex log information into a data element flow network.
[0005] However, these log records are typically discrete (each processing stage generates independent records), heterogeneous (different systems or tools generate logs in inconsistent formats), and have a low semantic level (primarily describing static dependencies, lacking representations of deeper semantics such as quality and context), and are largely factual (only recording operations that have occurred). More importantly, the runtime logs of scientific workflow systems often belong to the platform's backend maintenance data or internal assets. Influenced by factors such as access control, business strategies, privacy compliance, and security audits, ordinary users can usually only obtain limited task status information or result metadata, making it difficult to obtain complete, fine-grained execution logs. This is even more challenging in cross-team and cross-institutional collaboration scenarios, where log sharing and unified access are difficult to achieve. Therefore, although system logs can theoretically serve as the foundational data source for building a circulation network, in practical applications they often face the inherent constraint of "logs being unavailable or incomplete." Even if they can be obtained, issues such as the unified standardization of multi-source logs, semantic enhancement, and global connectivity modeling still need to be addressed to efficiently and automatically integrate low-level, scattered data sources into a unified, connected, and advanced intelligent analysis-supporting global network model.
[0006] Knowledge graph-based entity extraction construction roadmap:
[0007] In recent years, entity extraction based on knowledge graphs has been a research hotspot in related fields. This approach uses unstructured or semi-structured documents such as academic papers, technical reports, and standard documents as the main information source. Through natural language processing and information extraction techniques, the implicit "entity-relationship-entity" relationships in the text are transformed into structured triples. Then, combined with domain ontology, these relationships are normalized, aligned, and integrated to ultimately construct a domain knowledge graph. Specifically, in constructing a space-air element circulation network, experts often need to collect a large number of "data product papers" and related papers that cite them. Entity extraction is performed on each paper, forming small, isolated data circulation clusters. Then, based on the citation relationships of the papers, entities are merged and integrated to ultimately construct the data element circulation network.
[0008] However, the primary purpose of academic papers, technical reports, and standard documents is to elucidate "methodological innovations and experimental results." Their representation of the data processing chain is typically selective, abstract, and unstructured, leading to sparse, fragmented, and unreproducible relationships obtained from entity extraction. Specifically:
[0009] Key intermediate products and sub-steps are often omitted or merged. In papers, many steps in the remote sensing processing chain that determine the direction of quality propagation are often summarized in a single sentence, such as "image preprocessing was performed," "radiometric / geometric correction was completed," or "cloud shadow effects were removed." Entity extraction often only yields a single "preprocessing" node, failing to identify its internal sub-steps such as radiometric calibration, atmospheric correction, cloud detection / masking, registration, resampling, cropping, and temporal synthesis, as well as their inputs and outputs. Ultimately, this leads to a blurred data processing flow, making it impossible to accurately locate data outliers.
[0010] The system suffers from several shortcomings. Firstly, it lacks clear roles and operational environments within the processing chain. Secondly, academic texts often only implicitly define the "subject" in the author's affiliation information, failing to specify the platform used or whether third-party services are invoked. Thirdly, the extraction system tends to simplify the "subject" to simply "paper author." Furthermore, it cannot construct the necessary division of roles for factor circulation, including production, processing, service, and regulatory bodies, thus hindering subsequent needs for rights confirmation, liability determination, and audit traceability.
[0011] The lack of quality information is a significant issue. The aerospace data element circulation network needs to describe not only the flow relationships between data elements but also support the analysis of how data quality changes and the root causes of quality problems. However, the expression of data product quality in academic texts is often inconsistent. Therefore, even if entity extraction based on literature can yield some quality indicators, it is often difficult to stably map them to node-level quality fields in the "data-algorithm-subject" hierarchy, let alone construct a computable closed loop for quality propagation and root cause tracing.
[0012] In summary, both of the existing typical approaches face insurmountable obstacles in practical implementation: On the one hand, while the approach based on scientific workflow system logs can theoretically provide relatively complete input / output, parameter, and execution sequence information, such background logs are usually limited by factors such as platform access control, privacy compliance, and security auditing. Currently, the construction and application of my country's aerospace element data circulation platform are still in their initial stages, and a process circulation data and interface that is open, transparent, and queryable to users has not yet been formed. Complete and fine-grained log data cannot be obtained, resulting in the constraint of "data unavailability or incompleteness" in the practice of aerospace element circulation in my country. On the other hand, while knowledge graph entity extraction that relies entirely on academic literature is easy to obtain text resources and covers a wide range of topics, what it extracts is mostly "conceptual-level relationships." That is, it can usually identify high-level conceptual relationships such as "using a certain type of data," "adopting a certain type of method / model," and "oriented towards a certain type of application task," but it is difficult to recover the step-level input / output data products in the data processing chain. In contrast, the aerospace data element circulation network not only needs to describe conceptual relationships but also needs to meet the additional requirement of "engineering-level data element lineage," processing fine-grained decomposition of algorithms and clearly defined input and output data products; the network needs to include a branching and crossing topology structure under multi-source data fusion and multi-task reuse; and a quality field system that can be uniformly mapped to nodes, along with its propagation and tracing mechanisms. When key intermediate activities, role elements, and quality information are lacking, the constructed network is prone to problems such as broken links, non-closed topologies, and excessively coarse granularity, making it difficult to form a computable closed loop for quality propagation and root cause tracing, thus hindering subsequent research and application of aerospace data element circulation.
[0013] Therefore, there is an urgent need for a method to construct a remote sensing data element circulation network for scenarios where "evidence is available but incomplete," to effectively support subsequent advanced intelligent analysis applications such as quality propagation modeling and root cause tracing. This method must, under the premise of "real-world evidence being available but incomplete," ensure the credibility and traceability of key data starting points and core facts, while also enabling controllable completion and structured generation of missing links. This would result in a circulation network that is topologically connected, computationally granular, and semantically interpretable. Driven by these needs, this invention proposes a method that, based on real data anchors extracted from literature knowledge and prior knowledge, simulates and automatically verifies and repairs the network under explicit rule constraints, thereby obtaining a high-quality aerospace data element circulation network that combines authenticity and completeness. Summary of the Invention
[0014] To address the above technical problems, this invention proposes a multi-constraint aerospace data element circulation network construction system based on a large model. By introducing real-world evidence data constraints, structured network schemas, and topological constraints, and combining them with intelligent generation and automatic verification and repair mechanisms, this system solves the problems of existing technical approaches, thereby obtaining a connected, fine-grained, traceable, and full-lifecycle data element circulation network that can be used for quality propagation analysis. The specific technical solution is as follows:
[0015] A system for constructing a multi-constraint aerospace data element circulation network based on a large model includes:
[0016] The knowledge base module is used to store and output structured indexes. The input of the knowledge base module receives a literature database and an algorithm database. The knowledge base module converts coarse-grained algorithm names into fine-grained algorithm steps and associates them with research topics and data elements, outputting a structured index of research topic-data element-algorithm step set.
[0017] The rule constraint module, set up in parallel with the knowledge base module, is used to define and output the rule constraint set. The rule constraint module abstracts the rules into a unified rule constraint set.
[0018] The network generation module, whose input end is connected to the output end of the knowledge base module and the output end of the rule constraint module respectively, is used to perform structured graph construction within the feasible space defined by rules based on the retrieved prior knowledge, real data and structured index.
[0019] The inspection and correction module has its input end connected to the output end of the network generation module, and its output end is fed back to the network generation module, forming a modular closed-loop architecture. The inspection and correction module uses rule verification and semantic verification as the criteria to improve the circulation network output by the network generation module to an engineering usable state.
[0020] The present invention has the following beneficial effects:
[0021] 1. Enhance the authenticity and auditability of the network:
[0022] This invention uses real data anchors as the starting nodes of each connected subgraph and binds evidence pointers (source documents / entry fragments) to the anchors, while prohibiting the generation of original sensor or product names out of thin air. This significantly reduces the risk of unverifiable data sources, making key network facts verifiable and auditable, and facilitating their use in subsequent rights confirmation and regulatory scenarios.
[0023] 2. Improve the granularity of data processing workflows and anomaly localization capabilities:
[0024] By introducing a priori fragment library of processes and setting granularity rules (such as preprocessing must be broken down into multiple sub-steps and fusion must be explicitly aligned), the network generated by this invention can express step-level input-output boundaries, rather than remaining at conceptual nodes such as "preprocessing / modeling". This effect enables data anomalies or quality problems to be located at specific stages (such as cloud detection, registration, resampling, sample construction, etc.), improving the efficiency of problem investigation and reproduction.
[0025] 3. Enhance network topology connectivity and reusable expressive capabilities:
[0026] This invention forcibly introduces branch expansion and cross-reuse structures during the generation phase, and verifies and corrects the "number of reused nodes and branch cross-reuse structure" during the verification phase, thereby preventing the network from degenerating into linear chains or isolated clusters. This effect is more in line with the real-world patterns of multi-source fusion and multi-task reuse in remote sensing operations, and is beneficial for path querying, dependency analysis, and reuse evaluation.
[0027] 4. Ensure consistent structured output for easy data import and graph calculation:
[0028] This invention constrains fields, reference closure, ID uniqueness, and multiple inputs with single output through a rule module, and enforces these constraints during rule validation. This results in a stable data structure that can be directly written into a graph database or document database, supporting graph computation tasks such as source tracing, critical path analysis, and impact assessment, thereby reducing subsequent engineering integration costs.
[0029] 5. Provides a computational basis for quality communication and root cause analysis:
[0030] This invention injects local quality parameters and directional influence descriptions into data nodes, processing activity nodes, and subject nodes during the network construction phase, and decouples quality propagation calculation from the generation process, making quality propagation a repeatable and interpretable computational process. This allows for the execution of quality propagation, anomaly backtracking, and attribution analysis on the network, linking quality issues to specific data versions, specific processing steps, or specific implementing entities, providing support for quality monitoring and accountability.
[0031] 6. Reduce cross-platform dependencies and improve applicability and scalability:
[0032] Compared to approaches that rely solely on closed workflow system logs, this invention does not require complete background logs to function. Furthermore, compared to plain text entity extraction, this invention supplements missing links and forms a computable structure through prior fragments and rule constraints. Therefore, it is applicable to multiple platforms, teams, and thematic scenarios, facilitating continuous expansion to different application areas such as ecology, cities, vegetation, and atmosphere, and supporting the introduction of new data sources and new algorithmic processes.
[0033] 7. Improve the stability and usability of results by verifying and correcting the closed loop:
[0034] This invention employs a closed-loop mechanism of "rule hard verification + semantic review + minimal modification and repair," enabling the network to gradually converge from the initial generated result to a structurally sound and semantically reasonable version. This reduces the cost of repeated manual image editing and parameter tuning, and improves the stability and engineering usability of automated construction. Attached Figure Description
[0035] Figure 1 This is the element circulation network data structure of the present invention;
[0036] Figure 2 This is the overall system architecture of the present invention. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above objectives, this invention adopts the following technical solution.
[0038] For ease of explanation, this implementation plan uses the PROV approach to abstract the aerospace data element circulation network into three types of nodes and performs structured coding, such as... Figure 1 As shown.
[0039] Data nodes: Represent aerospace data elements, including "real data" and "derived data". Each Data node contains semantic attributes and an "initial quality score" field, which are used for subsequent quality propagation modeling.
[0040] Activity nodes represent algorithms or processing activities. Each Activity node must explicitly list the input and output data lists and adhere to the "multiple inputs, single output" structural constraint to break down complex processes into computable atomic activities. Each Activity node contains semantic attributes and an "initial quality score" field for subsequent quality propagation modeling.
[0041] Agent node: Represents the participating entity (roles such as production, processing, service, and governance), and provides local quality attributes for the entity's capabilities, which are used to subsequently associate the responsible entity with the quality impact.
[0042] The network output uses a unified JSON file format, with the top-level keys fixed as Data, Activities, and Agent, and other keys are prohibited to ensure that the structure can be directly inserted into the database and parsed by scripts.
[0043] This invention addresses the automated construction of a "space-air data element circulation network." The core challenge lies in two aspects: firstly, the network must possess engineering-grade computable lineage (fine-grained activities, explicit inputs and outputs, topological connectivity, and computable data quality); secondly, input information often originates from heterogeneous and incomplete external sources (fragments of closed workflow logs, abstracted expressions from papers / standard texts), making direct extraction insufficient for practical needs. Therefore, this invention employs a modular closed-loop architecture to achieve real-world data constraints, prior knowledge completion, and collaborative rule constraints, generation, and verification. This transforms the large-scale Language Model (LLM) from a random data generator into a controlled graph builder and semantic reviewer, ultimately outputting a data element circulation network with a fixed structure that can be directly imported into a database and used for graph computation. The overall structure of this invention's large-model-based multi-constraint space-air data element circulation network construction system is as follows: Figure 2 As shown.
[0044] 1. Knowledge Base Module:
[0045] The knowledge base module is designed to address issues such as insufficient authenticity and coarse granularity in data element circulation networks generated directly using traditional methods. The knowledge base module requires input from at least two types of data.
[0046] The first type is a literature database. This database uses web crawling technology to crawl data product papers related to geographic information science and their cited papers, extracting information such as research topics, data elements, and author affiliations. The extracted information is not stored randomly in the database but is interconnected. Specifically, the system extracts data element names, key attribute words, and author affiliations from paragraphs such as abstracts, data and methods, and experimental settings. It also calculates a confidence score for each data element (e.g., location weight, frequency of occurrence across documents, author authority) and binds it to a corresponding set of research topics. This confidence score will be used as the basis for subsequent data constraints during generation.
[0047] The second category is algorithm databases. In academic writing, authors often don't specifically describe the processing of simple data or the execution of algorithms, which is why traditional methods for building circulation networks cannot achieve fine-grained detail. To overcome this problem, this invention collects and organizes prior knowledge of common aerospace data element processing procedures. For example, preprocessing includes steps such as "radiometric calibration," "atmospheric correction," "cloud detection," "registration," "resampling," and "temporal synthesis," while change detection includes steps such as "geometric registration," "normalization," "semantic segmentation," "instance segmentation," and "temporal analysis." After converting coarse-grained algorithm names into fine-grained algorithm steps, it is also necessary to establish a connection with the research topics and data elements extracted from the paper.
[0048] Ultimately, the knowledge base module outputs a structured index of "research topic - data elements - algorithm step set", which can be used to enhance retrieval during the generation stage. For example, given a topic T, it can retrieve high-confidence real data D and the corresponding algorithm step set. Large models will prioritize information input from the knowledge base, thereby improving the authenticity and granularity of the circulation network.
[0049] 2. Rule Constraint Module:
[0050] The rule constraint module addresses issues such as misaligned attributes, unquantifiable quality, and insufficient network topology in data element circulation networks generated by traditional methods. Instead of simply pasting text into prompts, the module abstracts rules into a unified set of constraints, driving the larger model to generate high-quality data circulation networks. This set of constraints contains at least four categories.
[0051] 2.1 Structure and Field Rules:
[0052] This rule defines the required fields, types, enumeration ranges, allowed null values, and dependencies between fields for each Data, Activity, and Agent node. Specific constraints are shown in Table 1.
[0053] Table 1. Data structure rules and constraints for nodes
[0054]
[0055]
[0056] 2.2 Network topology and scalability constraints:
[0057] A true aerospace data element circulation network should be mesh-like, not chain-like. Topology rules are used to ensure the network is a graph structure, not a chain structure, and conforms to the typical form of remote sensing processing: each algorithm node is allowed to have multiple data nodes as input, but can only output one data node (if this is not met, the algorithm node is split into steps until it is); one data node can serve as input to multiple algorithm nodes, i.e., data reuse; one entity can execute multiple algorithm nodes. Furthermore, to ensure the scale of the circulation network, this module also explicitly specifies the total range of nodes and the approximate proportions of the three types of nodes.
[0058] 2.3 Quality Constraint Rules:
[0059] To achieve computable quality propagation, this invention represents data quality as a four-dimensional vector (extendable to more dimensions), as shown in formula (1):
[0060] (1)
[0061] in, Describing the accuracy of the data, Describe the integrity of the data. Describe the timeliness of the data. The document completeness of the data is described. For real data extracted from the knowledge base, the initial quality vector is evaluated by the corresponding descriptive text; for derived data, its quality is recursively calculated by the surrounding nodes through the quality propagation formula (5). This operation can ensure that the quality of the data is a repeatable, auditable, and verifiable calculation process, rather than a subjective description of the generative model.
[0062] For each algorithm node, the local quality parameters of the algorithm itself are defined from three dimensions: method rigor. Reproducibility Achieving maturity and the reliability coefficient of its synthesis algorithm. , A higher value indicates that the algorithm is less affected by the uncertainty of the data. The calculation method is shown in formula (2).
[0063] (2)
[0064] in, , , These represent weighted coefficients (defaulting to one-third for each) indicating the algorithm's rigor, reproducibility, and maturity. Additionally, to quantify the algorithm's impact on each dimension of data quality... This invention defines the gain value of the algorithm's impact on data quality. Among them, enumeration This indicates the directional impact of the algorithm on data quality; its value can be found in Table 2.
[0065] Table 2. Impact of Algorithms on Data Quality
[0066]
[0067] Similarly, for each subject node, based on the relevant descriptive information about the subject in the knowledge base, the subject's local quality parameters are defined from three dimensions: professionalism. Institutional Credibility Openness / auditability And synthesize the main body's credibility coefficient , The higher the value, the more reliable the subject is, and the higher the probability of obtaining high-quality data products. The calculation method is shown in formula (3).
[0068] (3)
[0069] in, , , The weighted coefficients representing the subject's professionalism, credibility, and openness (all defaulting to one-third) are specified in this invention. The data for each algorithm must be either multi-input single-output or single-input single-output: the input data set is... Output data is For each dimension First, the quality of the input data is aggregated into "input-side overall quality". Considering the "weakest link effect" in multi-source fusion, the weighted harmonic average can be used to calculate the result (more sensitive to low quality):
[0070] (4)
[0071] in, This indicates the number of data points input to the algorithm. Indicates the first The first data point Quality rating in each dimension Weights for each input data; To prevent division by zero of small constants (such as 1e-6).
[0072] Finally, the quality of the output data is affected by four factors, namely the overall quality of the input side. Algorithm reliability Credibility of the subject And the influence on gain value The quality calculation method for each dimension of the output data is shown in formula (5).
[0073] (5)
[0074] in, This represents the attenuation sensitivity coefficient (configurable by the user). Indicates the gain strength (determined by the type of algorithm).
[0075] 2.4 Output and Reference Consistency Rules:
[0076] The circulation network constructed by this invention does not allow for isolated nodes or isolated edge relationships. The references of each node (inputs / outputs / source_activity / target_activities / Agent.activities) must be mutually closed. Only in this way can the connectivity of the network be ensured, and the ID number must also be unique.
[0077] 3. Network generation module:
[0078] The key to the generation module is to prevent large models from generating data out of thin air. Instead, it involves "structured mapping" based on retrieved prior knowledge and real data, within a rule-defined feasible space. To reduce the error rate of outputting large JSON files all at once, this invention adopts a phased generation strategy.
[0079] 3.1 Large Model Input Prompt Word Set:
[0080] At the start of each task, the generation module constructs a set of prompt words based on the user's instructions (usually the research topic), the rule base, and the knowledge base, which must contain at least the following:
[0081] The research task description should include information such as research content, task type, research area, and time frame, and should clearly state the research results that the task aims to achieve.
[0082] The real data set allows users to specify the real data they want to use, or the set can be automatically selected based on the matching degree between the research topic and the real data in the knowledge base. Each real data set needs to provide authentic and reliable attribute field information as well as quality scores for each dimension with evidence summaries.
[0083] The algorithm process is a collection of fragments. The large model will analyze and understand the research task in advance and plan a suitable technical route. The complex research task will be broken down into several algorithm steps. In order to ensure that the granularity of each step meets the requirements, the system will also match the split algorithm steps with the knowledge base to ensure that the granularity of the technical route is small enough. Finally, the fragments are stored in the prompt words in the form of a list.
[0084] Output format requirements: output only JSON, with each element having a fixed top-level key.
[0085] It is particularly emphasized that the prompt will explicitly require that "the initial data nodes of each connected subgraph can only be selected from real data, and large models are not allowed to create their own sensor / product names." This is to ensure that the quality of the data nodes in the circulation network meets our defined quality constraint rules (simulating the real world), thereby enabling advanced intelligent analysis.
[0086] 3.2. Phased composition:
[0087] For the research topics provided by users, a set of corresponding prompt words is compiled and input into a large model to drive the large model to simulate and generate the corresponding data flow network. At the same time, in order to ensure the quality of the output data, this invention divides the generation process into three stages.
[0088] Phase 1: Skeleton Generation.
[0089] The large model will first automatically plan a technical route that can directly complete the research task based on the real data set and algorithm process fragment set recorded in the prompt words, and output a skeleton diagram containing node IDs and core attributes. This skeleton diagram does not require filling in all semantic and quality field information, but only requires the topology to meet the minimum computable requirements of "single output algorithm" and "reference closure".
[0090] Phase Two: Data Reanalysis and Expansion.
[0091] The skeleton diagram generated in Phase 1 is often just a chain-like data processing route and cannot yet be called a circulation network. Therefore, it is necessary to perform data re-analysis without destroying the skeleton diagram.
[0092] The system uses the "key intermediate data" in the framework as a new starting point for analysis and reprocessing. It can automatically select several reusable sets of intermediate data nodes according to rules. (e.g., surface reflectance, cloud mask, temporal composite image, feature stack, etc.) Selection principles include: located in the middle of the skeleton (has both upstream sources and can serve downstream tasks), semantically belonging to "general intermediate products" (can support multiple tasks), and of high quality or with sufficient evidence (easy to reuse).
[0093] Then the large model is required to be for each To plan a new branch, the branch must: reuse the existing... As input; output a derivative product that is different from the main task but related to the same domain (e.g., auxiliary variable, comparative product, error analysis product); specify the agent of the branch (which may be another entity) to reflect the flow.
[0094] It also forces the reuse of at least N intermediate data (e.g., the same reflectance / feature stack is used by different model branches) to form a true "circulation network" rather than a process chain.
[0095] Phase 3: Field instantiation, including attribute field generation and quality field generation.
[0096] This phase completes all nodes to the field level defined in the rule base, enabling them to be directly added to the database and used for subsequent quality propagation calculations. Based on the field table provided by the rule module, the large model completes the Data / Activity / Agent nodes one by one, following the principle of "real data first, fields interpretable".
[0097] 4. Inspection and Correction Module:
[0098] The verification and correction module is used to upgrade the circulation network output by the generation module from "seemingly reasonable" to an engineering-usable state that is "structurally computable, semantically reliable, traceable, and database-ready." This module uses a two-layer judgment based on rule verification and semantic verification, outputs structured errors and repair patches, and ensures stable convergence of the network through a minimal modification closed loop.
[0099] Rule validation (hard validation) is used to ensure that the structure is correct, the fields are complete, the topology meets the constraints, and the real data is not corrupted;
[0100] Semantic validation (soft validation) is used to ensure that the process conforms to common sense in remote sensing engineering, that local links are interpretable, and that there are no significant contradictions in the direction of quality impact.
[0101] Repair phase: Merge the two types of verification results into a repair plan, call the large model for local modifications, and prohibit overall regeneration.
[0102] 4.1 Rule Validation:
[0103] Rule verification is driven by the rule constraint module, which verifies whether the circulation network output by the generation module meets the requirements from four perspectives.
[0104] The algorithm step granularity check is to prevent situations where the algorithm steps are too granular, leading to uncomputable networks and unlocatable outliers. The main check method is to examine whether each algorithm node involves a single activity such as "preprocessing / spatial analysis / land use analysis" covering all algorithm steps.
[0105] Data structure verification is performed to ensure that the network can be parsed, stored in a database, queried, and used for graph computation. The main verification methods include: JSON files can be parsed correctly; each algorithm node outputs only one data point; references between nodes are closed; and the quality scores of each node are between zero and one.
[0106] Network topology complexity is checked to prevent the network from degenerating into a single chain, ensuring the existence of branches and cross-reuse, and guaranteeing that a "flow network" rather than a "processing flow" is generated. The main check method is to calculate the number of "critical data nodes" (used as input by at least two different Activities). The number of "critical data nodes" directly reflects the number of branches in the network; the more "critical data nodes," the more complex the network topology.
[0107] The authenticity check is to ensure that the source of the network's starting data is authentic and traceable, preventing the large model from fabricating sensor / product names. The main check method is to check whether the originating data node of each connected subgraph can be matched in the knowledge base.
[0108] 4.2 Semantic Validation:
[0109] Rule validation can only determine whether the structure is correct, but it cannot determine whether the process conforms to common sense in remote sensing engineering or whether local links are explainable. For example, "placing cloud detection after temporal composites" may be structurally valid, but it does not conform to common sense. Therefore, this invention further utilizes LLM for "local window review": each time, a local sub-map is extracted (containing only input data, algorithm steps, output data, and the corresponding subject), and LLM answers the following questions based on prior remote sensing knowledge and sub-map information:
[0110] Does the input-output order of this activity conform to common engineering sense?
[0111] Are there obvious errors such as "cloud detection after temporal synthesis" or "geometric correction after model inference"?
[0112] Are any necessary intermediate products or key steps missing?
[0113] Are there any instances of "empty data source / empty product name" that do not conform to anchor point constraints?
[0114] Do the changes in algorithm description and data quality conform to common sense?
[0115] 4.3 Rule semantic repair:
[0116] The validation results from rule validation and semantic validation are merged together, and the repair is performed following the principles of minimal modification and prohibition of overall regeneration: for deterministic issues, rule repair is prioritized (such as filling missing fields, correcting references, correcting enumerations, and splitting multi-output activities); for semantic issues, "local rewriting" is adopted: only the local subgraph that needs to be repaired is handed over to the LLM, which outputs the repaired nodes / edges or JSON patches, and overall regeneration is prohibited. After repair, it re-enters hard validation and semantic review until all passes or the iteration limit is reached (the limit is configurable, such as 3-5 rounds).
[0117] Through the above two-level verification and rule semantic repair, this invention can stably converge the network generated by the large model from "textually readable" to an engineering-usable state that is "structurally computable, semantically reliable, and with traceable anchor points," thereby meeting the application needs of the aerospace data element circulation network in advanced intelligent analysis tasks such as quality dissemination, root cause tracing, and asset management.
[0118] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
Claims
1. A system for constructing a multi-constraint aerospace data element circulation network based on a large model, characterized in that, include: The knowledge base module is used to store and output structured indexes. The input of the knowledge base module receives a literature database and an algorithm database. The knowledge base module converts coarse-grained algorithm names into fine-grained algorithm steps and associates them with research topics and data elements, outputting a structured index of research topic-data element-algorithm step set. The rule constraint module, set up in parallel with the knowledge base module, is used to define and output the rule constraint set. The rule constraint module abstracts the rules into a unified rule constraint set. The network generation module, whose input end is connected to the output end of the knowledge base module and the output end of the rule constraint module respectively, is used to perform structured graph construction within the feasible space defined by rules based on the retrieved prior knowledge, real data and structured index. The inspection and correction module has its input end connected to the output end of the network generation module, and its output end is fed back to the network generation module, forming a modular closed-loop architecture. The inspection and correction module uses rule verification and semantic verification as the criteria to improve the circulation network output by the network generation module to an engineering usable state. In the network generation module, at the start of each task, a set of prompt words is constructed based on the user's instructions, rule base, and knowledge base, which includes at least the following: The research task description should include the research content, task type, research area, time frame, and clearly state the research results that the task aims to achieve. Real datasets are selected by users, and each real dataset provides attribute field information and quality scores for each dimension with evidence summaries. A collection of algorithm process fragments is used to analyze and understand the research task in advance, and to plan a suitable technical route. The complex research task is broken down into several algorithm steps. The decomposed algorithm steps are matched with the knowledge base, and finally stored in the prompt words in the form of a list. The generation process was broken down into three stages; Phase 1: Skeleton Generation; First, based on the real data set and algorithm process fragment set recorded in the prompt words, an automatic technical route that can directly complete the research task is planned, and a skeleton diagram containing node IDs and core attributes is output. Phase Two: Data Reanalysis and Expansion; The intermediate data in the skeleton is used as a new starting point for analysis and reproduction is carried out, automatically selecting a set of several reusable intermediate data nodes according to rules. ; Then for each To plan a new branch, the branch must: reuse the existing... As input; output a derivative product that is different from the main task but related to the same domain; specify the Agent of the branch to reflect the flow; It also forces the reuse of at least N intermediate data to form a circulation network; Phase 3: Field instantiation, including attribute field generation and quality field generation; complete all nodes to the field level defined in the rule base, and complete the Data / Activity / Agent node by node according to the field table provided by the rule module and following the principle of "real data first and fields interpretable".
2. The system for constructing a multi-constraint aerospace data element circulation network based on a large model according to claim 1, characterized in that, The literature database contains papers on data products related to geographic information science and their cited papers, and extracts research topics, data elements, author affiliations and institutional information, as well as the confidence level calculated for each data element. The algorithm database stores prior knowledge of the aerospace data element processing flow.
3. The system for constructing a multi-constraint aerospace data element circulation network based on a large model according to claim 2, characterized in that, The confidence level calculation in the knowledge base module is based on the weight of the occurrence position, the frequency of occurrence across documents, and the author's authority, and is used as the basis for real data constraints in subsequent generation.
4. The system for constructing a multi-constraint aerospace data element circulation network based on a large model according to claim 1, characterized in that, The rule constraint set includes at least the following: structure and field rules, defining the required fields, data types, enumeration ranges, and dependencies between fields for Data nodes, Activity nodes, and Agent nodes; network topology and scale rules, ensuring the network is a graph structure rather than a chain structure, specifying that Activity nodes are multi-input single-output, data nodes are reusable, and the subject can execute multiple algorithm nodes, and specifying the total number of nodes and the proportion of the three types of nodes; quality constraint rules, representing data quality as a four-dimensional vector, defining the calculation methods for algorithm reliability coefficients and subject credibility coefficients, and establishing a quality propagation formula to achieve computable quality propagation; and output and reference consistency rules, ensuring that the references of each node are mutually closed and that the ID number is unique.
5. The system for constructing a multi-constraint aerospace data element circulation network based on a large model according to claim 4, characterized in that, The structure and field rules in the rule constraint module stipulate that: the Data node contains the fields id, label, description, intrinsic_quality, source_activity, and target_activities; the Activity node contains the fields id, label, description, quality_impact_pattern, inputs_data, outputs_data, and Quality, and the length of outputs_data must be 1; the Agent node contains the fields id, label, description, quality, and Activitylist.
6. The system for constructing a multi-constraint aerospace data element circulation network based on a large model according to claim 4, characterized in that, The quality constraint rules in the rule constraint module represent data quality as a four-dimensional vector. ; in, Describing the accuracy of the data, Describe the integrity of the data. Describe the timeliness of the data. The completeness of the documentation describing the data.
7. The system for constructing a multi-constraint aerospace data element circulation network based on a large model according to claim 4, characterized in that, The algorithm reliability coefficient in the quality constraint rules is: , It is calculated using the following formula: ; Among them, methodological rigor Reproducibility Achieving maturity , , , These are the weighting coefficients.
8. The system for constructing a multi-constraint aerospace data element circulation network based on a large model according to claim 4, characterized in that, The subject credibility coefficient in the quality constraint rule is: The calculation method is shown in the formula below: ; Among them, professionalism Institutional Credibility Openness / auditability , , , These are the weighting coefficients.
9. The system for constructing a multi-constraint aerospace data element circulation network based on a large model according to claim 1, characterized in that, The network generation module constructs a set of prompt words based on user instructions, rule base, and knowledge base. The set of prompt words includes at least a research task description, a set of real data, a set of algorithm flow fragments, and output format requirements. It adopts a phased generation strategy and executes them sequentially: In the skeleton generation phase, the technical route is planned based on the set of real data and the set of algorithm flow fragments, and a skeleton diagram containing node IDs and core attributes is output. In the data reanalysis expansion phase, without disrupting the skeleton graph, key intermediate data is used as a new starting point for analysis and reproduction is carried out, forcibly reusing at least N intermediate data nodes to form a circulation network; in the field instantiation phase, all nodes are completed to the field level defined in the rule base, so that they can be directly entered into the database and used for subsequent quality propagation calculations.
10. The system for constructing a multi-constraint aerospace data element circulation network based on a large model according to claim 1, characterized in that, The rule verification includes algorithm step granularity verification, data structure verification, network topology complexity verification, and authenticity verification. The semantic verification includes local window review to determine whether the process conforms to the common sense of remote sensing engineering. The verification and correction module merges the two types of verification results into a repair plan, follows the principle of minimum modification and prohibits overall regeneration, and performs the repair. After the repair, it re-enters the verification until all passes or the iteration limit is reached.
Citation Information
Patent Citations
Geospatial data element circulation method and system based on smart contract
CN120950616A
Large model information processing method based on knowledge graph
CN121960492A