A Method for Constructing a Catalog System of Water Transport Scientific Data Resources

By constructing a multidimensional orthogonal semantic star map based on knowledge cells, the problems of scalability and data silos in the water transport scientific data resource catalog system were solved, data semantic association and value quantification were realized, and maintenance costs were reduced.

CN122489684APending Publication Date: 2026-07-31TIANJIN RES INST FOR WATER TRANSPORT ENG M O T
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN RES INST FOR WATER TRANSPORT ENG M O T
Filing Date
2026-07-06
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

The existing catalog system of scientific data resources for water transport suffers from problems such as poor scalability, data silos, high manual maintenance costs, and inability to quantify data value when dealing with multi-source heterogeneous data management.

Method used

We use semantic decomposition rules and large language models to generate knowledge cells carrying DNA encoding and metadata, use graph neural networks to model semantic associations, construct a multidimensional orthogonal semantic star map, and generate a dynamic directory view through a user profile model, and combine it with a value transmission network to achieve self-updating.

Benefits of technology

It enables the classification tree to be reconstructed without the need to reconstruct new types of data, explicitly describes the semantic relationships of data, and has the ability to self-expand, strongly associate, be dynamic and quantify value in the catalog system, thereby reducing the cost of manual maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489684A_ABST
    Figure CN122489684A_ABST
Patent Text Reader

Abstract

This invention proposes a method for constructing a catalog system for water transport scientific data resources. The method includes: accessing a raw data resource pool; generating a set of knowledge cells using semantic decomposition and a large language model; constructing a multi-dimensional orthogonal semantic star map; constraining and mapping the cells in terms of subject domain, data form, and application scenario dimensions using a three-dimensional orthogonal coordinate system to generate an orthogonal projection index; converting user query intent into lens parameters to generate a dynamic catalog view oriented towards business scenarios; and calculating multi-level value transmission coefficients from cells to decision-making objectives through a value transmission network to generate self-updating instructions and value assessment results for the catalog. This invention generates a real-time dynamic view through a three-dimensional orthogonal index and dynamic lens parameters; quantifies the data support path for decision-making using a value transmission network; and achieves adaptive evolution of the catalog through a self-evolution engine, reducing maintenance costs and constructing a catalog system for water transport scientific data resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water transport technology, and in particular to a method for constructing a catalog system of scientific data resources for water transport. Background Technology

[0002] The catalog system for water transport scientific data resources is a core infrastructure for realizing the asset-based management and efficient utilization of data in the water transport sector. Currently, most existing catalog systems for water transport scientific data resources follow a traditional hierarchical classification model. Typical examples include tree-like hierarchical catalogs built based on international standards such as the ISO 19115 geographic information metadata standard and the DCAT (Data Catalog Vocabulary) data catalog vocabulary, or those with added extended fields to adapt to the specific needs of the water transport sector. These systems use a subject domain or data type as the root node, subdividing layer by layer to form a fixed classification tree structure. Each data resource is assigned a unique classification code and attached to the corresponding catalog node, thereby enabling the registration, retrieval, and location of data resources. However, with the explosive growth in the scale of data in the water transport science sector and the increasing diversification of data formats, the aforementioned traditional catalog systems have gradually revealed their inherent structural limitations when facing the unified management needs of multi-source heterogeneous data such as hydrological observation data, waterway survey data, port operation data, ship AIS trajectory data, remote sensing image data, numerical model output data, and text report data.

[0003] The existing catalog system for water transport scientific data resources suffers from a hierarchical structure that, once solidified, is difficult to expand. Adding new data types often requires reconstructing the entire classification tree, causing the maintenance cost of the catalog system to increase exponentially with the growth of data volume. Semantic relationships between data points are almost zero; the same hydrological data is repeatedly cataloged in both waterway management and port scheduling scenarios, resulting in severe data silos and resource waste. The catalog is a static snapshot, unable to reflect the derivative relationships and value changes that arise over time, making it difficult to support dynamic decision-making needs. Cataloging is highly dependent on manual labor; facing the massive and rapidly growing multi-source heterogeneous data in the water transport field, manual maintenance is extremely costly and error rates are difficult to control. Furthermore, it lacks a characterization of the data value transmission path, failing to answer the core question of "what decisions can this data support?", resulting in the inability to quantify and effectively convey the decision-making support value of data resources. Summary of the Invention

[0004] This invention aims to at least solve the technical problems existing in the prior art, and in particular, it innovatively proposes a method for constructing a catalog system of water transport scientific data resources.

[0005] To achieve the above-mentioned objectives of this invention, this invention provides a method for constructing a catalog system of water transport scientific data resources, the method comprising: S1. Access multi-source heterogeneous data in the field of water transport science, including hydrological observation data, waterway survey data, port operation data, ship AIS trajectory data, remote sensing image data, numerical model output data and text report data, to form a raw data resource pool for water transport science; S2. Based on the original data resource pool of water transport science, use semantic decomposition rules and large language models to generate a set of knowledge cells carrying DNA encoding and metadata; S3. Based on knowledge cell sets, using graph neural networks and a three-layer relationship discovery strategy of rules, statistics and reasoning, the semantic associations between knowledge cells are automatically modeled to generate a multi-dimensional orthogonal semantic star map. S4. Based on the multidimensional orthogonal semantic star map, the three-dimensional orthogonal coordinate system is used to perform structured constraints and coordinate mapping on all cells in three dimensions: subject domain, data form, and application scenario, to generate the orthogonal projection index of knowledge cells in three-dimensional space. S5. Based on the multidimensional orthogonal semantic star map and orthogonal projection index, the user profile model and scene lens parameter system are used to perform natural language parsing of user query intent and convert it into lens parameters to generate a dynamic directory view for business scenarios. S6. Based on the dynamic directory view and multidimensional orthogonal semantic star map, the value transmission network is used to calculate the multi-level value transmission coefficient from the knowledge cell to the decision target. The self-evolution engine drives the closed-loop iteration of relationship discovery, quality assessment and classification recommendation to generate the self-updating instructions and value assessment results of the directory system.

[0006] The beneficial effects of this invention are as follows: This invention atomizes multi-source heterogeneous data into knowledge cells carrying DNA encoding and metadata through semantic decomposition rules and a large language model. It automatically models seven semantic relationships—derivative, synonym, composition, strong correlation, evolution, support, and contradiction—using graph neural networks and a three-layer relationship discovery strategy of rules, statistics, and reasoning, constructing a multi-dimensional orthogonal semantic star map. This fundamentally breaks through the structural limitations of traditional hierarchical classification trees. The access of new data types does not require reconstructing the entire classification tree, and the semantic relationships of the same data in different business scenarios are explicitly characterized and interconnected, completely eliminating data silos. Secondly, based on a three-dimensional orthogonal coordinate system, all cells are subjected to structured constraints and orthogonal projection indexes in three dimensions: subject domain, data form, and application scenario. Combined with a user profile model and a scenario lens parameter system, user query intent is transformed into lens parameters, generating... The dynamic catalog view, tailored to business scenarios, upgrades the catalog from a static "snapshot" to a dynamic view that can change in real time according to user intent and business scenarios. Secondly, through a value transmission network, multi-level reachability traversal is performed along the fusion relationship edges with the decision target as the root node. This calculates the multi-level value transmission coefficient from each cell to the decision target and generates a value transmission map, accurately depicting the value transmission path of the data and quantitatively answering the core question, "What decisions can this data support?" Finally, a self-evolutionary engine drives a closed-loop iteration of relationship discovery, quality assessment, and classification recommendation, automatically generating self-updating instructions and value assessment results for the catalog system. This achieves adaptive evolution of the catalog system, significantly reducing manual maintenance costs, and thus constructing a water transport scientific data resource catalog system with five core capabilities: self-expansion, strong correlation, dynamism, automation, and quantifiable value.

[0007] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0008] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart of a method for constructing a catalog system of water transport scientific data resources according to the present invention. Detailed Implementation

[0009] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0010] like Figure 1As shown, a method for constructing a catalog system of scientific data resources for water transport is provided, the method comprising: S1. Access multi-source heterogeneous data in the field of water transport science, including hydrological observation data, waterway survey data, port operation data, ship AIS trajectory data, remote sensing image data, numerical model output data and text report data, to form a raw data resource pool for water transport science; In step S1, hydrological observation data (such as water level, flow velocity, sediment concentration, etc.) need to be standardized and converted according to the "Specifications for Hydrological Observation of Water Transport Engineering", and missing values ​​should be filled using linear interpolation. Channel survey data (such as water depth, topography, bottom sediment, etc.) need to be aligned with the WGS-84 spatial coordinate system to generate a timestamped Shapefile vector file. Port operation data (such as cargo throughput, loading and unloading efficiency, berth utilization rate, etc.) need to be synchronized with the port management system API in real time to the structured database tables, completing field mapping and redundant data filtering. Ship AIS trajectory data needs to be parsed using the NMEA0183 protocol message to extract the ship's MMSI. Key fields such as latitude, longitude, speed, and heading are aggregated and stored by ship ID and time series. Remote sensing image data needs to undergo radiometric correction, geometric correction, and band fusion processing to generate TIFF format multispectral images, and metadata indexes such as capture time, resolution, and coverage are established. Numerical model output data (such as water flow field, wave field, and water quality simulation results) needs to be converted from NetCDF or HDF5 format multidimensional arrays into structured data, and grid node coordinates and physical quantity values ​​are extracted. Text report data (such as waterway maintenance reports and accident analysis reports) needs to be scanned using OCR, and then key entities and structured fields are extracted using natural language processing tools. All accessed data needs to be tagged with metadata such as source identifier, acquisition time, and data accuracy, and a unique data fingerprint is generated using the existing SHA-256 hash algorithm to ensure data traceability and uniqueness.

[0011] S2. Based on the original data resource pool of water transport science, use semantic decomposition rules and large language models to generate a set of knowledge cells carrying DNA encoding and metadata; S3. Based on knowledge cell sets, using graph neural networks and a three-layer relationship discovery strategy of rules, statistics and reasoning, the semantic associations between knowledge cells are automatically modeled to generate a multi-dimensional orthogonal semantic star map. S4. Based on the multidimensional orthogonal semantic star map, the three-dimensional orthogonal coordinate system is used to perform structured constraints and coordinate mapping on all cells in three dimensions: subject domain, data form, and application scenario, to generate the orthogonal projection index of knowledge cells in three-dimensional space. S5. Based on the multidimensional orthogonal semantic star map and orthogonal projection index, the user profile model and scene lens parameter system are used to perform natural language parsing of user query intent and convert it into lens parameters to generate a dynamic directory view for business scenarios. S6. Based on the dynamic directory view and multidimensional orthogonal semantic star map, the value transmission network is used to calculate the multi-level value transmission coefficient from the knowledge cell to the decision target. The self-evolution engine drives the closed-loop iteration of relationship discovery, quality assessment and classification recommendation to generate the self-updating instructions and value assessment results of the directory system.

[0012] Optionally, generating a set of knowledge cells carrying DNA encoding and metadata includes: S201. Perform structured extraction and type determination on the multi-source heterogeneous data in the water transport science raw data resource pool to generate an initial data fragment set containing observation data, model data, remote sensing data, text data, rule data and metadata; In step S201, observation data needs to be acquired from equipment terminals such as hydrological stations and waterway monitoring points via data interfaces. Numerical indicators such as water level, flow velocity, and sediment concentration, along with their corresponding acquisition timestamps, are extracted. The criteria for this determination are that the data contains continuous or discrete physical quantity observations and time dimension labels. Model data needs to be parsed from NetCDF / HDF5 files output by numerical simulation software (such as MIKE, SWAN, etc.), extracting physical quantity parameters such as grid node coordinates, water flow velocity, and wave height, as well as the simulation time step. The criteria for this determination are that the data contains a spatial grid structure, multi-dimensional arrays, and simulation scene identifiers. Remote sensing data needs to undergo band extraction and feature interpretation of satellite imagery (such as Sentinel-2 and Landsat series) to obtain spatial information such as land cover type and waterway shoreline changes. The criteria for this determination are that the data contains a geographic reference coordinate system, image resolution, and spectral band attributes. Text data... The existing BERT pre-trained model should be used for entity recognition (such as port name, accident type), relation extraction (such as "port-throughput-value"), and event extraction to generate structured key-value pairs or triples. The judgment criteria are semantic units formed after unstructured text is processed by NLP. The rule data should extract constraints (such as waterway depth threshold) and calculation rules (such as berth utilization rate formula) from industry standards such as the "Waterway Engineering Design Code" and "Waterway Infrastructure Maintenance Technical Code". The judgment criteria are that the data contains logical judgment statements, mathematical operation expressions, or threshold parameters. The metadata should extract fields such as data source organization, collection accuracy, storage format, and update frequency from the header information and description documents of various types of data. The judgment criteria are the core element set that conforms to the DCAT-AP (Data Catalog Application Configuration File) or FGDC (Federal Geographic Data Council) metadata standards.

[0013] S202. Based on the initial set of data fragments, semantic understanding is performed using a large language model adjusted for the water transport domain. The subject domain affiliation, keywords, abstract, and semantic fingerprint vector of each fragment are automatically extracted to generate a fragment annotation set carrying semantic group metadata. In step S202, a pre-trained general-purpose large language model (such as BERT-base) is used as the base model. Fine-tuning is performed using a water transport domain-annotated dataset (containing over 100,000 corpora including standard texts of water transport engineering, academic papers, industry reports, and waterway maintenance records). This adjusts the model's attention mechanism weights and fully connected layer parameters, enhancing its semantic capture capabilities for professional terms such as "waterway depth threshold," "berth turnover rate," and "AIS trajectory analysis." Regarding subject domain extraction, the model automatically categorizes segments based on core terms by matching them to a pre-defined water transport science subject classification system (covering five major categories: waterway engineering, port engineering, water transport economics, maritime safety, and environmental monitoring). For example, segments containing "flow velocity measurement" and "siltation" are categorized under "waterway engineering," while segments containing "ship oil spill accidents" and "navigation safety assessments" are categorized under "maritime safety." The keyword extraction module combines the existing TF-IDF algorithm with the thesaurus of the "Basic Terminology Standard for Waterway Engineering" to select the top 5 weighted professional terms in the segment, such as "shoreline change," "spectral reflectance," and "spatial resolution" from remote sensing image data segments. In the abstract generation stage, the model generates a 50-100 word structured abstract based on the core content of the segment, including data type, core indicators, data source, and time range. For example, the abstract for a hydrological observation segment is "Daily water level and flow velocity observation data of a hydrological station in the middle and lower reaches of a certain river in 2023, collected once per hour, with a data accuracy of ±0.05m, sourced from the existing waterway monitoring system of a certain river." Semantic fingerprint vector generation is achieved by converting each segment into a fixed-length vector through the output of the last hidden layer of the model. The cosine similarity between vectors is used to quantify the semantic association between segments.

[0014] S203. Based on the fragment annotation set, using a hybrid decomposition strategy of rule engine and large language model, the initial data fragments are atomically decomposed according to the judgment criteria of semantic completeness, indivisibility, unique identification, quality assessability and association traceability, to generate a set of candidate cells that meet the knowledge cell judgment criteria. In step S203, the rule engine prioritizes executing preset decomposition rules on structured and semi-structured data (such as observation data, model data, and port operation data): For numerical observation data, it is decomposed at the smallest granularity of the collection timestamp (such as hourly), ensuring that each decomposed unit contains complete physical quantity values, time labels, and metadata; for spatial grid model data, it is decomposed at a single grid node, with each unit carrying grid coordinates, physical quantity parameters, and simulation time step; for structured port operation data, it is decomposed at a single record, with each record containing complete fields such as berth ID, operation time, and throughput. The large language model performs semantic segmentation on unstructured text data (such as accident analysis reports and waterway maintenance records): Based on pre-trained semantic understanding capabilities in the water transport domain, the model identifies the smallest unit with independent semantics in the text (such as a single event description or a rule clause). For example, it decomposes "Analysis of waterway environmental factors in a collision accident involving a vessel on a certain river in 2023" from an accident report as an independent candidate cell, ensuring that it meets the requirements of semantic completeness and indivisibility. The fusion phase of the hybrid strategy: The rule engine merges the structured data decomposition results with the unstructured data decomposition results output by the large model. Units that do not meet the judgment criteria are filtered out using quality verification rules (such as checking for unique data fingerprints and the completeness of metadata fields). In cases where there is overlap or conflict between the rule engine and the large model's decomposition results, rules from fields such as the "Water Transport Engineering Data Element Standard" are prioritized for judgment. If ambiguity persists, a manual review process is triggered. In the final candidate cell set, each cell is assigned a temporary unique identifier, and its decomposition source and association relationships are recorded.

[0015] The method for training the large language model in this embodiment is as follows: First, the labeled corpus of the water transport domain is divided into training set, validation set and test set in a ratio of 8:1:1. The initial learning rate is set to 5e-5. The AdamW optimizer is used to update the weights. The batch size is set to 16. The training rounds are set to 3 rounds. After each round of training, the entity recognition accuracy and semantic classification accuracy are calculated on the validation set. An early stopping mechanism is used to prevent the model from overfitting. Finally, the weight with the highest accuracy on the validation set is selected as the final parameters of the fine-tuned model.

[0016] S204. Based on the candidate unit cell set, using the domain code, subclass code, temporal granularity code, spatial granularity code, quality grade code, lineage tag code, semantic fingerprint hash and check bit encoding rules in the DNA coding system, a unique data DNA code is generated for each candidate unit cell, and a set of unit cell identifiers carrying the DNA code is generated. In step S204, the data DNA coding system adopts a hierarchical structure of "domain code - subclass code - time granularity code - spatial granularity code - quality level code - lineage tag code - semantic fingerprint hash - check bit". The coding rules for each part are as follows: Domain code: uses 2 digits, corresponding to the five core disciplines of water transport science (waterway engineering = 01, port engineering = 02, water transport economics = 03, maritime safety = 04, environmental monitoring = 05); Subclass code: uses 3 digits, representing the subdivisions under the domain, such as water depth measurement = 001, siltation monitoring = 002, and navigation mark status monitoring = 003 under the waterway engineering domain; Time granularity code: uses 1 digit, indicating the time acquisition / update granularity of the data (hourly = 1, daily = 2, monthly = 3, quarterly = 4, annual = 5); Spatial granularity code: uses 1 digit, indicating the data... Spatial coverage (basin level = 1, river section level = 2, port level = 3, berth level = 4, ship level = 5); Quality grade code: using a 1-digit number, divided into 5 levels based on data accuracy, completeness, and timeliness (Excellent = 1, Good = 2, Qualified = 3, Needs optimization = 4, Invalid = 5); Lineage tag code: using a 1-digit number, identifying the data source type (Observational data = 1, Model data = 2, Remote sensing data = 3, Text data = 4, Rule data = 5); Semantic fingerprint hash: using a 16-bit hexadecimal string, obtained by compressing the semantic fingerprint vector generated by S202 using the SHA-1 algorithm, used for quickly matching semantically similar cells; Check bit: using a 1-digit number, obtained by weighted summation of the first 7 parts of the code (weights are 2, 3, 5, 7, 11, 13, 17 respectively) and modulo 10, ensuring the integrity and correctness of the code.

[0017] After DNA encoding is completed, the encoding of each candidate cell is associated with its metadata and semantic group metadata and stored to form a complete set of knowledge cells containing unique identifiers, semantic attributes, quality characteristics and lineage relationships. This provides core data unit support for the subsequent construction of multidimensional orthogonal semantic star map and the generation of dynamic catalog view.

[0018] S205. Based on the cell identifier set, using a metadata model that includes identity group, semantic group, quality group, lineage group, permission group and value group, automatically fill each cell with identity metadata, semantic metadata, quality metadata, lineage metadata, permission metadata and value metadata to generate a knowledge cell set carrying complete metadata. In step S205, the identity metadata includes the cell's DNA code, a system-generated globally unique identifier, a temporary identifier, and a data fingerprint. The GUID ensures uniqueness across system environments, and the data fingerprint verifies that the cell content has not been tampered with. Semantic metadata uses subject domain affiliation, Top 5 professional keywords, a 50-100 character structured summary, and a semantic fingerprint vector, supplemented with data format labels (e.g., numerical, spatial grid, text semantic), which are automatically mapped based on data type determination results. Quality metadata uses quality level codes as its core, supplemented with specific quantitative indicators: the accuracy error range of numerical data annotation (e.g., water level data ±0.05m), the OCR recognition accuracy and entity extraction recall rate of text data annotation, the completeness (missing field percentage <1%) and timeliness (update interval ≤7 days) of all data annotations, and is linked to a quality verification log (recording verification time, execution entity, and results). The lineage metadata records the original data source ID of the cell (such as hydrological station equipment number, NetCDF file path), the disassembled parent cell ID, upstream and downstream related cell IDs (such as different band cells derived from the same original image), and the complete processing chain (such as "satellite image → radiometric correction → band fusion → atomization disassembly → cell"), forming a linear and traceable lineage relationship chain. The permission metadata sets the subject identifier based on the data ownership institution, assigns user role permissions (read-only / edit / download) based on the RBAC model, and configures the effective time and scope of permissions (such as only for a specific waterway planning project team). Sensitive data (such as accident analysis reports) is additionally anonymized. The value metadata generates an initial value score (0-10 points) by statistically analyzing the cell's historical call frequency and the number of times it is cited in decision reports, labels potential application scenarios (such as waterway maintenance, maritime emergency response, environmental assessment), and associates it with the output of a pre-trained value prediction model (such as the value weight for supporting waterway planning). After all metadata is populated, a consistency check (checking field integrity and logical conflicts) is performed through the rule engine. Cells that pass the check are added to the final knowledge cell set.

[0019] S206. Based on the knowledge cell set carrying complete metadata, the completeness, accuracy, timeliness and source credibility of each cell are quantitatively scored, and cells with scores below the preset threshold are marked as pending review and pushed to the domain expert review interface. After the review is passed, the final knowledge cell set carrying DNA encoding and metadata is generated.

[0020] In step S206, the quantitative scoring adopts a multi-dimensional weighted calculation method: the completeness score is based on the metadata field missing rate (the proportion of missing fields to the total number of fields is inversely mapped to 0-100 points, with no missing fields receiving full marks); the accuracy score combines the accuracy error range in the quality metadata (such as the accuracy compliance rate of numerical data), entity extraction accuracy, and text OCR recognition accuracy, taking the average of each sub-item; the timeliness score is based on the matching degree between the data update interval and the preset threshold (100 points for an update interval ≤ 7 days, 5 points deducted for each day exceeding the limit, with a minimum of 0 points); the source credibility score is based on the qualification level of the data source institution and the completeness of the lineage (100% completeness of the lineage receives full marks, 10 points deducted for each missing link). The weights of the four dimensions are 25% for completeness, 30% for accuracy, 20% for timeliness, and 25% for credibility, respectively, and the weighted sum is used to obtain the total cell score. The preset qualified threshold is 80 points, and cells with a total score below this value are marked as pending review.

[0021] The system automatically generates tasks to be reviewed, including the cell's DNA code, metadata details, scoring details, and issue prompts (such as "Insufficient timeliness: update interval is 15 days, exceeding the preset 7-day threshold"). Based on the cell's subject domain, it is pushed to the corresponding domain experts (e.g., maritime safety cells are pushed to maritime engineering experts). Experts review the cell's raw data fragments, processing logs, and scoring criteria through the review interface, and conduct manual verification: if the issue is confirmed to be correctable (e.g., supplementing missing metadata fields), the score is adjusted to qualified and marked as passed; if the data contains irreparable errors (e.g., severely distorted core indicators), it is marked as failed and the reason for rejection is provided.

[0022] Cells that pass the review are updated to "valid" and included in the final knowledge cell set; cells that fail are removed from the candidate set, and optimization suggestions are sent to the data source. The system synchronously generates review logs, recording the review time, expert ID, review results, and modifications, ensuring the review process is traceable. The final knowledge cell set possesses a unique DNA code, complete metadata, and reliable quality verification.

[0023] Optionally, generating a multidimensional orthogonal semantic star map includes: S301. Based on the knowledge cell set, a graph neural network is used to perform a preliminary traversal and feature extraction of the potential semantic associations between all cells, generating an initial candidate set of relations containing candidate relation edges and candidate edge confidence. In step S301, a Graph Attention Network (GAT) is used as the core model, mapping each knowledge cell to a graph node. Node features are fused from the numerical representation of each field encoded by DNA, the dimensionality reduction features of the semantic fingerprint vector, and the pre-trained word embedding vectors of the Top 5 keywords in the metadata. During the training phase, positive and negative sample sets are constructed: positive samples are cell pairs with a semantic fingerprint cosine similarity ≥ 0.8, and negative samples are randomly selected cell pairs from cross-disciplinary domains with a similarity < 0.3. The model parameters are optimized using the cross-entropy loss function. The model calculates the attention weights between nodes through a multi-head attention mechanism. When generating candidate relationship edges, only edges with an attention weight ≥ 0.5 are retained as candidates, and this weight is used as the confidence of the candidate edge. To improve traversal efficiency, a Local Sensitive Hash (LSH) algorithm based on semantic fingerprint vectors is used for approximate nearest neighbor search, reducing unnecessary node pair calculations. At the same time, candidate edges are forcibly generated for direct parent / child cell pairs in the kinship chain (with the initial confidence set to 0.9). Each edge in the initial candidate relation set records the source cell ID, target cell ID, confidence level, and preliminary association type (such as "semantic similarity" or "bloodline derivative").

[0024] S302. Based on the initial candidate relation set, use the rule matching strategy to automatically confirm the candidate edges that meet the requirements of bloodline metadata matching, spatial inclusion matching and semantic fingerprint cosine similarity exceeding the preset threshold for three relation types: derivation, synonymy and composition, and generate a rule-confirmed relation edge set. In step S302, the rule matching strategy formulates decision rules for the three types of candidate edges: Lineage metadata matching (derived relationship): If the candidate edge source cell and the target cell have a direct parent / child relationship or upstream / downstream link relationship in the lineage metadata, the rule engine will confirm the relationship type as "derived", supplement the lineage link processing step, and increase the confidence to 0.95.

[0025] Spatial inclusion matching (composition relationship): Two cells have the same domain code and temporal granularity code, and their spatial granularity codes have a hierarchical inclusion relationship. The rule engine determines this as a "composition" relationship, records the spatial inclusion level, and marks it with a confidence level of 0.93. For example, a daily port throughput cell and a berth-level operation data cell in a port engineering domain form a "composition" relationship.

[0026] Semantic fingerprint similarity matching (synonymous relationship): The preset semantic fingerprint cosine similarity threshold is 0.85. If the similarity of unit cell pairs is ≥ this threshold, the subject domains belong to the same category, and the overlap rate of the top 5 keywords is ≥ 60%, the rule engine marks it as a "synonymous" relationship, supplements the similarity value and the list of overlapping keywords, and sets the confidence level to 0.90.

[0027] Rule matching is performed in priority order: lineage metadata matching > spatial containment matching > semantic fingerprint similarity matching. If a candidate edge satisfies multiple rules simultaneously, the rule with the highest priority is used to determine the primary relation type, and the other satisfied rule types are recorded. After rule matching is completed, candidate edges that do not satisfy the rules are filtered out, and a set of rule-confirmed relation edges is generated. Each edge contains the source / target cell ID, relation type, relation attributes, and confidence level.

[0028] S303. Based on the rule-confirmed relation edge set and the remaining candidate edges in the initial relation candidate set that have not been confirmed by the rule, use statistical analysis strategies to calculate the Pearson correlation coefficient and evaluate the mutual information value of the unit cell pairs in the remaining candidate edges. Edges with an absolute value of correlation coefficient greater than n are confirmed as strong correlations, and edges with consistent time series trends are confirmed as evolutionary relationships, thus generating a statistically confirmed relation edge set. In step S303, the statistical analysis strategy targets the remaining candidate edges after rule confirmation, and performs precise analysis in two scenarios based on the cell data type and correlation characteristics: Pearson correlation coefficient and mutual information value assessment (strong correlation): For numerical cell pairs (such as channel depth and sediment deposition data), calculate the Pearson correlation coefficient with a preset threshold n of 0.7. If the absolute value of the correlation coefficient > 0.7, the rule engine confirms it as "strong correlation," supplements the correlation coefficient value and time range, and sets the confidence level to 0.85. For non-numerical or mixed-type cell pairs (such as maritime accident text and statistical data tables), calculate the mutual information value (using the point mutual information formula) with a preset threshold of 0.6. If the mutual information value ≥ 0.6, mark it as "strong correlation," record the mutual information value and core semantic topic, and set the confidence level to 0.83.

[0029] Time series trend consistency analysis (evolutionary relationship): For cell pairs with time granularity labels (such as port throughput and regional GDP data), time series trend features are extracted, and the similarity is calculated using the DTW algorithm, with a preset threshold of 0.8. If the similarity is ≥0.8, the rule engine confirms it as "evolutionary," supplements the similarity value and time period matching, and sets the confidence level to 0.82.

[0030] After statistical analysis, strongly correlated and evolving relationship edges are integrated into a set of statistically confirmed relationship edges. Each edge contains source cell ID, target cell ID, relationship type, relationship attribute, and confidence level.

[0031] S304. Based on the rule-confirmed relation edge set, the statistically confirmed relation edge set, and the remaining candidate edges that are still not confirmed in the initial relation candidate set, semantic reasoning is performed using a large language model adjusted for the water transport domain. The semantic determination of supporting and contradictory relations is performed on the cell pairs in the remaining candidate edges to generate a reasoning confirmed relation edge set. In step S304, the basic large language model (such as GPT-4) is first fine-tuned specifically for the water transport field. The fine-tuning corpus covers professional texts such as standard documents of the water transport industry, waterway maintenance technical reports, maritime emergency response cases, and port operation statistical annual reports, so that the model has the ability to understand semantics and make logical inferences specific to the water transport field. For the remaining candidate edges that have not been confirmed by rule matching and statistical analysis strategies, the core metadata information of the source cell and the target cell (semantic fingerprint vector, Top 5 professional keywords, structured summary, data form label) and the existing partial relation edge context (such as the adjacent cell relationship of the same bloodline) are organized into a structured prompt and input into the fine-tuned model.

[0032] The model makes two types of relationship judgments for candidate edges: Supporting relationship: If the core conclusion of the target cell can be directly supported by the data, conclusion or logic of the source cell, such as "the average water depth of a certain river section in 2023 is 8.2 meters" (source cell) and "the river section needs to be dredged in 2023 to meet the 10-meter navigation requirement" (target cell), the model judges the two to be "supporting" relationship, and the relationship attribute supplements the supporting evidence (such as "the source cell data proves that the water depth of the channel does not meet the navigation standard, supporting the dredging suggestion of the target cell"). The confidence level is mapped to 0.88 based on the reasoning confidence score output by the model (when the model score is ≥0.9).

[0033] Conflicting Relationship: If the model detects a conflict in the core semantics of two unit cells, such as "the container throughput of a certain port increased by 12% year-on-year in 2023" (source unit cell) and "the container throughput of the same port decreased by 5% year-on-year in 2023" (target unit cell), the model determines it as a "conflicting" relationship. The relationship attribute records the core point of the conflict (such as "the conclusions of throughput growth and decline conflict"), and the confidence level is mapped to 0.86 (when the model score is ≥0.9).

[0034] After the reasoning is completed, the rule engine verifies the relationship type and reasoning output by the model, filters out candidate edges with insufficient reasoning reasons, and generates a set of reasoning-confirmed relationship edges. Each edge contains source cell ID, target cell ID, relationship type (support / contradiction), relationship attribute (supporting basis / contradiction core point) and confidence level.

[0035] S305. Based on the rule-confirmed relation edge set, the statistically confirmed relation edge set, and the reasoning-confirmed relation edge set, the relation fusion engine is used to perform conflict detection and confidence-weighted fusion of the three types of relation edges, and each edge is given a relation type attribute, a weight value attribute, and a discovery source attribute to generate a fused relation edge set. In step S305, the relation fusion engine's processing is divided into three core stages: conflict detection and resolution, confidence optimization, and attribute enhancement. Conflict detection and resolution involves traversing three sets of relation edges: rule confirmation, statistical confirmation, and inference confirmation. Conflicts are identified in relation edges involving the same source cell ID-target cell ID pair. For relation type conflicts (such as the mutual exclusion of "derived" and "contradictory"), a "source priority + confidence" strategy is used: rule-confirmed relation (priority 1) > statistical confirmation relation (priority 2) > inference confirmation relation (priority 3). If priorities are the same, edges with higher confidence are retained. For attribute conflicts (such as inconsistent attribute descriptions under the same relation type), the attribute of the high-confidence edge is used as the benchmark, and a comparative explanation of the conflicting attributes is added to the relation attributes (e.g., "The original statistical attribute has a correlation coefficient of 0.81; the rule lineage link 'satellite image → radiometric correction' is retained").

[0036] Confidence optimization: For multiple relation edges of the same unit cell pair without conflict (such as the simultaneous existence of a "composition" primary relation and a "strong correlation" secondary relation), all valid relation types are retained, and the confidence of each edge is adjusted by weighting. The confidence weight of the rule source is set to 0.4, the statistical source to 0.3, and the inference source to 0.3. The final confidence of each edge is calculated (for example, if the confidence of the rule edge is 0.93, its final confidence = 0.93 × 0.4 = 0.372; if the confidence of the statistical edge is 0.85, its final confidence = 0.85 × 0.3 = 0.255).

[0037] Attribute enhancement: Add three types of structured attributes to each merged relation edge: ① Relationship attribute: Define the main relation type and the sub-relation type (if there are multiple non-conflicting relations); ② Weight value attribute: Record the weight of the source strategy corresponding to the relation edge (e.g., "rule 0.4"); ③ Discovery source attribute: Label the original discovery strategy (e.g., "rule matching" or "rule matching + statistical analysis").

[0038] After the above processing is completed, each edge in the fused relation edge set contains the source / target cell ID, relation type (primary / secondary), relation attribute, confidence level, weight value, and discovery source.

[0039] S306. Based on the set of fusion relationship edges and the set of knowledge cells, the graph database modeling method is used to treat all knowledge cells as nodes and fusion relationship edges as directed edges. Orthogonal constraints and coordinate mapping are performed along the three dimensions of subject domain, data form and application scenario to generate a multidimensional orthogonal semantic star map.

[0040] In step S306, the definitions and encoding system of the three orthogonal dimensions are first clarified: Subject domain dimension: Based on the professional classification of water transport scientific data, it is divided into five core subcategories: waterway engineering (code 1), port operation (code 2), maritime management (code 3), water transport economics (code 4), and shipbuilding engineering (code 5). Each subcategory is further subdivided into secondary subdomains (such as waterway surveying and dredging maintenance under waterway engineering), corresponding to a unique hierarchical code; Data format dimension: According to the data organization form, it is divided into structured data (encoding 1, such as database tables, CSV files), semi-structured data (encoding 2, such as XML, JSON configuration files), and unstructured data (encoding 3, such as satellite imagery, technical report text). Application scenario dimension: Based on the actual use of data, it is divided into four major scenarios: decision support (code 1, such as waterway planning scheme generation), resource retrieval (code 2, such as data asset positioning), risk warning (code 3, such as ship collision risk prediction), and scientific research analysis (code 4, such as navigation efficiency modeling).

[0041] Secondly, perform three-dimensional coordinate mapping: For each knowledge cell, extract the subject domain label, data form label, and application scenario label from the metadata, and map them to the corresponding dimension's encoded values ​​to form three-dimensional coordinates (subject domain encoding, data form encoding, and application scenario encoding). For example, the coordinates of a structured data table of waterway depth measurement in 2023 (used for waterway dredging decision-making) are (1-1, 1, 1), where "1-1" represents the secondary subdomain of waterway measurement under waterway engineering.

[0042] Then, orthogonal constraints are implemented: ensuring the independence of the three dimensions, meaning that changes in the value of any dimension do not affect the attributes of other dimensions. For example, all data forms and application scenarios can be covered within the same subject domain, and the same data form can also serve different subject domains and scenarios, avoiding coupling between dimensions.

[0043] Finally, graph database modeling and star map generation were completed: The Neo4j graph database was used, with knowledge cells as nodes (attributes including cell ID, DNA encoding, metadata summary, and 3D coordinates), and relational edges as directed edges (attributes including relation type, confidence level, weight value, and discovery source). Using graph visualization tools (such as Neo4jBrowser), the nodes were arranged in space according to 3D coordinates, forming an orthogonal semantic star map with the subject domain as the X-axis, data form as the Y-axis, and application scenario as the Z-axis. Users can quickly locate target cells and their relationships using any dimension of filtering conditions (such as "port operation + structured + risk warning"), enabling multi-dimensional semantic association queries and analysis. The generated multi-dimensional orthogonal semantic star map not only clearly presents the classification system of water transport scientific data resources but also intuitively demonstrates the derivation, composition, and strong correlation relationships between data through semantic association edges.

[0044] Optionally, the orthogonal projection index of the knowledge cell in three-dimensional space includes: S401. Based on the multidimensional orthogonal semantic star map, a three-dimensional orthogonal coordinate system is used to define the subject domain axis, data form axis and application scenario axis, and a discretized scale and threshold range are set for each axis to generate a three-dimensional orthogonal coordinate framework. In step S401, the subject domain axis (X-axis) uses the numerical values ​​corresponding to the hierarchical coding as the discretization scale. The main category coding is the integer part (1-5, corresponding to the five core subcategories of waterway engineering, port operation, maritime management, water transport economics, and shipbuilding engineering), and the secondary subdomain coding is the decimal part (e.g., 1.1 represents waterway surveying under waterway engineering, 1.2 represents dredging and maintenance, etc.). The threshold range is 1.0 to 5.9 (covering all main categories and secondary subdomains). Each scale point corresponds one-to-one with a specific subject subdomain, ensuring that the subject affiliation of the cell can be accurately mapped. The data form axis (Y-axis) uses the data form coding value as the discretization scale, with scale values ​​of 1 (structured data), 2 (semi-structured data), and 3 (unstructured data), and a threshold range of 1 to 3. Each scale point directly corresponds to a data organization form, clearly reflecting the data form characteristics of the cell. The application scenario axis (Z-axis) uses the application scenario encoding value as a discretized scale, with scale values ​​of 1 (decision support), 2 (resource retrieval), 3 (risk warning), and 4 (scientific research analysis), and a threshold range of 1 to 4. Each scale point corresponds to a practical application scenario, clearly defining the purpose and positioning of the unit cell.

[0045] By combining the discretized scales and threshold ranges of the three axes mentioned above, a three-dimensional spatial grid containing all dimensional combinations is constructed. Each grid node corresponds to a unique combination of (subject domain code, data form code, and application scenario code).

[0046] S402. Based on the three-dimensional orthogonal coordinate framework and the multi-dimensional orthogonal semantic star map, all knowledge cells are mapped to three-dimensional spatial coordinate points using coordinate mapping rules, thereby generating a set of three-dimensional coordinates for each knowledge cell. In step S402, firstly, each knowledge cell in the multidimensional orthogonal semantic star map is traversed, and the subject domain affiliation label, data form type label, and application scenario label are extracted from its metadata. Secondly, according to the encoding rules defined in step S401, each label is converted into a corresponding coordinate axis value: the subject domain label is mapped to the hierarchical encoding value of the X-axis (e.g., the ship navigation safety subdomain under maritime management corresponds to 3.1), the data form label is mapped to the encoding value of the Y-axis (e.g., unstructured data corresponds to 3), and the application scenario label is mapped to the encoding value of the Z-axis (e.g., risk warning corresponds to 3). Finally, the converted three coordinate axis values ​​are combined to form the three-dimensional spatial coordinate point (X, Y, Z) of the cell and recorded in the cell's three-dimensional coordinate set. For example, a structured data table of container throughput statistics for a port in 2023 (used for port operation decision support) has the subject domain label "Port Operation" (2), the data form label "Structured" (1), and the application scenario label "Decision Support" (1), with corresponding three-dimensional coordinates of (2.0, 1, 1); another example is an unstructured text report of a maritime accident (used for risk warning analysis), with the subject domain label "Accident Analysis Subdomain under Maritime Management" (3.2), the data form label "Unstructured" (3), and the application scenario label "Risk Warning" (3), with corresponding three-dimensional coordinates of (3.2, 3, 3). Through this mapping process, all knowledge cells are accurately located to the corresponding grid nodes of the three-dimensional orthogonal coordinate frame.

[0047] S403. Based on the three-dimensional coordinate set of the unit cell, the orthogonal projection algorithm is used to orthogonally project all coordinate points to the subject domain data form plane, the subject domain application scenario plane and the data form application scenario plane respectively, generating three sets of two-dimensional projection point sets. In step S403, the core of the orthogonal projection algorithm is to map three-dimensional coordinate points onto a specified two-dimensional plane, and the specific rules are as follows: Projecting onto the subject domain data form plane (XY plane): Retaining the X (subject domain hierarchy code) and Y (data form code) values ​​in the three-dimensional coordinates of the unit cell, ignoring the Z (application scenario code), and generating two-dimensional projection points (X, Y); Projecting onto the subject domain application scenario plane (XZ plane): Retaining X and Z values, ignoring Y, and generating two-dimensional projection points (X, Z); Project onto the data form application scenario plane (YZ plane): retain the Y and Z values, ignore X, and generate two-dimensional projection points (Y, Z).

[0048] Secondly, aggregation optimization is performed on each set of two-dimensional projection points: all projection points are traversed, projection points with identical coordinate values ​​are merged, all knowledge cell IDs associated with the projection point are recorded, and statistical information such as the average confidence level and relation type distribution of the associated cells are calculated. For example, in the XY plane, the projection point (1.1, 1) corresponds to the waterway survey (1.1) and structured data (1) under waterway engineering. After merging, the point is associated with all cells that conform to this subject domain and data form, such as "2023 water depth measurement data table of a certain river section" and "waterway topography structured dataset", and the average confidence level of these cells is calculated to be 0.89.

[0049] Finally, an orthogonal projection index is constructed, and an index table is created for each set of two-dimensional projection points. The index key is the coordinates of the two-dimensional projection points, and the index value is the list of associated cell IDs and statistical information. This index supports fast retrieval. For example, if a user needs to find the cell of "Port Operation (2.0) + Risk Warning (3)", they can directly locate the projection point of (2.0, 3) in the XZ plane index and obtain all related cells, which greatly shortens the retrieval time.

[0050] S404. Based on three sets of two-dimensional projection points, an index structure is established for each set of projection points using the R-tree spatial indexing method, generating three sets of R-tree indexes, including subject domain data form R-tree index, subject domain application scenario R-tree index, and data form application scenario R-tree index. In step S404, the R-tree, as a dynamic balanced tree structure based on the minimum bounding rectangle (MBR) of spatial objects, can efficiently handle range queries and nearest neighbor queries of two-dimensional spatial data, perfectly adapting to the retrieval requirements of multi-dimensional projection points in this invention, and can significantly reduce the time complexity of queries.

[0051] For the construction of the R-tree index (corresponding to the set of projection points on the XY plane) for subject domain data: all projection points on the XY plane are divided into several non-overlapping minimum boundary rectangles according to their spatial distribution, and each rectangle contains several projection points; Recursively construct R-tree nodes: leaf nodes store the coordinates (X, Y) of a single projection point, a list of associated cell IDs, and statistical information (such as average confidence); non-leaf nodes store the MBR of child nodes and pointers to child nodes; A quadratic split strategy is used to handle node overflow: when the number of projection points in a node exceeds a preset threshold, the two farthest projection points are selected as split centers, and the remaining projection points are assigned to the nodes where the centers are closer, ensuring that the MBR coverage of the split nodes is minimized.

[0052] The construction process for subject domain application scenario R-tree index (corresponding to the XZ plane projection point set) and data form application scenario R-tree index (corresponding to the YZ plane projection point set) is the same as above, only the projection point coordinates need to be replaced with the two-dimensional values ​​(XZ or YZ) of the corresponding plane.

[0053] Index Optimization and Maintenance: Addressing the dynamic updating characteristics of water transport scientific data, the R-tree index supports incremental updates. When adding a new cell, three sets of projection points are generated based on its 3D coordinates and inserted into the corresponding positions in each R-tree. When deleting a cell, associated projection points are simultaneously removed and the node structure is adjusted. The R-tree is periodically rebalanced to avoid query efficiency degradation caused by excessive tree depth by adjusting node levels and MBR distribution. When a user initiates a multi-dimensional combined query (e.g., "Port Operations (X=2.0) + Structured Data (Y=1) + Risk Warning (Z=3)"), the system responds quickly through the following steps: A range query is executed in the XY, XZ, and YZ R-tree indexes to obtain a set of projection points that meet the conditions; the intersection of the three sets of projection points is performed to obtain projection points that simultaneously satisfy all dimensional conditions; based on the cell ID list associated with the intersection projection points, the target knowledge cell and its associated relationships are directly returned, reducing query response time by more than 80% compared to traditional full table scans. The final generated set of three R-tree indexes is shown.

[0054] S405. Based on three sets of R-tree indexes, the three sets of R-tree indexes are associated and merged using an index fusion engine. Cell DNA encoding, quality grade and value weight are added as index attributes to each index record to generate a fused index set. The core logic of the index fusion engine is to achieve the association mapping of three sets of R-tree indexes based on the unique ID of the knowledge cell. The specific process is as follows: Traverse all knowledge cell IDs, extract the projection point information and association statistics corresponding to the cell ID from the subject domain data form R-tree index, the subject domain application scenario R-tree index, and the data form application scenario R-tree index, and establish a cross-index association table with the cell ID as the key; add three core attributes to the association table record corresponding to each cell ID: directly reference the unique DNA code in the knowledge cell metadata for quick identification of the cell's core characteristics and source; and quantify and score the data based on three dimensions: metadata completeness (e.g., field missing rate), data accuracy (e.g., error range), and update frequency (e.g., whether it has been updated in the last 6 months), categorized as A (excellent). The data is categorized into four levels: B (Good), C (Acceptable), and D (Needs Optimization). The scoring rules are defined by the Water Transport Scientific Data Quality Standard. A comprehensive weight value is calculated based on the frequency of the cell's application scenarios (e.g., the number of times it has been used in decision support scenarios in the past year), the importance of the subject domain (e.g., the weight coefficient for maritime management data is 1.2), and the scarcity of the data format (e.g., the weight coefficient for unstructured data is 1.1). The weight value ranges from 0 to 2, with higher values ​​indicating higher application value. Related table records and additional attributes are integrated into a single fused index record. Each record contains: cell ID, XY index pointer, XZ index pointer, YZ index pointer, DNA code, quality level, and value weight. The fused index uses a key-value pair storage structure, with the cell ID as the key and the integrated information as the value. Consistency verification is performed on the generated fused index set to ensure that all cell IDs have corresponding projection point information in the three R-tree indexes. If any are missing, they are marked as abnormal and a completion process is triggered. Simultaneously, based on user historical query logs, the fused index records corresponding to frequently queried cell IDs are cached and optimized to further improve query response speed. The fused index set generated through the above steps achieves deep association of the three sets of R-tree indexes. It not only retains the retrieval efficiency of the projection indexes of each dimension, but also enhances the semantic richness of the index through additional attributes. It supports users to perform precise filtering by combining conditions such as quality level and value weight when querying. For example, users can specify "port operation structured data with quality level A and value weight ≥ 1.5", and the system can quickly locate the knowledge cells that meet the conditions, thus meeting the needs of refined management and efficient utilization of water transport scientific data resources.

[0055] S406. Based on the fused index set, use coordinate consistency verification rules to perform conflict detection and deduplication on all index records, and sort the verified index records in descending order of value weight to generate an orthogonal projection index of knowledge cells in three-dimensional space.

[0056] The core of the coordinate consistency verification rule in step S406 is to verify whether the coordinate values ​​of the same knowledge cell are logically consistent in the three sets of projection indices. Specifically, this includes: extracting the coordinates of the projection points corresponding to the three sets of R-tree index pointers in the fused index record, and cross-verifying the X-axis value (the X values ​​of the XY and XZ projection points must be completely equal), Y-axis value (the Y values ​​of the XY and YZ projection points must be completely equal), and Z-axis value (the Z values ​​of the XZ and YZ projection points must be completely equal). If there is a mismatch in coordinate values ​​in any dimension (e.g., XY projection point X is 2.0, XZ projection point X is 2.1), it is marked as a conflict record, triggering an automatic tracing mechanism: the system traces back to the coordinate mapping step in step S402, checks whether there are errors in the metadata extraction and encoding conversion process, and if it is a metadata annotation error, it is pushed to the manual review queue for correction; if it is a deviation in the execution of the mapping rule, the mapping process is automatically re-executed.

[0057] The deduplication process addresses the following two scenarios: First, when multiple fused index records correspond to the same cell ID (due to duplicate insertions caused by index updates), the latest version record is directly retained. Second, when different cell IDs have identical 3D coordinates and highly similar DNA codes (similarity ≥ 95%), they are identified as duplicate data. The record with the highest value weight and best quality level is retained, while the rest are marked as "to be archived" and stored in the backup index library.

[0058] The validated index records are sorted in descending order of value weight. If weights are the same, records with higher quality levels are prioritized. If quality levels are the same, they are sorted by cell update timestamp (latest update takes precedence). The final generated orthogonal projection index adopts a hierarchical storage structure: the top layer is the index list sorted by value weight, the middle layer is a mapping table between 3D coordinates and cell IDs, and the bottom layer associates the metadata and semantic relationships of the original knowledge cells. This index supports real-time incremental updates; when a new cell is added, process segments S401 to S406 are automatically executed to ensure the synchronization of the index and data resources.

[0059] Optionally, generating dynamic catalog views tailored to specific business scenarios includes: S501. Based on the multidimensional orthogonal semantic star map and orthogonal projection index, the user profile model is used to load the current user's role identity, subject domain preference, data form preference, scenario preference and historical behavior record to generate a user profile set containing the user's multidimensional preference weight vector. In step S501, the role identity is quantitatively mapped: user roles (such as maritime safety researcher, port operation decision-maker, waterway engineering designer, etc.) are mapped to role weight coefficients. For example, a maritime safety researcher corresponds to a safety domain weight coefficient of 1.3, and a port operation decision-maker corresponds to an operation domain weight coefficient of 1.2. Secondly, the preference dimension is quantitatively calculated: subject domain preference is calculated based on the weighted average of the proportion of each subject domain in the user's queries over the past 3 months. For example, if 70% of the user's queries are concentrated in port operations, then the port operations preference weight is 0.7. Data format preference is determined according to the format distribution of the user's downloaded / accessed data. For example, if 80% of the access is structured data... For data, the structured data preference weight is 0.8; scenario preferences are based on the frequency of scenario tags in the user's historical queries, such as a risk warning scenario accounting for 60%, in which case the risk warning preference weight is 0.6; historical behavior records are adjusted for timeliness using a time decay function (e.g., behavior from the most recent month has a weight of 1.0, while behavior from 3 months ago has a weight of 0.5); finally, the weight values ​​of the above dimensions are integrated into a multi-dimensional preference weight vector, for example, the vector for a maritime safety researcher is (port operation: 0.6, maritime management: 0.8, structured data: 0.5, unstructured data: 0.9, risk warning: 0.9, decision support: 0.4). Meanwhile, the user profile set supports real-time dynamic updates. Every time a user completes a query, download, or annotation operation, the system automatically recalculates the preference weight vector to ensure that the profile highly matches the user's current needs.

[0060] S502. Based on the user profile set, the query intent text input by the user is segmented, entity recognized and intent classified, and the query keywords, target subject domain, target data form and target application scenario are automatically extracted to generate query parsing results carrying intent feature vectors. In step S502, the word segmentation process employs a word segmentation algorithm enhanced with terminology from the field of water transport science. This algorithm, combined with a professional dictionary constructed based on the "Water Transport Science and Technology Terminology Standard," accurately segments the query text, preventing missegmentation of technical terms. For example, "I need structured data for port operation risk warning in 2023" is segmented into "I / need / 2023 / port operation / risk warning / of / structured data".

[0061] Secondly, the entity recognition step is based on a pre-trained BiLSTM-CRF model, using a customized training dataset (containing over 100,000 labeled samples) tailored for the water transport field, to automatically identify key entities and their categories in the query. For example, it extracts entities such as "port operation" (subject domain), "risk warning" (application scenario), and "structured data" (data format) from the above query, and marks their location and confidence level.

[0062] Next, the intent classification stage uses a finely tuned BERT model, combined with a water transport business scenario labeling system for multi-classification. For example, "statistically analyze the distribution of accident data of a certain waterway in the past five years" is classified as "statistical analysis" intent, and "find unstructured reports of port operations in 2024" is classified as "data retrieval" intent, with a classification accuracy rate of over 92%.

[0063] Finally, the system generates query parsing results carrying intent feature vectors: keywords are converted into weight vectors using the TF-IDF algorithm, entity categories are encoded using one-hot encoding, and intent types are converted into label encodings, integrating these to form a 256-dimensional intent feature vector. For fuzzy queries, the system uses a context completion mechanism, combined with user profile preferences, to supplement entity information, ensuring the parsing results are complete and accurate.

[0064] S503. Based on the query parsing results and the scene lens parameter system, the intent feature vector is transformed into a lens parameter set containing subject domain filtering parameters, data form filtering parameters, scene weight parameters, quality threshold parameters, time window parameters, spatial range parameters, relationship depth parameters, sorting basis parameters, view form parameters, aggregation granularity parameters, language parameters, and privacy filtering parameters using intent parameter mapping rules. In step S503, the subject domain filtering parameters are as follows: based on the subject domain entity set in the query parsing results, select the 1-2 entities with the highest confidence as the core filtering conditions; if not explicitly specified, refer to the user profile subject domain preference weight vector and select the subject domains with the top two weights as the default filtering conditions (e.g., if the user prefers port operation 0.6 and maritime management 0.8, the default filtering conditions are maritime management and port operation).

[0065] Data form filtering parameters: Extract data form entities from the query parsing results; if not explicitly specified, the form with the higher weight is selected as the default filtering condition based on the user profile data form preference weight; cross-form queries include multiple forms.

[0066] Scenario weight parameter: Set the basic weight of the application scenario entity in the query parsing results to 1.0, multiply it by the scenario preference weight corresponding to the user profile to obtain the comprehensive scenario weight; if there are multiple scenario entities, take the one with the highest weight as the main scenario weight, and the rest as auxiliary weights to participate in the subsequent sorting.

[0067] Quality threshold parameter: If the query implies data quality requirements, it will be automatically mapped to quality level A; if the user has no explicit requirements, it will be based on the user profile's historical data access records. If the user frequently accesses level A data, the default threshold will be A, otherwise it will be B; decision support intents will be forcibly set to A.

[0068] Time window parameter: Extracts time entities from query parsing results and converts them into time intervals; if the query has no time information, it refers to the user's historical query time range or the system's default time window of the past 2 years; supports automatic conversion between relative and absolute time.

[0069] Spatial range parameter: Identifies geospatial entities in the query and maps them to predefined geographic boundary coordinates; if the query has no spatial information, a default range is set according to the user profile spatial preferences; cross-regional queries are extended to the national water transport sector; fuzzy spatial entity semantic matching is supported.

[0070] Relationship depth parameter: Set according to the intent classification results. Data retrieval intent corresponds to depth 1, association analysis intent corresponds to depth 3, and statistical report intent corresponds to depth 2. If users frequently view multi-level related data, appropriately increase the relationship depth (maximum not exceeding 5).

[0071] Sorting parameters: Prioritize query sorting requirements; if no specific requirements are specified, set a default sort based on user profile sorting preferences; if preferences are unclear, sort by value weight in descending order by default, and if weights are the same, sort by quality level and update time in that order.

[0072] View format parameters: Map the view according to the intent type. Data retrieval corresponds to the list view, statistical analysis corresponds to the table view, trend prediction corresponds to the line / bar chart view, and correlation analysis corresponds to the graph view. If the user prefers a visual view, the corresponding visual format will be selected first. Manual switching of view format is supported.

[0073] Aggregation granularity parameters: determined based on time window and intent type. The time window is year for annual aggregation and month for monthly aggregation. Statistical analysis intents use medium granularity by default, while trend prediction intents use fine granularity. If the user explicitly requires a specific granularity, it can be used. An interface for dynamically adjusting aggregation granularity parameters is supported.

[0074] Language parameters: Automatically recognizes the language of the query text or refers to the user's historical language settings; Chinese is the default, and international users can switch to English; supports multilingual mixed queries with automatic adaptation.

[0075] Privacy filtering parameters: Based on the permission levels set for user roles, internal researchers can access the full data, external collaborators can filter confidential data (such as core parameters for port security), and ordinary users can only access public data; at the same time, secondary filtering is combined with privacy tags (public, internal, confidential, etc.) to ensure data access compliance; for sensitive data, some fields are automatically hidden (de-identification processing).

[0076] Finally, all the above parameters are integrated into a lens parameter set, with each parameter carrying a source identifier (query parsing / user profile / default rule) so that the parameter weights can be dynamically adjusted based on user feedback.

[0077] S504. Based on the lens parameter set, multidimensional orthogonal semantic star map and orthogonal projection index, the graph traversal query engine performs multidimensional constraint filtering and association expansion on the cell nodes and relation edges in the semantic star map according to the lens parameters, and generates a set of candidate cell results that satisfy the lens parameter constraints. The entire process is divided into four core stages: initial screening, spatiotemporal filtering, correlation expansion and comprehensive sorting. Each stage combines lens parameters to achieve accurate querying.

[0078] Initial screening phase: Based on the subject domain, data form, and quality threshold parameters of the lens parameters, initial unit cell nodes are quickly screened from the top-level sorted list of the orthogonal projection index. By matching subject domain labels, data form, and quality level, an initial candidate unit cell set is obtained. Using a hierarchical storage structure, the screening time complexity is controlled to the O(logN) level.

[0079] Spatiotemporal filtering stage: For each cell in the initial candidate cell set, extract the spatial range and timestamp information, match it with the spatial range and time window parameters of the lens parameters, filter out cell that does not meet the spatiotemporal conditions, and ensure the spatiotemporal accuracy of the result.

[0080] Association Expansion Phase: Based on the relationship depth parameters of the lens parameters, the spatiotemporally filtered unit cells are association-extended. Starting from the target unit cell, the relationship edges of the multidimensional orthogonal semantic star map are traversed to extend to the associated unit cells at the specified depth. During the extension, spatiotemporal filtering and quality threshold parameters are applied simultaneously to ensure the validity of the results, and the associated unit cell path information is recorded.

[0081] Comprehensive Ranking Stage: Combining the scene weights of the lens parameters with the ranking criteria parameters, the expanded candidate cell set is comprehensively ranked. First, the scene matching score of each cell is calculated, and then the cells are ranked according to the ranking criteria parameters, taking into account multiple dimensions such as scene score, quality level, and update timestamp.

[0082] Finally, a candidate cell result set is generated, containing information such as unique ID, metadata core summary, association path, and comprehensive ranking score, providing data support for the dynamic catalog view. Meanwhile, the graph traversal query engine supports parallel computing optimization, controlling the response time of large-scale semantic star map queries to within 1 second through distributed node sharding processing, meeting real-time interaction requirements.

[0083] S505. Based on the candidate cell result set and view shape parameters, the view rendering engine renders the candidate cell result set into one of the following: list view, tree view, network graph view, timeline view, or map view, according to the view shape parameters, to generate a dynamic directory view oriented to the business scenario.

[0084] In step S505, firstly, in the list view, the view rendering engine arranges the core metadata (unique ID, subject domain, etc.) of candidate cells in a table format, with each row corresponding to one cell, supporting custom column operations; for data retrieval queries, download links and preview buttons are also displayed for easy resource retrieval. Secondly, the tree view is rendered based on the cell hierarchical relationship, extending child nodes from the subject domain as the root node and attaching candidate cells, allowing users to browse and locate resources layer by layer. Next, the network graph view adopts a force-oriented layout, with cells as nodes and relationships as edges. Node colors distinguish subject domains, and edge thickness reflects the relationship strength. Hovering the mouse displays a metadata summary, and clicking a node highlights the relationship path. Then, the timeline view generates a bar chart with the time window interval as the horizontal axis and the number of cells as the vertical axis. Clicking a bar expands the cells within that time unit and sorts them by update time, suitable for queries such as trend prediction. The map view maps the cell spatial range parameters to a GIS layer, using icons to distinguish data forms. The icon size corresponds to the comprehensive ranking score, and clicking an icon displays metadata details, supporting filtering of specific spatial resources. Finally, the view rendering engine supports responsive adaptation, automatically adjusting the layout and interaction methods according to the device, providing a view switching interface, retaining the filtering and sorting states during switching, and generating view description documentation for complex queries.

[0085] S506. Based on the dynamic directory view, collect user clicks, favorites, downloads and ratings of the dynamic directory view, and send the behavioral data back to the user profile model to update the user's multi-dimensional preference weight vector in real time, generating an updated user profile set.

[0086] In step S506, the behavior data acquisition module captures user interaction behavior from multiple dimensions: click records include target cell ID, view shape, click location, and dwell time (over 5 seconds is considered a valid interest behavior); collection of associated cell lists, timestamps, and notes; download statistics include data shape, file format, and frequency; and evaluations include star ratings, text content, and object. All data is timestamped and device-identified to ensure traceability.

[0087] User profile updates employ a weighted incremental learning strategy, assigning differentiated weight coefficients to different behaviors (0.3 for download, 0.2 for favorites, and 0.1 for clicks), mapping behaviors to preference dimensions. For example, downloading structured data for "Maritime Management" increases the corresponding subject domain and data format preference weights by 0.05 and 0.03, respectively; a 5-star rating increases the preference weight of the scenario to which the cell belongs by 0.02.

[0088] Real-time updates are achieved using streaming frameworks (such as Flink), triggering updates as soon as valid behavioral data is generated, avoiding delays. Anomaly detection is performed during updates (e.g., frequent clicks on the same cell are considered invalid operations) to ensure accuracy, while privacy information is anonymized to comply with security standards.

[0089] The updated profiles are stored in a distributed columnar database, supporting low-latency queries and high-concurrency access. Each profile includes a multi-dimensional preference weight vector, historical behavior summaries, and the latest interaction time. The system regularly manages profile versions, retaining versions from the past three months to facilitate preference trend analysis and optimize personalized recommendations.

[0090] Optionally, the self-updating instructions and value assessment results for generating the directory system include: S601. Based on dynamic directory view and multidimensional orthogonal semantic star map, the value transmission network takes the decision target as the root node and performs multi-level accessibility traversal of all knowledge cells along the fusion relationship edge in the multidimensional orthogonal semantic star map. The value transmission coefficient from each cell to the decision target is calculated, and a value transmission map containing multi-level value transmission paths and transmission coefficients is generated. The expression for calculating the value transmission coefficient from each unit cell to the decision objective is as follows: in, Representing a knowledge cell To the decision-making goal The value transmission coefficient, Indicates the first A knowledge cell This represents the decision-making target node, which is the root node of the value transmission network. This represents a single directed propagation path. Indicates from the decision-making goal To knowledge cell The set of all directed paths, Representing an edge The weight value attribute reflects the confidence and importance of the relation edge. Representing a path The first A fusion relationship edge, Representing an edge Relationship types, This represents the path length attenuation coefficient, which controls the value attenuation during multi-stage propagation. Representing a path The number of hops (number of edges), i.e., the conduction series. This represents the influence factor of relation type, which modifies the transmission direction and intensity of different semantic relations.

[0091] In step S601, when calculating the value transmission coefficient, the path length attenuation coefficient is set to 0.8-0.95, dynamically adjusted according to the accuracy requirements of the decision-making scenario: 0.95 (weak attenuation) is used for directly related data, and 0.8 (strong attenuation) is used for mining deep associations. The relationship type influence factor is set with different values ​​for different fusion relationship edges: 1.2 for causal relationship edges (enhanced transmission), 0.9 for subordinate relationship edges (weakened transmission), 1.0 for similar relationship edges (neutral transmission), and 1.1 for cross-validation relationship edges (moderate enhancement), ensuring that the impact of different semantic relationships on value transmission conforms to business logic.

[0092] The value transmission graph is constructed using a three-layer verification process: the first layer verifies path validity, eliminating paths with invalid relation edges (weight values ​​< 0.3); the second layer verifies the transmission direction, ensuring unidirectional transmission from the decision-making target node to the knowledge cell node; and the third layer verifies data consistency, filtering paths with a matching degree below 60%. After the graph is generated, it supports sorting cell nodes in descending order of transmission coefficient, facilitating the location of core data resources that contribute the most to the decision-making target.

[0093] S602. Based on the value transmission graph, the self-evolution engine drives the closed-loop iteration of relationship discovery, quality assessment and classification recommendation with a preset iteration cycle. The relationship edge confidence, knowledge cell quality score and recommendation ranking weight in the multidimensional orthogonal semantic star graph are linked and corrected to generate an iterative optimization set containing the relationship update set, quality correction set and recommendation optimization set. In step S602, the closed-loop iterative process of the self-evolutionary engine includes three core stages: relationship discovery, quality assessment, and classification recommendation, as well as iteration cycle control. Relationship discovery stage: Based on paths with abnormal transmission coefficients in the value transmission graph, the engine initiates an association rule mining algorithm to analyze potential semantic relationships between cells. For example, if "port cargo throughput" and "channel capacity" co-occur frequently in multiple high transmission coefficient paths and there are no direct relationship edges in the semantic star graph, "influence association" candidate edges are generated. The initial confidence is calculated based on the co-occurrence frequency and the relevance to the decision-making scenario. For existing relationship edges where the transmission coefficient and edge weight do not match, they are marked as "to be verified" and manual review is triggered.

[0094] Quality assessment phase: Cell quality scores are adjusted by combining value transmission coefficients and user behavior data. If a cell's transmission coefficient is in the Top 10% but the user rating is less than 3 stars, the quality level is downgraded by one level; if the transmission coefficient is in the Top 20% and the number of downloads is ≥100, the quality level is upgraded by one level; for low-quality cells that appear ≥3 times in the transmission path, a quality check task is pushed to the data administrator.

[0095] In the classification and recommendation stage: the recommendation model weights are updated using iteratively optimized relationship edges and quality scores. For example, the recommendation weight of the subject domain to which the cell with the high conductivity coefficient belongs is increased, and the ranking rules are adjusted according to user scenario preferences, so that closely related cells are ranked higher. The generated optimized recommendation set includes a scenario weight coefficient update table and other content.

[0096] Iteration cycle control: The preset cycle can be daily (for scenarios with high real-time requirements) or weekly (for regular scenarios). After the cycle ends, the engine automatically executes the relationship update set, quality correction set, and recommendation optimization set, and outputs an iteration report, recording key indicators.

[0097] S603. Based on the iterative optimization set, the self-updating rule engine is used to automatically identify and mark expired cells, low-value cells and conflict records in the knowledge cell set, multidimensional orthogonal semantic star map and orthogonal projection index, and generate a self-updating instruction set of the directory system containing new instructions, correction instructions and deletion instructions. In step S602, the core identification logic of the self-updating rule engine is based on the iterative optimization set and preset threshold triggering: For expired cells, the engine compares the cell metadata update time with the current system time, and marks cells that have exceeded the preset validity period (such as 12 months without updates and no associated update plan) as "expired"; For low-value cells, identification is based on the value transmission graph transmission coefficient being less than 0.2 for three consecutive iterations and the lack of effective interaction in user behavior data (downloads, collections are 0 and click dwell time is <5 seconds); For conflict records, the consistency of the multidimensional orthogonal semantic star map relationship edges, knowledge cell metadata and orthogonal projection index is cross-validated, and cases where the subject domain to which the cell belongs contradicts the semantic star map affiliation relationship are marked as "conflict".

[0098] Post-marking processing rules are executed differently based on type: expired cells automatically generate deletion instructions, which are removed from the knowledge cell set and orthogonal projection index after administrator confirmation; low-value cells generate deletion instructions pending review, with identification criteria for administrator decision-making, and the semantic star map association edges are updated upon confirmation of deletion; conflict records generate correction instructions, specifying conflict fields and suggested correction values, and are pushed to data maintenance personnel for manual verification and updates.

[0099] The self-updating instruction set generation and iterative optimization set are linked: the relation update set adds fused relation edges to the multidimensional orthogonal semantic star map, generating relation addition instructions; the quality correction set adjusts the cell quality level and updates it to the knowledge cell metadata, generating quality correction instructions; the recommendation optimization set changes the sorting rules and updates the orthogonal projection index sorting weights, generating index optimization instructions. All instructions include a unique identifier, target object, operation type, and execution priority, with the priority being conflict correction > quality update > expired deletion > low-value processing.

[0100] The self-update instruction execution process adopts a combined "automatic + manual" mode: high-priority conflict correction and quality update instructions are executed automatically by the system (with a preset rule matching degree of ≥90%) and an operation log is generated; low-priority deletion instructions require manual review, during which the administrator can view the historical value contribution of the cell and user feedback. After the instruction is executed, the system verifies the integrity of the directory system to ensure that there are no missing or redundant records, and feeds the update results back to the self-evolution engine as the input for the next round of iteration.

[0101] S604. Based on the self-updating instruction set and value transmission map, the decision support value of all knowledge cells is quantitatively calculated and classified into levels, and value assessment results for various business scenarios are generated by combining the recommended optimization set in the iterative optimization set.

[0102] In step S604, the quantitative calculation of decision support value adopts a multi-dimensional weighted fusion model, expressed as: V=λ1×C+λ2×U+λ3×Q. Where V is the decision support value score of the knowledge cell; C is the value transmission coefficient calculated in step S601; U is the comprehensive user behavior score (calculated based on the behavioral data from step S506, by normalizing the download frequency × 0.4 + the number of collections × 0.3 + the average rating × 0.2 + the number of effective clicks × 0.1, with all normalized values ​​mapped to the [0,1] interval); Q is the corrected cell quality score from step S602; λ1, λ2, and λ3 are weighting coefficients, with default values ​​of 0.5, 0.3, and 0.2 respectively, which can be dynamically adjusted according to the business scenario (e.g., when the decision-making scenario places greater emphasis on data quality, λ3 is increased to 0.3).

[0103] The value level is divided into five levels based on the score range of V: ​​Level S (V≥0.8) is core decision support resource, Level A (0.6≤V<0.8) is important support resource, Level B (0.4≤V<0.6) is routine support resource, Level C (0.2≤V<0.4) is potential value resource, and Level D (V<0.2) is low value resource. After the level is divided, the system marks Level S cells with the "core resource" label and increases their storage priority; Level D cells are included in the low value cell identification range of step S603; for cells that are Level S across multiple business scenarios, a global core resource library is established to support cross-scenario sharing.

[0104] The generation of value assessment results for various business scenarios requires integration with the recommended optimization set in the iterative optimization set. First, based on the scenario weight coefficients of the recommended optimization set, knowledge cells suitable for the current scenario are selected. Second, cells are sorted in descending order of value score, and their transmission paths, quality levels, and user feedback are labeled. Finally, a scenario-specific assessment report is generated, including statistics on the distribution of cells at each level, a list of core cells, and resource optimization suggestions (such as supplementing D-level cell association data and strengthening S-level cell relationship mining). The value assessment results are applied in three ways: first, they are synchronized to the dynamic directory view recommendation module, prioritizing the display of S-level and A-level cells; second, they are fed back to the self-evolution engine as a key basis for the next round of iterative relationship discovery and quality assessment; and third, they are pushed to data administrators and business decision-makers in the form of visual reports to support their development of data resource optimization strategies.

[0105] All user data involved in this application has been authorized, and the acquisition, processing, and transmission comply with legal and regulatory requirements, and necessary confidentiality measures have been taken.

[0106] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A waterborne science data resource catalog system construction method, characterized in that, The method includes: Access to multi-source heterogeneous data in the field of water transport science, including hydrological observation data, waterway survey data, port operation data, ship AIS trajectory data, remote sensing image data, numerical model output data, and text report data, to form a raw data resource pool for water transport science; Based on the original data resource pool of water transport science, semantic decomposition rules and large language models are used to generate a set of knowledge cells carrying DNA encoding and metadata; Based on knowledge cell sets, using graph neural networks and a three-layer relationship discovery strategy of rules, statistics and reasoning, we automatically model the semantic associations between knowledge cells to generate a multi-dimensional orthogonal semantic star map. Based on the multidimensional orthogonal semantic star map, the structured constraints and coordinate mapping of all cells in three dimensions—subject domain, data form, and application scenario—are performed using a three-dimensional orthogonal coordinate system to generate an orthogonal projection index of knowledge cells in three-dimensional space. Based on the multidimensional orthogonal semantic star map and orthogonal projection index, the user profile model and scene lens parameter system are used to perform natural language parsing of user query intent and transform it into lens parameters to generate a dynamic directory view for business scenarios. Based on dynamic directory views and multidimensional orthogonal semantic star maps, the value transmission network is used to calculate the multi-level value transmission coefficients from knowledge cells to decision targets. The self-evolution engine drives the closed-loop iteration of relationship discovery, quality assessment and classification recommendation to generate self-updating instructions and value assessment results for the directory system.

2. The method for constructing a catalog system of water transport scientific data resources as described in claim 1, characterized in that, The generation of knowledge cell sets carrying DNA encoding and metadata includes: The multi-source heterogeneous data from the original data resource pool of water transport science are structured and type-determined to generate an initial set of data fragments containing observation data, model data, remote sensing data, text data, rule data, and metadata. Based on the initial set of data fragments, a large language model adjusted for the water transport domain is used for semantic understanding. The subject domain affiliation, keywords, abstract and semantic fingerprint vector of each fragment are automatically extracted to generate a fragment annotation set carrying semantic group metadata. Based on the fragment annotation set, a hybrid decomposition strategy of rule engine and large language model is used to atomize the initial data fragments according to the judgment criteria of semantic completeness, indivisibility, unique identification, quality assessment and correlation traceability, and generate a set of candidate cells that meet the knowledge cell judgment criteria. Based on the candidate unit cell set, using the domain code, subclass code, temporal granularity code, spatial granularity code, quality grade code, lineage tag code, semantic fingerprint hash and check bit encoding rules in the DNA coding system, a unique data DNA code is generated for each candidate unit cell, and a set of unit cell identifiers carrying the DNA code is generated. Based on the cell identifier set, using a metadata model that includes identity group, semantic group, quality group, lineage group, permission group and value group, the identity metadata, semantic metadata, quality metadata, lineage metadata, permission metadata and value metadata of each cell are automatically filled, generating a knowledge cell set carrying complete metadata; Based on a set of knowledge cells carrying complete metadata, the completeness, accuracy, timeliness, and source credibility of each cell are quantitatively scored. Cells with scores below a preset threshold are marked as pending review and pushed to the domain expert review interface. After review and approval, the final set of knowledge cells carrying DNA encoding and metadata is generated.

3. The method for constructing a catalog system of water transport scientific data resources as described in claim 1, characterized in that, Generating a multidimensional orthogonal semantic star map includes: Based on the knowledge cell set, a graph neural network is used to perform a preliminary traversal and feature extraction of the potential semantic associations between all cells, generating an initial candidate set of relations containing candidate relation edges and candidate edge confidence. Based on the initial candidate relation set, the rule matching strategy is used to automatically confirm the three relation types of candidate edges that meet the criteria of bloodline metadata matching, spatial inclusion matching, and semantic fingerprint cosine similarity exceeding the preset threshold, namely, derivation, synonymy, and composition, and generate a rule-confirmed relation edge set. Based on the rule-confirmed relation edge set and the remaining candidate edges in the initial relation candidate set that were not confirmed by the rules, a statistical analysis strategy is used to calculate the Pearson correlation coefficient and evaluate the mutual information value of the unit cell pairs in the remaining candidate edges. Edges with an absolute value of correlation coefficient greater than n are confirmed as strong correlations, and edges with consistent time series trends are confirmed as evolutionary relationships, thus generating a statistically confirmed relation edge set. Based on the rule-confirmed relation edge set, the statistically confirmed relation edge set, and the remaining candidate edges that are still unconfirmed in the initial relation candidate set, semantic reasoning is performed using a large language model adjusted for the water transport domain. The semantic determination of supporting and contradictory relations is performed on the cell pairs in the remaining candidate edges to generate a reasoning confirmed relation edge set. Based on the rule-confirmed relation edge set, the statistically confirmed relation edge set, and the reasoning-confirmed relation edge set, a relation fusion engine is used to perform conflict detection and confidence-weighted fusion of the three types of relation edges. Relationship type attribute, weight value attribute, and discovery source attribute are added to each edge to generate a fused relation edge set. Based on the set of fused relation edges and the set of knowledge cells, a graph database modeling method is used to treat all knowledge cells as nodes and fused relation edges as directed edges. Orthogonal constraints and coordinate mapping are performed along three dimensions: subject domain, data form, and application scenario to generate a multidimensional orthogonal semantic star map.

4. The method for constructing a catalog system of water transport scientific data resources as described in claim 1, characterized in that, The orthogonal projection index of the generated knowledge cell in three-dimensional space includes: Based on the multidimensional orthogonal semantic star map, a three-dimensional orthogonal coordinate system is used to define the subject domain axis, data form axis and application scenario axis, and a discretized scale and threshold range are set for each axis to generate a three-dimensional orthogonal coordinate framework. Based on all knowledge cells in the three-dimensional orthogonal coordinate framework and multi-dimensional orthogonal semantic star map, coordinate mapping rules are used to map the subject domain affiliation, data form type and application scenario label of each knowledge cell to three-dimensional spatial coordinate points, generating a set of three-dimensional coordinates of the cell. Based on the three-dimensional coordinate set of the unit cell, the orthogonal projection algorithm is used to orthogonally project all coordinate points onto the subject domain data form plane, the subject domain application scenario plane and the data form application scenario plane, respectively, generating three sets of two-dimensional projection point sets. Based on three sets of two-dimensional projection points, an index structure is established for each set of projection points using the R-tree spatial indexing method, generating three sets of R-tree indexes, including subject domain data form R-tree index, subject domain application scenario R-tree index, and data form application scenario R-tree index. Based on three sets of R-tree indexes, the three sets of R-tree indexes are associated and merged using an index fusion engine. Each index record is then assigned cell DNA encoding, quality grade, and value weight as index attributes to generate a fused index set. Based on the fused index set, the coordinate consistency verification rules are used to perform conflict detection and deduplication on all index records. The index records that pass the verification are then sorted in descending order of value weight to generate an orthogonal projection index of knowledge cells in three-dimensional space.

5. The method for constructing a catalog system of water transport scientific data resources as described in claim 1, characterized in that, Generating dynamic catalog views tailored to specific business scenarios includes: Based on the multidimensional orthogonal semantic star map and orthogonal projection index, the user profile model is used to load the current user's role identity, subject domain preference, data form preference, scenario preference and historical behavior records to generate a user profile set containing the user's multidimensional preference weight vector. Based on user profile sets, the system performs word segmentation, entity recognition, and intent classification on the user's input query intent text, automatically extracts query keywords, target subject domains, target data forms, and target application scenarios, and generates query parsing results carrying intent feature vectors. Based on the query parsing results and the scene lens parameter system, the intent feature vector is transformed into a lens parameter set that includes subject domain filtering parameters, data form filtering parameters, scene weight parameters, quality threshold parameters, time window parameters, spatial range parameters, relationship depth parameters, sorting basis parameters, view form parameters, aggregation granularity parameters, language parameters, and privacy filtering parameters using intent parameter mapping rules. Based on the lens parameter set, multidimensional orthogonal semantic star map and orthogonal projection index, the graph traversal query engine performs multidimensional constraint filtering and association expansion on the cell nodes and relation edges in the semantic star map according to the lens parameters, and generates a set of candidate cell results that satisfy the lens parameter constraints. Based on the candidate cell result set and view shape parameters, the view rendering engine renders the candidate cell result set into one of the following: list view, tree view, network graph view, timeline view, or map view, generating a dynamic directory view tailored to business scenarios.

6. A method for constructing a catalog system of water transport scientific data resources as described in claim 5, characterized in that, Generating dynamic catalog views tailored to specific business scenarios also includes: Based on the dynamic directory view, user clicks, favorites, downloads and ratings of the dynamic directory view are collected, and the behavioral data is sent back to the user profile model to update the user's multidimensional preference weight vector in real time, generating an updated user profile set.

7. The method for constructing a catalog system of water transport scientific data resources as described in claim 1, characterized in that, The self-updating instructions and value assessment results for generating the directory system include: Based on dynamic directory view and multidimensional orthogonal semantic star map, the value transmission network is used to perform multi-level accessibility traversal of all knowledge cells along the fusion relationship edge in the multidimensional orthogonal semantic star map with the decision target as the root node. The value transmission coefficient from each cell to the decision target is calculated, and a value transmission map containing multi-level value transmission paths and transmission coefficients is generated. Based on the value transmission graph, a self-evolutionary engine is used to drive the closed-loop iteration of relationship discovery, quality assessment and classification recommendation with a preset iteration cycle. The relationship edge confidence, knowledge cell quality score and recommendation ranking weight in the multidimensional orthogonal semantic star graph are linked and corrected to generate an iterative optimization set containing a relationship update set, a quality correction set and a recommendation optimization set. Based on iterative optimization sets, the self-updating rule engine is used to automatically identify and mark expired cells, low-value cells and conflict records in knowledge cell sets, multidimensional orthogonal semantic star maps and orthogonal projection indexes, and generate a self-updating instruction set of the directory system containing addition instructions, correction instructions and deletion instructions. Based on the self-updating instruction set and value transmission graph, the decision support value of all knowledge cells is quantitatively calculated and classified into levels. Combined with the recommended optimization set in the iterative optimization set, value assessment results for various business scenarios are generated.

8. A method for constructing a catalog system of water transport scientific data resources as described in claim 7, characterized in that, The expression for calculating the value transmission coefficient from each unit cell to the decision objective is as follows: in, Representing a knowledge cell To the decision-making goal The value transmission coefficient, Indicates the first A knowledge cell Indicates the decision-making target node. This represents a single directed propagation path. Indicates from the decision-making goal To knowledge cell The set of all directed paths, Representing an edge The weight value attribute, Representing a path The first A fusion relationship edge, Representing an edge Relationship types, This represents the path length attenuation coefficient. Representing a path The number of jumps, Indicates the relationship type influencing factor.