Water affair data representation and governance method and system based on large model

By generating a unified semantic representation vector based on a large model, the problem of difficulty in integrating multimodal water data is solved, enabling intelligent detection, quality diagnosis and lineage tracing, and improving the automation of data governance and the accuracy of decision-making.

CN122153388APending Publication Date: 2026-06-05NINGBO DONGHAI GRP CORP +2

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NINGBO DONGHAI GRP CORP
Filing Date
2026-02-12
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies lack a unified semantic representation space for deep integration of multimodal water data, making it difficult to semantically align and correlate different modalities of data, thus failing to meet the needs of smart water management for deep data integration and intelligent governance.

Method used

By employing a large model-based approach, a unified semantic representation vector is generated through modality recognition, feature extraction, cross-modal mapping and alignment. Based on this vector, automated governance instructions are generated to achieve intelligent exploration, quality diagnosis and lineage tracing of water data.

Benefits of technology

It enables intelligent detection, quality diagnosis, and lineage tracing of water data, improving the automation level and decision-making accuracy of data governance, and meeting the needs of smart water management for deep data integration and intelligent governance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122153388A_ABST
    Figure CN122153388A_ABST
Patent Text Reader

Abstract

The application relates to a large model-based water affair data representation and treatment method and system, and relates to the field of water affair data processing, which comprises the following steps: collecting multi-source heterogeneous data in a water affair system; performing modal recognition on the multi-source heterogeneous data to determine the data modal type corresponding to each data; matching a corresponding special encoder based on the data modal type; calling the matched special encoder to perform feature extraction on each modal data to generate each modal feature vector; inputting each modal feature vector into a preset shared projection layer to perform cross-modal mapping and alignment to generate a unified semantic representation vector; generating an automatic treatment instruction based on the unified semantic representation vector; and executing the automatic treatment instruction to complete intelligent exploration, quality diagnosis or bloodline tracing of the water affair data. The application has the effect of meeting the needs of intelligent water affairs for data deep fusion and intelligent treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water data processing, and in particular to a method and system for water data representation and management based on a large model. Background Technology

[0002] Water affairs data refers to monitoring data, pipeline data, work order data, and image and video data that originate from SCADA, GIS, business systems, and inspection terminals in the water affairs industry and have different forms and structures such as time series, spatial coordinates, natural language text, and pixel matrices.

[0003] Currently, data governance in this field mainly relies on traditional ETL tools to process structured data and attempts to aggregate multi-source data through data platforms. For heterogeneous data with multiple modalities such as time series, spatial, text, and images, existing technical solutions mostly adopt independent processing or simple correlation models.

[0004] However, such methods lack a unified semantic representation space for deep fusion of multimodal data. This makes it difficult to achieve true semantic alignment and association between data from different modalities, hindering cross-modal joint queries, in-depth analysis, and automated governance. Consequently, they fall short of meeting the needs of smart water management for deep data fusion and intelligent governance, and require further improvement. Summary of the Invention

[0005] To meet the needs of smart water management for deep data fusion and intelligent governance, this invention provides a method and system for water data representation and governance based on a large model.

[0006] Firstly, the present invention provides a method for water resources data characterization and management based on a large model, employing the following technical solution: A water data representation and governance method based on a large model includes: Collect multi-source heterogeneous data from the water system; Modality identification is performed on multi-source heterogeneous data to determine the data mode type corresponding to each data; Based on the data mode type, match the corresponding dedicated encoder; The matching dedicated encoder is invoked to extract features from each modal data to generate feature vectors for each modality; Each modal feature vector is input into a pre-defined shared projection layer for cross-modal mapping and alignment to generate a unified semantic representation vector; Generate automated governance instructions based on a unified semantic representation vector; Execute automated governance commands to complete intelligent detection, quality diagnosis, or lineage tracing of water data.

[0007] By employing the aforementioned technical solution, feature extraction is performed on data from various modalities using a dedicated encoder. Then, cross-modal semantic alignment and mapping are achieved through a shared projection layer, fusing multi-source heterogeneous data into a unified semantic representation vector. This enables automated understanding of data semantics, discovery of correlations, problem diagnosis, and generation of governance instructions. Consequently, intelligent exploration, quality diagnosis, and lineage tracing of water data are realized, improving the automation level and decision-making accuracy of data governance and meeting the needs of smart water management for deep data fusion and intelligent governance.

[0008] Optionally, methods for generating unified semantic representation vectors are also included: Based on the feature vectors of each modality, linear transformation and dimensionality reduction are performed through a shared projection layer to generate dimensionality-reduced feature vectors for each modality; Based on the dimensionality-reduced feature vectors of each modality, a contrastive learning algorithm is used to calculate the cross-modal feature similarity and cross-modal feature difference. The transformation parameters of the shared projection layer are adjusted by combining cross-modal feature similarity and cross-modal feature difference to generate optimized transformation parameters; Based on the optimized transformation parameters and the feature vectors of each modality, a unified semantic representation vector is generated by mapping and fusing them through a shared projection layer.

[0009] By adopting the above technical solution, linear transformation and dimensionality reduction processing are performed on the feature vectors of each modality to obtain the dimensionality-reduced feature vectors of each modality. Then, the similarity and difference of cross-modal features are calculated to optimize the transformation parameters and finally generate a unified semantic representation vector, thereby improving the accuracy and efficiency of water data governance.

[0010] Optionally, methods for generating automated governance instructions may also be included: Perform semantic understanding on the unified semantic representation vector to generate initial semantic labels; Based on the initial semantic tags, a pre-defined water affairs knowledge graph is retrieved to obtain related business rules and historical cases; Data quality diagnostic results are generated by combining related business rules, historical cases, initial semantic tags, and unified semantic representation vectors. Based on the data quality diagnostic results, conduct a governance impact assessment to determine governance strategies; The governance strategy is input into a preset decision module for calculation and optimization to generate automated governance instructions.

[0011] By adopting the above technical solution, initial semantic labels are obtained by semantic understanding of the unified semantic representation vector. Then, based on the initial semantic labels, relevant business rules and historical cases are retrieved from the water affairs knowledge graph. This information is then combined to generate data quality diagnosis results. Based on the diagnosis results, the governance impact is assessed and governance strategies are determined. Finally, the decision-making module transforms the strategies into automated governance instructions, thereby realizing the automation and intelligence of water affairs data governance.

[0012] Optional, also includes: Based on a unified semantic representation vector, modal data subvectors and associated water event identifiers are determined; Based on the water event identifier, the modal data sub-vectors associated with the same water event identifier are aggregated into a multimodal sub-vector set for the current event; Based on the water event identifier, retrieve the set of causal rules and key modality definitions corresponding to the event type from the water knowledge graph; The theoretical consistency relationships between different modal subvectors in a multimodal subvector set are calculated based on a set of causal rules. The actual semantic correlation between the subvectors within a multimodal subvector set is calculated. The consistency deviation value is obtained by combining the actual semantic relevance with the theoretical consistency relationship. When the consistency deviation value exceeds the preset conflict threshold, or when a necessary sub-vector in the set is detected to be missing according to the definition of the critical modality, a data conflict or critical missing is determined to have occurred.

[0013] By adopting the above technical solution, the modal data sub-vectors and water affairs event identifiers are obtained by analyzing the unified semantic representation vector. Then, the sub-vectors of the same event are aggregated into a multimodal sub-vector set, and the corresponding causal rules and key modal definitions are retrieved from the water affairs knowledge graph. The theoretical consistency relationship and the actual semantic correlation are calculated, and the consistency deviation value is obtained. Finally, data conflicts or key missing data are determined, thereby realizing the automated and accurate discovery and location of water affairs multimodal data consistency problems.

[0014] Optional, emergency decision-making methods may also be included: In response to the determination of data conflict or key missing data, the credibility of evidence is evaluated through an adversarial game mechanism based on the multimodal sub-vector set and the causal rule set to generate credibility weights; The multimodal subvector set is weighted according to credibility weights to generate a weighted evidence vector set; Generate cross-modal virtual evidence based on a weighted set of evidence vectors and a set of causal rules; Integrate weighted evidence vector sets with virtual evidence to generate probabilistic path graphs; The minimum regret decision criterion is applied based on probabilistic path graphs to generate emergency governance scripts, and the emergency governance scripts are executed to complete emergency response operations for events corresponding to water incident identifiers.

[0015] By adopting the above technical solutions, dynamic credibility assessment and weighting of multimodal evidence are performed, and virtual evidence is generated in combination with causal rules to supplement or correct conflicting and missing data. Then, real and virtual evidence are integrated to construct a probabilistic event development path map, and the minimum regret value decision criterion is applied to generate the optimal emergency action script, thereby improving the response speed to water emergency events under conditions of data uncertainty or conflict.

[0016] Optionally, a method for generating credibility weights may also be included: Input the multimodal subvector set into the preset corresponding evidence claim module to generate the evidence statement vector for each subvector; Based on the evidence statement vector, a self-assessment is performed through the corresponding evidence claim module to generate a self-assessment credibility score; The evidence statement vector is aggregated and input into the preset evidence adjudication module, and combined with the causal rule set for logical verification to obtain the logical verification result; Based on the logical verification results, the evidence adjudication module generates a global quality score and intermodal conflict indicators for each piece of evidence. Credibility weights are generated based on self-assessed credibility scores, global quality scores, and intermodal conflict indicators.

[0017] By adopting the above technical solution, multimodal subvectors are input into the evidence assertion module to generate evidence statement vectors and self-evaluate their credibility. These vectors are then aggregated into the evidence adjudication module and logically verified using causal rules to assess the overall quality of the evidence and identify intermodal conflicts. Finally, credibility weights are obtained, thereby improving the accuracy of evidence fusion and decision-making in scenarios with data conflicts and missing data.

[0018] Optionally, methods for generating cross-modal virtual evidence are also included: The target modality of the virtual evidence to be generated is determined based on water event identifiers, causal rule sets, and credibility weights, and high-credibility evidence is extracted from the weighted evidence vector set. Based on highly credible evidence and a set of causal rules, counterfactual conditional propositions about the target modality are constructed. Virtual raw data is generated based on counterfactual conditional propositions; Feature extraction and semantic space alignment are performed on virtual raw data to generate and fuse cross-modal virtual evidence.

[0019] By adopting the above technical solution, counterfactual conditional propositions are constructed by screening target modalities and high-credibility evidence, and virtual raw data of the target modalities are generated based on the propositions. Then, through feature extraction and semantic alignment, these data are transformed into fusionable cross-modal virtual evidence. Thus, when data conflicts or omissions occur, the system can proactively reason and generate supplementary information that conforms to domain logic, thereby enhancing the robustness and completeness of decision-making.

[0020] Optional, security measures may also be included: Based on multi-source heterogeneous data and a pre-set water safety database, mimicry bait data is generated; The mimicry decoy data is injected into the data bus and transmitted in a mixed manner with real business data, and the sequence of access requests to the data on the data bus is collected. Behavioral semantic analysis is performed on access request sequences to obtain abnormal semantic features; When the abnormal semantic features meet the preset attack judgment conditions, security protection is carried out using the preset protection processing method.

[0021] By adopting the above technical solution, the system first generates mimicry decoy data and injects it into the data bus, then collects access request sequences from the mixed data stream, performs behavioral semantic analysis on the sequences to extract abnormal semantic features, and finally triggers active protection when the attack judgment conditions are met, thereby achieving security protection for the water data system.

[0022] Optionally, the protective treatment method includes: Attack sources are traced based on abnormal semantic features and access request sequences to generate attack source profiles; Based on the attack source profile and water security database, fake business data streams were generated; The level of security enhancement is determined based on the semantic features of anomalies and the water security reservoir. The control data bus directs the fake business data stream back to the attack source, while simultaneously raising the security level of subsequent access requests from the attack source based on the security enhancement level, and recording this attack in the water security database.

[0023] By adopting the above technical solution, the attack source is tracked and a profile is generated by analyzing abnormal semantic features and access request sequences. Based on the profile, a fake business data stream is dynamically generated to target and interfere with the attacker. At the same time, the corresponding security enhancement level is determined and implemented according to the attack characteristics. Finally, the fake data stream is fed back to the attack source and the security control level of its subsequent access is improved, thereby improving data security.

[0024] Secondly, this application provides a water data characterization and governance system based on a large model, employing the following technical solution: A water data representation and governance system based on a large model includes: The acquisition module is used to collect multi-source heterogeneous data; The memory is used to store the program that implements a water data representation and governance method based on a large model; The processor is used to load and execute programs stored in memory.

[0025] In summary, this application includes at least one of the following beneficial technical effects: 1. By calling a dedicated encoder to extract features from data of each modality, and then using a shared projection layer to achieve cross-modal semantic alignment and mapping, multi-source heterogeneous data is fused into a unified semantic representation vector. This enables automated understanding of data semantics, discovery of correlations, diagnosis of problems, and generation of governance instructions, thereby achieving intelligent exploration, quality diagnosis, and lineage tracing of water data. This improves the automation level and decision-making accuracy of data governance, meeting the needs of smart water management for deep data fusion and intelligent governance. 2. By analyzing the unified semantic representation vector, modal data sub-vectors and water affairs event identifiers are obtained. Then, the sub-vectors of the same event are aggregated into a multimodal sub-vector set. The corresponding causal rules and key modal definitions are retrieved from the water affairs knowledge graph to calculate the theoretical consistency relationship and the actual semantic correlation, and obtain the consistency deviation value. Finally, data conflicts or key missing data are determined, thereby realizing the automated and accurate discovery and location of water affairs multimodal data consistency problems. 3. By analyzing abnormal semantic features and access request sequences, the attack source is traced and a profile is generated. Based on this profile, fake business data streams are dynamically generated to target and interfere with the attacker. At the same time, the corresponding security enhancement level is determined and implemented according to the attack characteristics. Finally, the fake data streams are directed back to the attack source and the security control level of its subsequent access is improved, thereby improving data security. Attached Figure Description

[0026] Figure 1 This is a flowchart of a method for water data representation and governance based on a large model. Detailed Implementation

[0027] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0028] Reference Figure 1 This application discloses a method for water data characterization and management based on a large model, comprising the following steps: S10: Collect multi-source heterogeneous data from the water system.

[0029] Multi-source heterogeneous data refers to data sets originating from different business systems within the water industry and possessing different formats and structural characteristics. In this embodiment, the following four types of data are specifically collected: Time-series monitoring data refers to a sequence of monitoring values ​​with timestamps that are acquired in real time through the Modbus / OPC protocol interface of the SCADA system, such as pipeline pressure values, flow values, and pump current values.

[0030] Spatial vector data refers to the spatial element information of the pipeline network obtained through the WFS service interface that conforms to the OGC standard of the GIS platform, including the geometric coordinates, attributes and topological connections of facilities such as pipeline segments, valves, and water meters.

[0031] Unstructured text data refers to text information obtained by calling business systems through the enterprise service bus, such as complaint descriptions in customer service work orders and text notes in inspection records.

[0032] Image and video data refers to binary media files obtained through the WebSocket interface of the streaming media service, such as photos of pipe corrosion and videos of water accumulation in wells uploaded by inspection personnel using mobile terminals.

[0033] To ensure secure data transmission and meet the protection requirements of critical information infrastructure, this embodiment integrates the national cryptographic algorithm system at the data acquisition and transmission layer. Specifically, all types of data transmitted through the data bus are symmetrically encrypted using the SM4 algorithm; when establishing a secure transmission channel, the SM2 algorithm is used for asymmetric encryption and authentication. This encryption mechanism is implemented throughout the entire data chain from SCADA, GIS, and business systems to the system's data bus, achieving unified and high-strength security protection for both structured data and unstructured media files, ensuring that the data transmission process meets the requirements of Level 3 Information Security Protection.

[0034] S11: Perform modality identification on multi-source heterogeneous data to determine the data modality type corresponding to each data.

[0035] Data modality type refers to the basic category divided according to the essential characteristics and processing methods of data. In this step, the system automatically identifies and classifies the collected data into four preset data modality types by parsing the data's protocol header, metadata tags, and content characteristics: Time series type corresponds to monitoring data organized in the form of a time sequence; spatial type corresponds to vector data with geographic coordinates and topological relationships; text type corresponds to unstructured data composed of natural language characters; and image type corresponds to visual data composed of pixel matrices.

[0036] Specific modality recognition algorithms (such as rule-based matching or lightweight classification models) are well known to those skilled in the art and will not be elaborated here.

[0037] S12: Match the corresponding dedicated encoder based on the data modality type.

[0038] A dedicated encoder is a neural network model specifically trained for the features of a particular data modality, used to transform raw data into high-dimensional feature vectors. The system has a pre-set encoder resource pool and performs matching based on the data modality type determined in S11. Match a time-series encoder to time-series data, a spatial encoder to spatial data, a text encoder to text-type data, and an image encoder to image-type data.

[0039] The specific encoder selection (such as choosing an architecture like Transformer, GNN, BERT, or CNN) and model parameters are determined by those skilled in the art based on actual business needs, and will not be elaborated here.

[0040] S13: Call the matching dedicated encoder to extract features from each modal data to generate feature vectors for each modality.

[0041] Each modal feature vector refers to a fixed-dimensional array of real numbers output by a dedicated encoder, which contains deep semantic information of the original data. In this step, the system calls the encoders matched by S12 for parallel processing to extract features from the original data: the temporal encoder processes the temporal monitoring data to generate temporal feature vectors, the spatial encoder processes the spatial vector data to generate spatial feature vectors, the text encoder processes the unstructured text data to generate text feature vectors, and the image encoder processes the image and video data to generate image feature vectors. Feature extraction techniques are well known to those skilled in the art and will not be elaborated upon here.

[0042] The dimensions of each feature vector (e.g., 768 dimensions) are determined by the encoder design and will not be elaborated here.

[0043] S14: Input the feature vectors of each modality into a preset shared projection layer for cross-modal mapping and alignment to generate a unified semantic representation vector.

[0044] A shared projection layer is a pre-trained neural network layer whose function is to map feature vectors from different modalities to a unified vector space.

[0045] A unified semantic representation vector is a vector that integrates multimodal semantic information after mapping and alignment.

[0046] In this step, the system inputs the modal feature vectors generated in S13 into the shared projection layer to obtain a unified semantic representation vector.

[0047] The specific method for generating the unified semantic representation vector will be explained in detail in subsequent sections S20 to S23, and will not be repeated here.

[0048] S15: Generate automated governance instructions based on unified semantic representation vectors.

[0049] Automated governance instructions refer to a set of task instructions generated intelligently by the system that can drive downstream business systems or guide manual operations.

[0050] The specific method for generating automated governance instructions will be explained in detail in subsequent sections S30 to S34, and will not be repeated here.

[0051] S16: Execute automated governance instructions to complete intelligent detection, quality diagnosis, or lineage tracing of water data.

[0052] By calling the corresponding external system application interface through the process execution engine, the automated governance instructions are transformed into specific, executable operation sequences, thereby completing the intelligent exploration, quality diagnosis, or lineage tracing of water affairs based on a certain characterization vector data.

[0053] The specific execution process includes: Intelligent Detection: By executing pre-defined data query and pattern analysis tasks in the instructions, the system automatically discovers abnormal patterns, potential risks, or new knowledge correlations from multi-source data. For example, if the instructions include a trend analysis command for historical pressure data of a specific area, the system will automatically output a detection report on whether there is persistent low pressure in that area after execution.

[0054] Quality Diagnosis: By executing pre-defined rule verification, causal reasoning, or comparative analysis tasks within the instructions, the system performs root cause analysis and quality assessment on discovered data anomalies. For example, if the instructions include a command to call a hydraulic model for simulation verification, after execution, the system automatically generates a diagnostic conclusion and confidence level regarding whether a pressure anomaly at a certain point is due to pipeline leakage or scheduling operations.

[0055] Lineage tracing: By executing pre-defined data link tracing tasks in the command, the system automatically locates the source of specific data, records the processing steps it has undergone, and analyzes its downstream impact. For example, if the command includes a reverse tracing command for a key indicator in a report, after execution, the system automatically generates a complete processing link map of that indicator from the original sensor data to the final report.

[0056] The pre-defined data query and pattern analysis tasks, rule verification, causal reasoning or comparison analysis tasks, and data link tracing tasks in the instructions are all a series of executable logical units pre-defined and encapsulated by those skilled in the art based on common scenarios and business objectives of water data governance. These task logics define the specific technical actions (such as querying a specific database, calling a certain analysis algorithm, or verifying a certain business rule) required to achieve governance objectives such as exploration, diagnosis, and tracing. Their specific content and implementation can be designed and configured by those skilled in the art according to actual governance needs, and will not be elaborated here.

[0057] The interface integration method between the process execution engine and external systems (such as SCADA, GIS, and work order systems) is a conventional technical approach in this field and will not be elaborated here.

[0058] It also includes methods for generating unified semantic representation vectors: S20: Based on the feature vectors of each modality, linear transformation and dimensionality reduction are performed through a shared projection layer to generate dimensionality-reduced feature vectors for each modality.

[0059] The dimensionality-reduced feature vectors for each modality refer to the feature vectors with lower dimensions (e.g., reduced from 768 dimensions to 256 dimensions) output after the high-dimensional feature vectors generated by S13 are input into the shared projection layer and subjected to linear transformation matrix calculation and compression in this layer. This vector is a preliminary projection representation of the feature vectors of each modality in a unified semantic space. The specific mathematical processes of linear transformation and dimensionality reduction are common knowledge in this field and will not be elaborated here.

[0060] S21: Based on the dimensionality reduction feature vectors of each modality, the cross-modal feature similarity and cross-modal feature difference are calculated using a contrastive learning algorithm.

[0061] Cross-modal feature similarity refers to the degree of closeness between the dimensionality-reduced feature vectors of the same water entity (such as the same pipeline) in different modalities (such as the time-series pressure data and spatial location data of the pipeline) in a unified semantic space.

[0062] Cross-modal feature dissimilarity refers to the degree of distance between modal dimensionality-reduced feature vectors of different water entities (such as different pipelines) or unrelated entities in a unified semantic space.

[0063] In this step, the system calculates cross-modal feature similarity and cross-modal feature difference based on preset sample pairs (positive sample pairs: vectors of different modalities of the same entity; negative sample pairs: vectors of different entities), and constructs a loss function (such as InfoNCELoss) accordingly. The specific process of the contrastive learning algorithm is common knowledge in this field and will not be elaborated here.

[0064] S22: Combine cross-modal feature similarity and cross-modal feature difference to adjust the transformation parameters of the shared projection layer to generate optimized transformation parameters.

[0065] The optimized transformation parameters refer to the weight matrix updated in the shared projection layer after optimization algorithms such as backpropagation and gradient descent. The optimization objective is to maximize the similarity of cross-modal vectors of the same entity, while minimizing the similarity (or maximizing the difference) of vectors of different entities.

[0066] In this step, the gradient is calculated through backpropagation based on the loss function calculated in S21, and the weights of the shared projection layer are updated. After multiple iterations, when the model converges or reaches the preset training objective (set in advance by those skilled in the art, and will not be elaborated here), the final weights are the optimized transformation parameters. The specific parameter optimization algorithm (such as the Adam optimizer) is common knowledge in the art and will not be elaborated here.

[0067] S23: Based on the optimized transformation parameters and the feature vectors of each modality, a shared projection layer is used to map and fuse them to generate a unified semantic representation vector.

[0068] The original modal feature vectors generated in S13 are input into a shared projection layer loaded with the optimized transformation parameters obtained in S22. After final mapping by this layer, a unified-dimensional feature vector that deeply integrates multimodal semantic information is output, namely the unified semantic representation vector of S14.

[0069] It also includes methods for generating automated governance instructions: S30: Perform semantic understanding on the unified semantic representation vector to generate initial semantic labels.

[0070] Initial semantic labels refer to the structured semantic description fragments output by the governance agent after decoding and initially understanding the unified semantic representation vector generated by S23. They are expressed in the form of entity-attribute-state or event-type-parameter. For example, for a given representation vector, the agent can generate labels such as: [Entity: DMA-Partition-05], [Attribute: Minimum Nighttime Traffic], [State: Continuously Abnormally High]. The specific process of semantic understanding (such as vector decoding based on attention mechanisms) is an internal working mechanism of the governance agent and will not be elaborated upon here.

[0071] The governance agent is a software agent based on a large language model architecture, with deep knowledge enhancement and functional modifications tailored to the water sector. Its core is a domain-adaptive large language model foundation. This foundation is built upon general-purpose large models (such as the Qwen series) and continuously pre-trained and fine-tuned using massive amounts of water sector professional literature, technical regulations, historical work orders, and case reports. This allows it to deeply understand water sector professional terminology and business logic, such as production-sales differences, DMA zoning, and hydraulic models.

[0072] This intelligent agent embeds a Retrieval Enhanced Generation (RAG) module, enabling real-time queries of the water resources knowledge graph to obtain accurate domain knowledge. Simultaneously, it is structurally integrated with multiple sub-agent modules oriented towards governance tasks, primarily including: Data exploration agent: responsible for automatically identifying the business meaning of data fields and performing semantic annotation; Quality diagnosis agent: responsible for combining generative adversarial networks and rule engines to discover and diagnose data anomalies; Lineage tracing agent: responsible for automatically tracing the source, processing process and impact chain of data.

[0073] The aforementioned intelligent agent modules work collaboratively to complete the complex task from semantic understanding to decision generation. The specific architectural design and internal module collaboration mechanism of the governance intelligent agent are to be implemented by those skilled in the art according to system requirements, and will not be elaborated here.

[0074] S31: Retrieve the preset water affairs knowledge graph based on the initial semantic tags to obtain related business rules and historical cases.

[0075] The water affairs knowledge graph is a structured domain knowledge base that pre-stores relationships among water network entities, equipment attributes, business rules, physical causal logic, and historical governance cases. The construction of this knowledge graph is based on publicly available standards, technical procedures, and historical data in the water industry. It was pre-constructed by those skilled in the art according to specific business scopes. The specific methods for knowledge extraction, fusion, and storage are standard techniques for constructing knowledge graphs in this field and will not be elaborated upon here.

[0076] Related business rules refer to the standardized processing principles retrieved from the water knowledge graph that are related to the state described by the initial semantic tags. For example, the retrieved rule is: if the minimum nighttime flow in the DMA partition exceeds 20% of the daytime average for three consecutive days, the leakage verification process is triggered.

[0077] Historical cases refer to records of past governance events, their handling methods, and results that are retrieved from the water affairs knowledge graph and are similar to or related to the current initial semantic tags in terms of business scenarios.

[0078] In this step, the water resources knowledge graph is retrieved using the initial semantic tags as query criteria to obtain relevant business rules and historical cases that can be referenced to support decision-making. The knowledge graph retrieval technology is a standard practice in this field and will not be elaborated upon here.

[0079] S32: Combine related business rules, historical cases, initial semantic labels, and unified semantic representation vectors to generate data quality diagnostic results.

[0080] Data quality diagnostic results refer to the judgments made by the governance agent regarding the data itself or the business problems it reflects, after comprehensively considering the initial semantic labels of S30, the related business rules retrieved by S31, and historical cases, and then further associating them with the original unified semantic representation vector for deep reasoning. These results include problem localization, root cause analysis, and confidence assessment.

[0081] The governance agent first selects the most relevant rules from associated business rules based on the problem type described by the initial semantic labels, and determines whether the current state (represented by a unified semantic representation vector) conforms to the rule, thus arriving at a rule determination conclusion. Simultaneously, the agent calculates the similarity between the unified semantic representation vector and the feature vectors of each historical case, identifying the most similar case and obtaining a case reference conclusion. Subsequently, the agent compares and merges the rule determination conclusion with the case reference conclusion: if they match, the conclusion is directly adopted; if there is a conflict or complementarity, a comprehensive preliminary diagnostic conclusion is generated according to a preset priority. Finally, based on the degree of rule matching, the level of case similarity, and the consistency of evidence, the agent calculates the comprehensive confidence level of the preliminary diagnostic conclusion through a built-in confidence assessment module. Ultimately, the preliminary diagnostic conclusion and the comprehensive confidence level are combined and output as the data quality diagnostic result. The specific algorithmic models for rule matching, case similarity calculation, conclusion fusion, and confidence assessment within the governance agent can be selected or implemented by those skilled in the art using suitable machine learning models, and will not be elaborated upon here.

[0082] S33: Conduct a governance impact assessment based on the data quality diagnosis results to determine governance strategies.

[0083] Governance strategy refers to the specific action guidelines and plans formulated based on the data quality diagnosis results and after assessing the potential business impact of the problem.

[0084] Based on the problem identification, root cause analysis, and confidence assessment in the data quality diagnosis results, a pre-set strategy knowledge base can be queried to obtain governance strategy entries corresponding to the current diagnosis conclusion.

[0085] The strategy knowledge base contains pre-set action guidelines, resource allocation plans, and execution priorities for different water-related problem scenarios. Its content is formulated and maintained in advance by those skilled in the art based on business management standards and historical experience, and will not be elaborated here.

[0086] For example, for the data quality diagnosis result "Problem location: DMA-partition-05 suspected physical leakage; confidence level: 91%", after querying the strategy knowledge base, the matching governance strategy is: "Strategy type: emergency leakage handling; core action: immediately dispatch personnel to the site to listen for and locate the leakage; execution priority: high; required time limit: within 24 hours; responsible team: leakage detection team A". S34: Input the governance strategy into the preset decision module for calculation and optimization to generate automated governance instructions.

[0087] The decision-making module refers to the functional component in the system responsible for transforming abstract governance strategies into specific, executable, and resource- and logic-optimized automated governance instructions. It can be implemented based on a rule engine and optimization algorithms.

[0088] The decision-making module receives the governance strategy output by S33, and, in conjunction with the real-time status of the current system (such as personnel scheduling and equipment availability), execution costs, and timeliness requirements, refines and optimizes the actions in the strategy, ultimately outputting structured automated governance instructions. The specific implementation method of the decision-making module is to be selected by those skilled in the art based on the system architecture, and will not be elaborated upon here.

[0089] Also includes: S40: Determine modal data subvectors and associated water event identifiers based on unified semantic representation vectors.

[0090] Modal data sub-vectors refer to feature vector fragments corresponding to a single original data source, separated from a unified semantic representation vector. By utilizing the data source metadata carried during the generation of the unified semantic representation vector (which records the correspondence between each feature dimension and the original sensor ID, data table record ID, etc.), the system can reverse-engineer the vector into multiple independent modal data sub-vectors. The recording and use of data source metadata are standard techniques in data processing within this field and will not be elaborated upon here.

[0091] A water event identifier is a unique number assigned by the system to a business anomaly requiring attention and handling (such as a pipe burst or water quality fluctuation), used to link all data related to this anomaly. A pre-defined event detection model is used to analyze the unified semantic representation vector in real time. This model comprehensively determines whether a new business anomaly event (such as a pipe burst or water pollution) has occurred based on the mutability, correlation, and similarity to historical anomaly patterns of the vector's features in the unified semantic space. When an anomaly is determined to have occurred, a new unique number is generated as the water event identifier, and all modal data sub-vectors decomposed from S40 that are semantically highly related to the anomaly are automatically associated with this identifier. The event detection model is pre-defined by those skilled in the art and will not be elaborated upon here.

[0092] S41: Based on the water event identifier, aggregate the modal data sub-vectors associated with the same water event identifier into a multimodal sub-vector set for the current event.

[0093] A multimodal subvector set refers to a vector group formed by bringing together all modal data subvectors identified in S40 that are associated with the same water event identifier.

[0094] Using the established relationships in S40, and with the water event identifier as the query key, all modal data sub-vectors bound to it are retrieved from storage and aggregated to form a multimodal sub-vector set corresponding to the event.

[0095] S42: Based on the water event identifier, retrieve the set of causal rules and key modality definitions corresponding to the event type from the water knowledge graph.

[0096] A causal rule set refers to a set of rules retrieved from a water knowledge graph that describes the causal or temporal logical relationships that should be followed between various physical quantities or states during the occurrence and development of such events (such as "pipe burst" or "water pollution").

[0097] By using the event type corresponding to the water event identifier as the query condition, all causal logic rules related to that type of event are retrieved from the water knowledge graph and returned, thus obtaining a set of causal rules.

[0098] The definition of key modalities refers to several data modal types and their expected roles that are essential for the effective analysis and judgment of such events, retrieved from the water knowledge graph.

[0099] By identifying the event type corresponding to the water event identifier, a predefined list of data modalities and their functional descriptions that must be relied upon when analyzing such events are obtained from the water knowledge graph, thereby obtaining the definition of key modalities.

[0100] The water resources knowledge graph stores the mapping relationship between different event types and their corresponding causal logic rules and key modality definitions. By querying this mapping relationship, the set of causal rules and key modality definitions can be retrieved.

[0101] S43: Calculate the theoretical consistency relationship between different modal subvectors in a multimodal subvector set based on a set of causal rules.

[0102] Theoretical consistency relationship refers to the expected correlation that the data sub-vectors of each modality of the event should satisfy in terms of value, state or trend after logical deduction based on the causal rule set retrieved by S42.

[0103] By transforming each logical rule in the causal rule set (e.g., "If a pipe bursts at point X, the pressure at point Y downstream should decrease by ΔP after time T") into a mathematical description of the specific quantitative relationships, state order, or trend of change that should exist between related sub-vectors (such as the pressure sub-vector at point X and the pressure sub-vector at point Y) in the multimodal sub-vector set, the theoretical consistency relationship that should be satisfied between all related sub-vectors can be calculated.

[0104] The specific methods for rule transformation and relation derivation are designed and implemented by those skilled in the art based on the rule expression forms (such as natural language and structured logic), and will not be elaborated here.

[0105] S44: Calculate the actual semantic correlation between subvectors within a multimodal subvector set.

[0106] Actual semantic relevance refers to the strength of the actual relevance of each vector in the multimodal sub-vector set aggregated by S41 at the current moment, obtained by statistical analysis.

[0107] By invoking a preset correlation calculation algorithm, specific sub-vector pairs to be examined in the multimodal sub-vector set are processed to obtain the actual correlation strength value between them, i.e., the actual semantic correlation degree. The specific selection and implementation of the correlation calculation algorithm are well known to those skilled in the art and will not be elaborated here.

[0108] S45: Combine the actual semantic relevance with the theoretical consistency relationship to obtain the consistency deviation value.

[0109] The consistency deviation value refers to the numerical result obtained by quantitatively comparing the actual semantic correlation calculated in S44 with the theoretical consistency relationship derived in S43 (such as calculating Euclidean distance, KL divergence, or other discrepancy measures, which should be selected by those skilled in the art and will not be elaborated here). The larger the value, the more serious the deviation between the actual observation data and the physical laws of the domain.

[0110] S46: When the consistency deviation value exceeds the preset conflict threshold, or when a necessary sub-vector in the set is detected to be missing according to the definition of the key modality, it is determined that a data conflict or key missing data has occurred.

[0111] The conflict threshold is a pre-set numerical limit used to determine whether the consistency deviation has reached an unacceptable level and conflict handling needs to be initiated.

[0112] Data conflict refers to a situation where the facts or states reflected by different modalities contradict each other, and the degree of contradiction (reflected by the consistency deviation value) exceeds a reasonable range (conflict threshold).

[0113] Critical missing refers to the complete absence or invalidation of data for one or more modalities that are necessary for the current event analysis, according to the definition of critical modalities.

[0114] When the consistency deviation value exceeds the conflict threshold, it is determined that a data conflict has occurred.

[0115] The consistency deviation value is compared with the conflict threshold. If the deviation value is greater than the threshold, it is determined that there is a significant contradiction between the multimodal data of the current event, that is, a data conflict has occurred.

[0116] Simultaneously, based on the definition of critical modalities, the system checks one by one whether the sub-vectors corresponding to each necessary modality listed in the definition exist in the multimodal sub-vector set. If a sub-vector of any necessary modality is found to be missing, it is determined that a critical missing condition has occurred.

[0117] Based on the above comparison and verification results, the status of the current event is marked. The size of the conflict threshold is set by those skilled in the art according to business fault tolerance requirements, and will not be elaborated here.

[0118] It also includes emergency decision-making methods: S50: In response to the determination of data conflict or key missing data, the credibility of evidence is evaluated through an adversarial game mechanism based on the multimodal sub-vector set and the causal rule set to generate credibility weights.

[0119] Credibility weight refers to the weight value assigned to each sub-vector in the multimodal sub-vector set after evaluation through an adversarial game mechanism. The higher the weight, the more credible the evidence is in the current conflict or missing context.

[0120] In this step, the system initiates an adversarial game mechanism (details in S60 to S64), using the causal rule set as the arbitration basis to dynamically evaluate each sub-vector in the multimodal sub-vector set and assign a credibility weight to each piece of evidence. The specific training and operation process of the adversarial game mechanism is a well-known technique implemented by those skilled in the art based on the combination of game theory and machine learning, and will not be elaborated here.

[0121] S51: The multimodal subvector set is weighted according to the credibility weight to generate a weighted evidence vector set.

[0122] The weighted evidence vector set refers to the new vector set obtained by multiplying each sub-vector in the multimodal sub-vector set by its corresponding confidence weight.

[0123] Each subvector in the multimodal subvector set is numerically weighted to generate a weighted evidence vector set.

[0124] S52: Generate cross-modal virtual evidence based on a weighted set of evidence vectors and a set of causal rules.

[0125] Cross-modal virtual evidence refers to the feature vectors corresponding to virtual data generated through counterfactual reasoning based on a weighted set of evidence vectors and a set of causal rules, used to fill in missing information or explain conflicts. The specific generation method of cross-modal virtual evidence will be explained in detail in subsequent sections S70 and S73, and will not be repeated here.

[0126] S53: Integrate weighted evidence vector sets with virtual evidence to generate probabilistic path graphs.

[0127] A probabilistic path graph is a graph that describes multiple possible future development paths of an event and their corresponding probabilities, generated by a time-series causal inference model after integrating a weighted evidence vector set and virtual evidence.

[0128] The weighted evidence vector set and the generated virtual evidence are input into a time-series causal inference model. This model infers multiple possible future development paths of events based on the physical laws of the domain and assigns a probability of occurrence to each path, thereby forming a probabilistic path graph. The specific construction method of the time-series causal inference model is a conventional technical choice for those skilled in the art based on probabilistic graphical models or deep time-series models, and will not be elaborated here.

[0129] S54: Based on the probabilistic path graph, apply the minimum regret decision criterion to generate an emergency governance script, and execute the emergency governance script to complete the emergency response operation for the event corresponding to the water incident identifier.

[0130] An emergency response script is a sequence of instructions that contains a series of specific emergency response steps, generated based on a probabilistic path graph and using the minimum regret decision criterion.

[0131] For each possible future in the probabilistic path graph, the potential consequences and losses (i.e., regret values) of different emergency action plans are evaluated. The action sequence that maximizes the regret value among all possible futures and minimizes it is selected and packaged into an executable emergency governance script. Subsequently, the system's process execution engine automatically executes the various instructions in this script (such as dispatching personnel, controlling equipment, and issuing notifications) to complete the emergency response to the real events corresponding to the water event identifiers defined by S40. The specific calculation method of the minimum regret value decision criterion is common knowledge in the field of decision theory and will not be elaborated here.

[0132] It also includes methods for generating credibility weights: S60: Input the multimodal sub-vector set into the preset corresponding evidence claim module to generate the evidence statement vector for each sub-vector.

[0133] The corresponding evidence assertion module refers to a set of pre-trained neural network sub-modules, each specifically adapted to a data modality (temporal, spatial, text, image). Its function is to receive the feature sub-vectors of the corresponding modality and, through an internal forward inference network, decode these vectors into a structured intermediate vector containing specific factual assertion semantics. The specific network structure and training method of the corresponding evidence assertion module are determined by those skilled in the art based on actual needs and will not be elaborated here.

[0134] An evidence statement vector is a fixed-dimensional array of real numbers output by the corresponding evidence claim module. This vector is used to encode the core semantic claim implied by the specific modality of data represented by its input subvectors.

[0135] The evidence statement vector is generated by invoking the forward reasoning process of the corresponding evidence claim module. This process maps the modal feature sub-vectors to a semantic space dedicated to expressing the claim, thereby forming a structured semantic encoding. The forward reasoning process is common knowledge in neural network models in this field and will not be elaborated here.

[0136] S61: Based on the evidence statement vector, perform self-assessment through the corresponding evidence claim module to generate a self-assessment credibility score.

[0137] The self-assessed credibility score is a scalar value used to quantify the reliability of the evidence statement vectors it outputs.

[0138] The self-assessed credibility score is generated through a confidence assessment sub-network embedded within the corresponding evidence claim module. This sub-network takes the evidence statement vector and the module's internal intermediate states as input, and outputs a scalar score after forward computation. The specific structure of the confidence assessment sub-network and its parameter training method are determined by those skilled in the art based on the module design, and will not be elaborated here.

[0139] S62: Aggregate the evidence statement vector and input it into the preset evidence adjudication module, and perform logical verification in combination with the causal rule set to obtain the logical verification result.

[0140] The evidence adjudication module is an arbitration model built on attention mechanisms and graph neural networks, used to comprehensively evaluate the consistency of multimodal evidence.

[0141] The logical verification result refers to the judgment information output by the evidence adjudication module, which is used to indicate the logical conformity between each evidence statement and the set of causal rules, as well as the conflict between different evidence statements.

[0142] The logical verification result is generated through the forward reasoning process of the evidence adjudication module. This process takes the aggregated set of evidence statement vectors and the set of causal rules as input, calculates the correlation weights between different pieces of evidence through its internal attention mechanism, and uses a graph neural network combined with rules for logical reasoning and conflict detection, ultimately outputting a structured verification conclusion. The specific mechanisms involved in the forward reasoning process, such as attention weight calculation and graph reasoning, are common knowledge in this field and will not be elaborated here.

[0143] The specific algorithm flow used to generate logical verification results within the evidence adjudication module is designed and implemented by those skilled in the art based on actual needs, and will not be elaborated here.

[0144] S63: Based on the logical verification results, the evidence adjudication module generates a global quality score and intermodal conflict indicators for each piece of evidence.

[0145] The global quality score refers to the score assigned by the evidence adjudication module to each evidence statement vector, reflecting its relative credibility in the global evidence set.

[0146] Intermodal conflict indicators are indicators used to quantify the degree of contradiction between different modal evidence.

[0147] Based on the logical verification results, the pre-defined scoring function and conflict detection algorithm within the evidence adjudication module are invoked to calculate and output a global quality score for each evidence statement vector. Simultaneously, inter-modal conflict indicators are generated to characterize the contradictory relationships between different modalities of evidence. The specific implementation and parameters of the scoring function and conflict detection algorithm are determined by those skilled in the art based on actual needs and will not be elaborated upon here.

[0148] S64: Generate credibility weights based on self-assessed credibility scores, global quality scores, and intermodal conflict indicators.

[0149] By using a preset weight fusion function, the self-assessed credibility score, global quality score, and intermodal conflict indicators are comprehensively calculated to generate a corresponding credibility weight for each subvector in the multimodal subvector set.

[0150] The construction of the weight fusion function is well known to those skilled in the art and will not be elaborated here.

[0151] It also includes methods for generating cross-modal virtual evidence: S70: Based on water event identifiers, causal rule sets, and credibility weights, determine the target modality of the virtual evidence to be generated, and extract high-credibility evidence from the weighted evidence vector set.

[0152] A target modality refers to a data modality type that is essential for event analysis but is currently missing or whose evidence has low credibility (i.e., its weight is below a preset threshold, which is set in advance by those skilled in the art and will not be elaborated here). By comparing the event type corresponding to the water event identifier with the key modality definition obtained in step S42, modalities that are crucial to the analysis of the event but do not have a corresponding sub-vector in the current multimodal sub-vector set, or whose credibility weight of the corresponding sub-vector is below the preset threshold, are identified and thus determined as target modalities.

[0153] High-credibility evidence refers to one or more evidence sub-vectors selected from the weighted evidence vector set whose corresponding credibility weights exceed a preset credibility threshold (this threshold is set in advance by those skilled in the art and will not be elaborated here). By comparing the credibility weights attached to each sub-vector in the weighted evidence vector set with the credibility threshold, all sub-vectors with weights higher than the threshold are selected as high-credibility evidence for subsequent generation of virtual evidence.

[0154] S71: Construct counterfactual conditional propositions about the target modality based on a set of highly credible evidence and causal rules.

[0155] Counterfactual propositions are hypothetical logical statements that describe the state or value that a target modality should present under hypothetical conditions, based on known, highly credible evidence and domain causal rules.

[0156] By parsing the numerical parameters represented by high-confidence evidence and matching and substituting these parameters with the corresponding logical rules in the causal rule set, a complete hypothetical logical statement is constructed. The specific logic of the matching and substitution is designed and implemented by those skilled in the art based on the rule expression form, and will not be elaborated here.

[0157] S72: Generate virtual raw data based on counterfactual conditional propositions.

[0158] Virtual raw data refers to simulated data that conforms to the data format and statistical characteristics of the target modality. For example, if the target modality is an image, then the virtual raw data is a simulated image; if the target modality is time series data, then it is a set of simulated time series values.

[0159] By using counterfactual conditional propositions as input, a pre-trained generative model corresponding to the target modality is invoked to synthesize virtual raw data that conforms to the logical constraints of the propositions and possesses the data characteristics and distribution of the target modality. The generative model can be a conditional generative adversarial network, a diffusion model, or other data generation model suitable for this modality. Its specific architecture and training method are selected and implemented by those skilled in the art based on the characteristics of the target modality, and will not be elaborated here.

[0160] S73: Perform feature extraction and semantic space alignment on virtual raw data to generate and fuse cross-modal virtual evidence.

[0161] A dedicated encoder matching the target modality is invoked to extract features from the virtual raw data, obtaining a virtual feature vector for the target modality. This virtual feature vector is then input into a shared projection layer, mapped to a unified semantic space, generating a cross-modal virtual evidence vector. Finally, this cross-modal virtual evidence vector is integrated into a weighted evidence vector set for subsequent probabilistic path deduction.

[0162] It also includes safety protection methods: S80: Generates mimicry decoy data based on multi-source heterogeneous data and a pre-set water safety database.

[0163] A water security database is a pre-built knowledge base that stores the network topology of water systems, business data characteristics, known attack patterns, and defense strategies. The water security database is pre-set by those skilled in the art and will not be elaborated upon here.

[0164] Mimicry decoy data refers to highly realistic copies of data that contain non-sensitive or false information, used to actively attract and detect potential attacks.

[0165] Mimicry decoy data is generated by analyzing the characteristic patterns (such as data structure, numerical range, temporal patterns, and communication protocols) of real business data stored in the water security database, and combining this with known attack patterns. Data synthesis techniques are used to generate data that is highly similar in form to real data streams but harmless or misleading in content. Specific data synthesis algorithms and simulation strategies are designed by those skilled in the art based on actual needs and will not be elaborated upon here.

[0166] S81: Inject mimicry decoy data into the data bus and transmit it in a mixed manner with real business data, and collect the access request sequence of data on the data bus.

[0167] A data bus is a common communication channel in a system used to transmit data between different modules or services.

[0168] Real business data refers to the raw or intermediate data collected from the water business system and used in the actual treatment process.

[0169] An access request sequence refers to a series of read, write, or query operations initiated on data transmitted on the data bus, arranged chronologically. By utilizing the bus's built-in logging function, all access requests to data transmitted on the bus (including real business data and decoy data) are captured and recorded in real time. Key fields such as request time, source address, target data identifier, and operation type are extracted and arranged and stored chronologically to form the access request sequence. The methods for collecting and parsing bus logs are standard techniques in this field and will not be elaborated upon here.

[0170] First, mimicry decoy data is injected into the data bus and transmitted in a mixed manner with real business data. This is to actively deceive and detect potential attacks, induce attackers to access the mimicry decoy data, thereby exposing their attack intentions and behavioral characteristics, and collect the access request sequence of data on the data bus for subsequent steps.

[0171] S82: Perform behavioral semantic analysis on the access request sequence to obtain abnormal semantic features.

[0172] Anomaly semantic features refer to a set of structured metrics or vectors used to quantify covert probing attack patterns. These features not only identify explicit malicious requests, but also reveal the attacker's logical intent to semantically probe and correlate different data modalities (such as time-series monitoring, spatial topology, and ticket text).

[0173] The abnormal semantic features are obtained through the following steps: First, the system detects whether there are access records for the mimicry decoy data in the access request sequence. Any such access is directly marked as an attack and corresponding confirmatory features are generated. Second, the semantic association analysis module, built based on the shared projection layer and contrastive learning mechanism trained in the aforementioned steps (S10 to S23), is invoked to map the data objects pointed to by all requests in the sequence to an aligned unified semantic space. The association transition paths and patterns between requests in different modal dimensions such as time sequence, space, and text are analyzed. By calculating the frequency, order, and business logic rationality of cross-modal semantic transitions within a specific time window, quantitative indicators representing probing reconnaissance intentions are extracted. Finally, the above analysis results are combined to generate a structured set of abnormal semantic features.

[0174] The implementation of the semantic association analysis module relies on the unified semantic representation capability trained by this method. Its specific model architecture and algorithm parameters are determined by those skilled in the art based on security analysis requirements, and will not be elaborated here.

[0175] S83: When the abnormal semantic features meet the preset attack judgment conditions, security protection is carried out using the preset protection processing method.

[0176] Attack determination criteria refer to a pre-defined set of rules or model output thresholds used to determine whether abnormal semantic features have reached the level required to be identified as an attack. These attack determination criteria are pre-defined by those skilled in the art and will not be elaborated upon here.

[0177] The protective measures refer to the set of defensive operations performed after an attack has been confirmed. Specific protective measures will be detailed in S90 to S93, and will not be elaborated upon here.

[0178] When abnormal semantic features meet the attack judgment conditions, security protection should be implemented using protective measures.

[0179] Protective measures include: S90: Based on abnormal semantic features and access request sequences, the attack source is traced to generate an attack source profile.

[0180] An attack source profile is a structured document used to comprehensively describe the attacker's identity, technical characteristics, behavioral patterns, and potential intentions.

[0181] Attack source profiling is obtained by aggregating and correlating abnormal semantic features with source addresses, attack behavior patterns, and probing paths in access request sequences. The specific models and algorithms relied upon in the analysis process are configured by those skilled in the art according to security requirements and will not be elaborated here.

[0182] S91: Generate fake business data streams based on attack source profiling and water security database.

[0183] Fake business data streams refer to highly simulated continuous business data sequences used to proactively respond to attacks, consume attackers' resources, and induce them to make misjudgments.

[0184] The fake business data stream is dynamically generated by querying the water security database for fake data templates and generation strategies that match the attack patterns reflected in the attack source profile, and then instantiating parameters based on that profile. The specific content of the data templates, generation strategies, and their matching logic are designed by those skilled in the art based on security countermeasure requirements, and will not be elaborated here.

[0185] S92: Determine the security enhancement level based on abnormal semantic features and the water security database.

[0186] Security enhancement level refers to the system security protection strategy level with different levels of strictness, which is dynamically set according to the severity and characteristics of the attack.

[0187] The security enhancement level can be directly determined by querying a protection strategy level mapping table stored in the water security database that matches the attack patterns and severity represented by the anomaly semantic features. The logic for setting the mapping table and the specific level definitions are formulated by those skilled in the art based on the system security baseline requirements, and will not be elaborated here.

[0188] S93: The control data bus directs the fake business data stream back to the attack source, while simultaneously raising the security level of subsequent access requests from the attack source based on the security enhancement level, and recording this attack in the water security database.

[0189] The control data bus precisely and directionally sends the generated fake business data streams to the attack source address. Simultaneously, based on the determined security enhancement level, the security verification level and review intensity for all subsequent access requests from this attack source are immediately increased. Finally, the attack source profile, abnormal semantic features, handling actions, and effects of this attack event are fully recorded and updated to the water security database for policy optimization and knowledge accumulation. The specific implementation mechanisms of the targeted feedback and security level enhancement strategies are designed by those skilled in the art based on the system architecture and will not be elaborated here.

[0190] Based on the same inventive concept, embodiments of the present invention provide a water data representation and management system based on a large model, comprising: The data acquisition module is used to collect multi-source heterogeneous data and access request sequences; The memory is used to store the program that implements a water data representation and governance method based on a large model; The processor is used to load and execute programs stored in memory.

[0191] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0192] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for water data representation and governance based on a large model, characterized in that, include: Collect multi-source heterogeneous data from the water system; Modality identification is performed on multi-source heterogeneous data to determine the data mode type corresponding to each data; Based on the data mode type, match the corresponding dedicated encoder; The matching dedicated encoder is invoked to extract features from each modal data to generate feature vectors for each modality; Each modal feature vector is input into a pre-defined shared projection layer for cross-modal mapping and alignment to generate a unified semantic representation vector; Generate automated governance instructions based on a unified semantic representation vector; Execute automated governance commands to complete intelligent detection, quality diagnosis, or lineage tracing of water data.

2. The water data representation and management method based on a large model according to claim 1, characterized in that, It also includes methods for generating unified semantic representation vectors: Based on the feature vectors of each modality, linear transformation and dimensionality reduction are performed through a shared projection layer to generate dimensionality-reduced feature vectors for each modality; Based on the dimensionality-reduced feature vectors of each modality, a contrastive learning algorithm is used to calculate the cross-modal feature similarity and cross-modal feature difference. The transformation parameters of the shared projection layer are adjusted by combining cross-modal feature similarity and cross-modal feature difference to generate optimized transformation parameters; Based on the optimized transformation parameters and the feature vectors of each modality, a unified semantic representation vector is generated by mapping and fusing them through a shared projection layer.

3. The water data representation and management method based on a large model according to claim 1, characterized in that, It also includes methods for generating automated governance instructions: Perform semantic understanding on the unified semantic representation vector to generate initial semantic labels; Based on the initial semantic tags, a pre-defined water affairs knowledge graph is retrieved to obtain related business rules and historical cases; Data quality diagnostic results are generated by combining related business rules, historical cases, initial semantic tags, and unified semantic representation vectors. Based on the data quality diagnostic results, conduct a governance impact assessment to determine governance strategies; The governance strategy is input into a preset decision module for calculation and optimization to generate automated governance instructions.

4. The water data representation and management method based on a large model according to claim 3, characterized in that, Also includes: Based on a unified semantic representation vector, modal data subvectors and associated water event identifiers are determined; Based on the water event identifier, the modal data sub-vectors associated with the same water event identifier are aggregated into a multimodal sub-vector set for the current event; Based on the water event identifier, retrieve the set of causal rules and key modality definitions corresponding to the event type from the water knowledge graph; The theoretical consistency relationships between different modal subvectors in a multimodal subvector set are calculated based on a set of causal rules. The actual semantic correlation between each subvector within a multimodal subvector set is calculated. The consistency deviation value is obtained by combining the actual semantic relevance with the theoretical consistency relationship. When the consistency deviation value exceeds the preset conflict threshold, or when a necessary sub-vector in the set is detected to be missing according to the definition of the critical modality, a data conflict or critical missing is determined to have occurred.

5. The water data representation and management method based on a large model according to claim 4, characterized in that, It also includes emergency decision-making methods: In response to the determination of data conflict or key missing data, the credibility of evidence is evaluated through an adversarial game mechanism based on the multimodal subvector set and the causal rule set to generate credibility weights; The multimodal subvector set is weighted according to credibility weights to generate a weighted evidence vector set; Generate cross-modal virtual evidence based on a weighted set of evidence vectors and a set of causal rules; Integrate weighted evidence vector sets with virtual evidence to generate probabilistic path graphs; The minimum regret decision criterion is applied based on probabilistic path graphs to generate emergency governance scripts, and the emergency governance scripts are executed to complete emergency response operations for events corresponding to water incident identifiers.

6. The water data representation and management method based on a large model according to claim 5, characterized in that, It also includes methods for generating credibility weights: Input the multimodal subvector set into the preset corresponding evidence claim module to generate the evidence statement vector for each subvector; Based on the evidence statement vector, a self-assessment is performed through the corresponding evidence claim module to generate a self-assessment credibility score; The evidence statement vector is aggregated and input into the preset evidence adjudication module, and combined with the causal rule set for logical verification to obtain the logical verification result; Based on the logical verification results, the evidence adjudication module generates a global quality score and intermodal conflict indicators for each piece of evidence. Credibility weights are generated based on self-assessed credibility scores, global quality scores, and intermodal conflict indicators.

7. The water data representation and management method based on a large model according to claim 5, characterized in that, It also includes methods for generating cross-modal virtual evidence: The target modality of the virtual evidence to be generated is determined based on water event identifiers, causal rule sets, and credibility weights, and high-credibility evidence is extracted from the weighted evidence vector set. Based on highly credible evidence and a set of causal rules, counterfactual conditional propositions about the target modality are constructed. Virtual raw data is generated based on counterfactual conditional propositions; Feature extraction and semantic space alignment are performed on virtual raw data to generate and fuse cross-modal virtual evidence.

8. The water data representation and governance method based on a large model according to claim 1, characterized in that, It also includes safety protection methods: Based on multi-source heterogeneous data and a pre-set water safety database, mimicry bait data is generated; The mimicry decoy data is injected into the data bus and transmitted in a mixed manner with real business data, and the sequence of access requests to the data on the data bus is collected. Behavioral semantic analysis is performed on access request sequences to obtain abnormal semantic features; When the abnormal semantic features meet the preset attack judgment conditions, security protection is carried out using the preset protection processing method.

9. A water data representation and governance method based on a large model according to claim 8, characterized in that, The protective treatment method includes: Attack sources are traced based on abnormal semantic features and access request sequences to generate attack source profiles; Based on the attack source profile and water security database, fake business data streams were generated; The level of security enhancement is determined based on the semantic features of anomalies and the water security reservoir. The control data bus directs the fake business data stream back to the attack source, while simultaneously raising the security level of subsequent access requests from the attack source based on the security enhancement level, and recording this attack in the water security database.

10. A water data representation and governance system based on a large model, characterized in that, include: The acquisition module is used to acquire multi-source heterogeneous data; A memory for storing a program that implements a water data characterization and governance method based on a large model as described in any one of claims 1 to 9; The processor is used to load and execute programs stored in memory.