Weak model joint enhancement method and system for digital twin hydraulic engineering large model
By integrating multiple weak models and knowledge graphs, the problem of insufficient accuracy and generalization ability of water conservancy digital twin models in complex environments is solved, and more efficient water conservancy system prediction and decision support are achieved.
Patent Information
- Application Number
- CN202511177817.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-12-23
AI Technical Summary
Current water conservancy digital twin models suffer from insufficient model accuracy and weak generalization ability when faced with complex and ever-changing water conservancy environments and massive amounts of data. Traditional single models are unable to comprehensively and accurately describe the complex behavior of water conservancy systems.
By integrating multiple relatively simple weak models and utilizing pseudo-label generation, knowledge graphs, and attention mechanisms, a water conservancy knowledge graph is constructed to achieve information fusion between large and weak models and improve the overall performance of the model.
It improves the prediction accuracy and generalization ability of water conservancy digital twin models for complex water conservancy phenomena, reduces training time and computational resource consumption, improves model development efficiency, and enhances the understanding and processing ability of complex relationships in water conservancy systems.
Smart Images

Figure CN121189132A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water conservancy engineering technology, and in particular to a method and system for joint enhancement of weak models for digital twin water conservancy large models. Background Technology
[0002] As an emerging technology, water conservancy digital twins can monitor, predict, and optimize the operation status of water conservancy projects in real time by constructing a virtual model that corresponds to the real water conservancy system.
[0003] However, current digital twin models for water conservancy often suffer from insufficient model accuracy and weak generalization ability when faced with complex and ever-changing water conservancy environments and massive amounts of data. Traditional single models are insufficient to comprehensively and accurately describe the complex behavior of water conservancy systems, while developing a powerful model that can cover all water conservancy elements faces enormous technical challenges and high costs.
[0004] Therefore, there is a need for a method and system for joint enhancement of weak models in digital twin water conservancy models to effectively integrate multiple relatively simple weak models and improve the overall performance of water conservancy digital twin models. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for joint enhancement of weak models for digital twin water conservancy models. The main technical problem to be solved is to effectively integrate multiple relatively simple weak models through the method and system for joint enhancement of weak models for digital twin water conservancy models, thereby improving the overall performance of the water conservancy digital twin model.
[0006] To achieve the above objectives, this invention provides a joint enhancement method for weak models in a digital twin hydraulic system model. The weak model refers to a subsystem-level prediction model that has been pre-trained on the corresponding hydraulic subsystem and constructed using expert knowledge rules, statistical modeling, or small-scale supervised learning. The method specifically includes the following steps:
[0007] S1. Unified Coding and Input Alignment: The raw input data from different water conservancy subsystems (such as hydrology, water quality, and hydrodynamics) are standardized and coded (e.g., unit conversion, time alignment, and missing data completion). Spatiotemporal alignment is completed through interpolation, resampling, and other methods to obtain a structurally consistent and unlabeled aligned raw dataset. This dataset serves as the input basis for generating pseudo-labels in the weak model in S2.
[0008] S2. Generate pseudo-labels and expand the dataset using weak models: The aligned original data obtained in S1 is divided into different water conservancy subsystems and then input into the corresponding subsystem pre-trained weak models for prediction to obtain subsystem-level pseudo-label outputs; these pseudo-label data are then merged with the original labeled dataset to construct an expanded training set for use by subsequent large models.
[0009] The pre-trained weak model of the subsystem refers to the subsystem-level prediction model that is pre-trained on the corresponding water conservancy subsystem and constructed based on expert knowledge rules, statistical modeling or small-scale supervised learning; the original labeled dataset refers to the measured data of historical field monitoring points or the key variables of the subsystem and their spatiotemporal labels labeled after manual inspection and expert evaluation.
[0010] S3. Constructing a water resources knowledge graph or causal structure to guide learning: Constructing a water resources knowledge graph or causal graph based on domain knowledge to clarify the logical and physical relationships between variables in each subsystem (e.g., upstream flow affects downstream water level, pollution sources affect water quality indicators, etc.); assisting the large model to more effectively absorb the pre-trained weak model and expert knowledge of each subsystem in S2.
[0011] S4. Integrating information from weak and large models for joint modeling: During the training of the large model, an attention mechanism or cue learning structure is introduced to embed the output of the pre-trained weak model in the subsystem of S2 (such as pseudo-labels and hidden features) and the knowledge graph structure into the input of the large model, calculate its attention weight for the current task, and realize the fusion of weak model knowledge and large model representation.
[0012] A further technical solution of the present invention: The original input data from different water conservancy subsystems in step S1 refers to the original input data collected from the hydrological, hydraulic engineering, water regulation, and water quality water conservancy subsystems, such as the time series of hydrology, water quality, and hydrodynamics at each monitoring point; the input data is uniformly encoded using the Z-score normalization method, and the input of the pre-trained weak model of the subsystem is aligned using linear transformation, thereby obtaining a uniformly encoded original dataset with consistent structure and spatiotemporal alignment. The original dataset output in S1 will be used in S2 to generate pseudo-labels and expand the dataset for the weak model;
[0013] Specifically, the unified encoding refers to the encoding from the first... Input data of each subsystem The Z-score standardization method is used for unified coding.
[0014] ①
[0015] In formula ①, This is the original data. For the first Subsystem The mean of each feature,
[0016] Standard deviation, The data is standardized.
[0017] The input alignment specifically refers to the first... Input of a weak model Perform a linear transformation to align it with the target dimension of the larger model:
[0018] ②
[0019] In formula ②, This is the weight matrix. For bias vectors, This is the aligned output.
[0020] A further technical solution of the present invention: the input in step S2 is the aligned original dataset output by S1. Align the original dataset obtained from S1. After dividing into subsystems, the corresponding pre-trained weak models of each subsystem are input into them. Perform predictions and generate subsystem-level pseudo-labels. And by setting a confidence threshold High-confidence pseudo-labels are selected and merged with the original labeled dataset to obtain expanded data. Extended dataset of S2 It will be used in S3 to build knowledge graphs or causal graph structures, and in S4 for use in large models;
[0021] The pseudo-label generation step in step S2 is as follows: align the original dataset obtained in step S1. Using weak models Generate pseudo tags:
[0022] ③
[0023] In formula ③, Pseudo-label This represents a pre-trained weak model whose output is a probability distribution or a direct label.
[0024] A further technical solution of the present invention: the input of S3 is an extended dataset. Based on domain expert knowledge, the knowledge of weak models is transferred and enhanced by constructing a causal graph structure and injecting adapter parameters, thereby obtaining the causal graph adjacency matrix. and the corrected large model output The causal graph structure and adapter parameters obtained in step S3 will be used in step S4 to introduce an attention mechanism, calculate the attention weights between the output of the weak model and the internal features of the large model, and achieve information fusion between the two.
[0025] The causal graph structure learning step in step S3 is as follows: Constructing the adjacency matrix of the hydraulic causal graph. The objective function is defined as follows:
[0026] ④
[0027] In formula ④, For observation data matrices (such as multi-dimensional features like water level, flow velocity, and temperature);
[0028] The Frobenius norm measures the error in data fitting. The L1 regularization term constrains the sparsity of the causal graph. This is an acyclic constraint (refer to DAG constraint) to ensure that the causal graph is acyclic.
[0029] A further technical solution of the present invention: In step S4, the causal graph structure and adapter parameters obtained in step S3 are jointly embedded with the output of the pre-trained weak model of the subsystem (such as pseudo-labels and hidden features) into the input of the large model, and the weights are dynamically allocated through an attention mechanism to obtain the joint features. This achievement will enable information complementarity between large and weak models, thereby improving the prediction accuracy and generalization ability of large-scale digital twin models for water conservancy.
[0030] The attention weight calculation steps for S4 are as follows: Define the internal features of the large model as... The weak model output is Weights are dynamically allocated through an attention mechanism:
[0031] ⑤
[0032] In formula ⑤, A learnable query and key weight matrix; This is a scaling factor to prevent the inner product from becoming too large; The function ensures that attention weights are normalized.
[0033] A preferred technical solution of the present invention: the pseudo-label screening step in step S2 is as follows: setting a confidence threshold. Only retain high-confidence pseudo-labels:
[0034]
[0035] The dataset expansion step included in S2 is as follows: [The original labeled dataset is then expanded]. Compared with the filtered pseudo-label dataset merge: .
[0036] A preferred technical solution of the present invention: The adapter parameter injection step included in S3 is: inserting an adapter layer into the large model. Its parameters are derived from the features of the weak model. Dynamically generated:
[0037]
[0038] The output of the large model is corrected as follows:
[0039] ⑥
[0040] In formula ⑥, This is the original output of the large model. These are prediction results from a weak model.
[0041] The optimization steps included in S3 are: defining a prompt template for the water conservancy field. Fuse the output of the weak model as context:
[0042] The large model generates causal logic based on prompts: .
[0043] A preferred technical solution of the present invention: The information joining step of S4 is as follows: the weak model output is passed through a weight matrix. After mapping, it is weighted and fused with attention weights:
[0044] ⑦
[0045] In formula ⑦, The combined features enable information complementarity between the large and weak models;
[0046] The attention residual enhancement step in S4 is as follows: To mitigate information loss, residual connections are introduced:
[0047] .
[0048] The present invention also provides a weak model joint enhancement system for a digital twin hydraulic system, for executing the aforementioned weak model joint enhancement system for a digital twin hydraulic system. The enhancement system includes a weak model module, a data unification module, an enhancement training module, a knowledge management module, and a feedback and evaluation module: the weak model module, the data unification module, and the enhancement training module are used to execute steps S1 and S2, and the feedback and evaluation module and the knowledge management module are used to execute steps S3 and S4.
[0049] The weak model module consists of multiple water conservancy subsystem modelers. Each water conservancy subsystem modeler is pre-trained on the corresponding water conservancy subsystem and constructed into a corresponding subsystem-level predictive weak model based on expert knowledge rules, statistical modeling, or through small-scale supervised learning. Different water conservancy subsystems include hydrology, hydraulic engineering, water regulation, and water quality.
[0050] The data unification module includes input / output standardization, data preprocessing, and a label fusion unit, which are used to standardize input data from different hydraulic subsystem modelers, remove noise and outliers, and fuse labels to generate an unlabeled dataset.
[0051] The enhanced training module includes a pseudo-label generator, a hint builder, and a fusion optimizer. The pseudo-label generator accepts unlabeled datasets and predicts and generates pseudo-labels for the unlabeled datasets, which are then merged with the original labeled data to form an expanded dataset. The hint builder generates hint information based on knowledge graphs or causal graph structures. The fusion optimizer optimizes the absorption of weak model knowledge during the training of large models.
[0052] The feedback and evaluation module consists of a simulator, an error monitor, and a reinforcement learning optimizer. The simulator simulates the operation of the water conservancy system, the error monitor monitors the error between the model output and the actual situation, and the reinforcement learning optimizer adjusts the model parameters based on the error feedback.
[0053] The knowledge management module includes graph construction, causal relationship modeling, and domain knowledge base, which are used to build and maintain knowledge graphs and causal relationship models in the water conservancy field, and store relevant knowledge.
[0054] By employing the above technical solutions, the present invention provides a method and system for joint enhancement of weak models in digital twin hydraulic systems, which has at least the following beneficial effects:
[0055] 1. This invention integrates information from multiple weak models, fully utilizes the advantages of each subsystem modeler, and compensates for the shortcomings of a single model, thereby improving the accuracy of the large-scale digital twin model for water conservancy in predicting and simulating complex water conservancy phenomena.
[0056] 2. This invention utilizes weak models to generate pseudo-labels to expand the dataset, increasing data diversity and helping to improve the generalization ability of large models in different water conservancy scenarios, enabling them to better cope with various complex situations in actual engineering.
[0057] 3. This invention enables large models to quickly absorb knowledge from weak models through prompting optimization or adapter methods, reducing the training time and computational resource consumption of large models and improving model development efficiency.
[0058] 4. This invention constructs a knowledge graph or causal graph structure, combining professional knowledge in the field of water conservancy with model learning, thereby improving the efficiency of knowledge utilization and enabling the model to better understand and handle complex relationships in the water conservancy system. Attached Figure Description
[0059] Figure 1This is a flowchart of the weak model joint enhancement method for digital twin hydraulic large model according to the present invention. Detailed Implementation
[0060] The present invention will be further described below with reference to the accompanying drawings and embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0061] This invention provides a method for joint enhancement of weak models in a digital twin hydraulic system model. The weak model in this embodiment refers to a subsystem-level prediction model that has been pre-trained on the corresponding hydraulic subsystem and constructed based on expert knowledge rules, statistical modeling, or small-scale supervised learning. Figure 1 The present invention specifically includes the following steps:
[0062] S1. Unified Encoding and Alignment of Multi-Source Data: Raw input data are collected from different water conservancy subsystems (such as hydrology, water quality, and hydrodynamics). These data differ in terms of temporal granularity, spatial resolution, and physical quantity units. To achieve unified modeling, the data is first standardized and encoded (such as unit conversion, time alignment, and missing measurement completion). Then, spatiotemporal alignment is completed through interpolation, resampling, and other methods. Finally, a structurally consistent aligned raw dataset is obtained. This dataset is still unlabeled and not included in the model prediction. It serves as the input basis for generating pseudo-labels for the weak model in S2.
[0063] The input S1 is the original dataset from different water conservancy subsystems (such as time series of hydrology, water quality and hydrodynamics at each monitoring point). The input data is uniformly encoded using the Z-score normalization method, and the input of the weak model is aligned using linear transformation, thereby obtaining a uniformly encoded original dataset with consistent structure and spatiotemporal alignment. The original dataset output by S1 will be used in S2 to generate pseudo-labels and expand the dataset for the weak model.
[0064] Specifically, the unified coding is for the data from the first... Input data of each subsystem The Z-score standardization method is used for unified coding.
[0065] ①
[0066] In formula ①, This is the original data. For the first Subsystem The mean of each feature, Standard deviation, The data is standardized.
[0067] Input alignment specifically refers to the first... Input of a weak model Perform a linear transformation to align it with the target dimension of the larger model:
[0068] ②
[0069] In formula ②, This is the weight matrix. For bias vectors, The output after alignment;
[0070] Table 1. Reference Table for Unified Coding of Input Parameters of Hydraulic Subsystem
[0071] .
[0072] S2 weak model generates pseudo-labels to construct an expanded training set: Each subsystem has a pre-trained independent weak model (obtained based on expert rules, statistical modeling, or small-scale supervised training). To introduce more diverse supervision information into the training of the large model, the aligned raw data obtained in S1 is divided according to subsystems and input into the corresponding pre-trained weak models for prediction, resulting in subsystem-level pseudo-label outputs. Subsequently, these pseudo-label data are merged with the original labeled dataset (the original labeled dataset refers to historical field monitoring point measured data or subsystem key variables and their spatiotemporal labels labeled after manual inspection or expert evaluation), constructing an expanded training set for subsequent use by the large model.
[0073] The input in S2 is the aligned original dataset output by S1. The aligned dataset obtained from S1 After dividing by subsystem, input the corresponding weak models respectively. Perform predictions and generate subsystem-level pseudo-labels. And by setting a confidence threshold High-confidence pseudo-labels are selected and merged with the original labeled dataset to obtain an expanded dataset. Extended dataset of S2 It will be used in S3 to build knowledge graphs or causal graph structures, and in S4 for use in large models;
[0074] The pseudo-label generation step in step S2 is as follows: align the original dataset obtained in step S1. Using weak models Generate pseudo tags:
[0075] ③
[0076] In formula ③, It is a pseudo-tag. This represents a pre-trained weak model whose output is a probability distribution or a direct label.
[0077] Step S2 includes a pseudo-label filtering step: setting a confidence threshold. Only retain high-confidence pseudo-labels: ;
[0078] S2 includes the following dataset expansion steps: Expanding the original labeled dataset... Compared with the filtered pseudo-label dataset merge: ;
[0079] Table 2 Examples of pseudo-label generation and dataset expansion parameters
[0080] .
[0081] S3. Constructing a Water Resources Knowledge Graph or Causal Structure to Guide Learning: To enable the large model to better grasp the logical and causal relationships between variables in water resources subsystems, a water resources knowledge graph or causal graph is constructed based on domain knowledge to clarify the logical and physical relationships between variables in each subsystem (e.g., upstream flow affects downstream water level, pollution sources affect water quality indicators, etc.). This structure, as an explicit knowledge constraint, is integrated into the large model in the form of prompts or structural guidance to help it more effectively absorb the pre-trained weak model and expert knowledge from S2 (the weak model is defined as above).
[0082] The input to S3 is the extended dataset. Based on domain expert knowledge, the knowledge of weak models is transferred and enhanced by constructing a causal graph structure and injecting adapter parameters, thereby obtaining the causal graph adjacency matrix. and the corrected large model output The causal graph structure and adapter parameters obtained in step S3 will be used in step S4 to introduce an attention mechanism, calculate the attention weights between the output of the weak model and the internal features of the large model, and achieve information fusion between the two.
[0083] The causal graph structure learning step in step S3 is as follows: Constructing the adjacency matrix of the hydraulic causal graph. The objective function is defined as follows:
[0084] ④
[0085] In formula ④, For observation data matrices (such as multi-dimensional features like water level, flow velocity, and temperature);
[0086] The Frobenius norm measures the error in data fitting. The L1 regularization term constrains the sparsity of the causal graph. These are acyclic constraints (refer to DAG constraints) to ensure that the cause-effect graph is free of loops;
[0087] The adapter parameter injection step in step S3 is as follows: inserting an adapter layer into the large model. Its parameters are derived from the features of the weak model. Dynamically generated:
[0088]
[0089] The output of the large model is corrected as follows:
[0090] ⑥
[0091] In formula ⑥ This is the original output of the large model. These are prediction results from a weak model.
[0092] The optimization steps included in S3 are: defining a prompt template for the water conservancy field. Fuse the output of the weak model as context:
[0093] The large model generates causal logic based on prompts: .
[0094] Table 3 Key parameters and adapter configuration for water conservancy cause-effect diagram
[0095] .
[0096] S4 integrates information from the weak model and the large model for joint modeling: During the training of the large model, an attention mechanism or cue learning structure is introduced. The output of the weak model in S2 (such as pseudo-labels and hidden features) and the knowledge graph structure are jointly embedded into the input of the large model. The attention weights for the current task are calculated, realizing the fusion of weak model knowledge and large model representation (the weak model is defined as in the previous step). Finally, the large model learns the dynamic relationships of multiple subsystems through joint modeling and has stronger generalization ability and interpretability. In step S4, the causal graph structure and adapter parameters obtained in step S3 are jointly embedded into the input of the large model along with the output of the weak model (such as pseudo-labels and hidden features). The weights are dynamically allocated through the attention mechanism to obtain the joint features. This achievement will enable information complementarity between large and weak models, thereby improving the prediction accuracy and generalization ability of large-scale digital twin models for water conservancy.
[0097] The attention weight calculation step in step S4 is as follows: Define the internal features of the large model as...
[0098] The output of the weak model is Weights are dynamically allocated through an attention mechanism:
[0099] ⑤
[0100] In formula ⑤, A learnable query and key weight matrix; This is a scaling factor to prevent the inner product from becoming too large; The function ensures that attention weights are normalized.
[0101] The information integration step in step S4 is as follows: passing the output of the weak model through the weight matrix. After mapping, it is weighted and fused with attention weights:
[0102] ⑦
[0103] In formula ⑦, The combined features enable information complementarity between the large and weak models;
[0104] The attention residual enhancement steps of S4 are as follows: To mitigate information loss, residual connections are introduced:
[0105] ;
[0106] Table 4. Core parameters of the attention mechanism and application examples in water conservancy scenarios.
[0107] .
[0108] The embodiment also provides a weak model joint enhancement system for a digital twin hydraulic large model. The enhancement system includes a weak model module, a data unification module, an enhancement training module, a knowledge management module, and a feedback and evaluation module. The weak model module, the data unification module, and the enhancement training module are used to perform steps S1 and S2, and the feedback and evaluation module and the knowledge management module are used to perform steps S3 and S4.
[0109] The weak model module consists of multiple water conservancy subsystem modelers, which are used to model different water conservancy subsystems, including hydrology, hydraulic engineering, water regulation and water quality.
[0110] The data unification module includes input / output standardization, data preprocessing, and a label fusion unit, which are used to standardize input data from different hydraulic subsystem modelers, remove noise and outliers, and fuse labels to generate an unlabeled dataset.
[0111] The augmented training module includes a pseudo-label generator, a hint builder, and a fusion optimizer. The pseudo-label generator takes an unlabeled dataset and predicts and generates pseudo-labels for the unlabeled dataset, which are then merged with the original labeled data to form an expanded dataset. The hint builder generates hint information based on the knowledge graph or causal graph structure. The fusion optimizer optimizes the absorption of weak model knowledge during the training of large models.
[0112] The feedback and evaluation module consists of a simulator, an error monitor, and a reinforcement learning optimizer. The simulator simulates the operation of the water conservancy system, the error monitor monitors the error between the model output and the actual situation, and the reinforcement learning optimizer adjusts the model parameters based on the error feedback.
[0113] The knowledge management module includes graph construction, causal relationship modeling, and domain knowledge base, which are used to build and maintain knowledge graphs and causal relationship models in the water conservancy field, and store relevant knowledge.
[0114] The working principle of this invention is as follows: After completing steps S1, S2, S3, and S4, a weak model joint enhancement system for a large-scale digital twin model of water conservancy is successfully constructed. Through unified encoding and output alignment of input data (step S1), pseudo-label generation and dataset expansion (step S2), knowledge graph and adapter knowledge transfer (step S3), and information joint using attention mechanisms (step S4), the technical effects of cross-model feature alignment, knowledge transfer, and dynamic weight allocation are achieved. The joint features in step S4 can dynamically identify key risk factors in scenarios such as water conservancy monitoring and disaster early warning, enabling efficient prediction and intelligent decision support for complex water conservancy systems. Ultimately, this significantly improves the generalization ability and prediction accuracy of the large-scale digital twin model of water conservancy, providing intelligent scientific basis and efficient tools for water conservancy management and disaster prevention and mitigation.
[0115] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A weak model joint enhancement method for a digital twin water conservancy large model, characterized in that, Specifically, the following steps are included: S1. Unified Encoding and Input Alignment: The original input data from different water conservancy subsystems are standardized and aligned in time and space to obtain an aligned original dataset with consistent structure and no labels. This dataset serves as the input basis for the weak model to generate pseudo-labels in S2. S2. Generate pseudo-labels and expand the dataset using weak models: The aligned original data obtained in S1 is divided into different water conservancy subsystems and then input into the corresponding subsystem pre-trained weak models for prediction to obtain subsystem-level pseudo-label outputs; these pseudo-label data are then merged with the original labeled dataset to construct an expanded training set for use by subsequent large models. The pre-trained weak model of the subsystem refers to the subsystem-level prediction model that is pre-trained on the corresponding water conservancy subsystem and constructed based on expert knowledge rules, statistical modeling or small-scale supervised learning; the original labeled dataset refers to the measured data of historical field monitoring points or the key variables of the subsystem and their spatiotemporal labels labeled after manual inspection and expert evaluation. S3. Constructing a water conservancy knowledge graph or causal structure to guide learning: Constructing a water conservancy knowledge graph or causal graph based on domain knowledge to clarify the logical and physical relationships between variables in each subsystem; assisting the large model to more effectively absorb the pre-trained weak model and expert knowledge of each subsystem in S2; S4. Integrating information from weak and large models for joint modeling: During the training of the large model, an attention mechanism or cue learning structure is introduced to embed the output of the pre-trained weak model in the subsystem of S2 and the knowledge graph structure into the input of the large model, calculate their attention weights for the current task, and realize the fusion of weak model knowledge and large model representation.
2. The weak model joint enhancement method for a digital twin water conservancy large model according to claim 1, characterized in that: The original input data from different water conservancy subsystems in step S1 refers to the original input data collected from the hydrological, hydraulic engineering, water regulation, and water quality water conservancy subsystems. The input data is uniformly encoded using the Z-score normalization method, and the input of the pre-trained weak model of the subsystem is aligned using linear transformation, thereby obtaining a uniformly encoded original dataset with consistent structure and spatiotemporal alignment. The original dataset output from S1 will be used in S2 to generate pseudo-labels and expand the dataset for the weak model. Wherein, the uniform coding is specifically for input data from the first subsystem , using Z-score standardization method for uniform coding: ① In formula 1, is the original data, is the mean value of the first subsystem of the first feature, is the standard deviation, is the normalized data; The input alignment is specifically for the input of the weak model linear transformation to align it with the target dimension of the large model: ② In formula 2, is a weight matrix, is a bias vector, is the aligned output.
3. The weak model joint enhancement method for a digital twin hydraulic system according to claim 1, characterized in that: The input in the S2 step is the aligned original data set output by S1 The aligned original data set obtained by S1 is input into S2 The corresponding subsystem pre-training weak model is input respectively after the division of the subsystem Prediction is performed to generate pseudo-labels at the subsystem level And by setting a confidence threshold High-confidence pseudo-labels are screened out, and these pseudo-label data are combined with the original labeled data to obtain an expanded data set The expanded data set of S2 Will be used in S3 to build a knowledge graph or causal graph structure and in S4 for large models The pseudo-label generation step in step S2 is as follows: align the original dataset obtained in step S1. Using weak models Generate pseudo tags: ③ In formula ③, It is a pseudo-tag. This represents a pre-trained weak model whose output is a probability distribution or a direct label.
4. The method for joint enhancement of weak models in a digital twin hydraulic system according to claim 1, characterized in that: The input to S3 is the extended dataset. Based on domain expert knowledge, the knowledge of weak models is transferred and enhanced by constructing a causal graph structure and injecting adapter parameters, thereby obtaining the causal graph adjacency matrix. and the corrected large model output The causal graph structure and adapter parameters obtained in step S3 will be used in step S4 to introduce an attention mechanism, calculate the attention weights between the output of the weak model and the internal features of the large model, and achieve information fusion between the two. The causal graph structure learning step in step S3 is as follows: Constructing the adjacency matrix of the hydraulic causal graph. The objective function is defined as follows: ④ In formula ④, For the observation data matrix; The Frobenius norm measures the error in data fitting. The L1 regularization term constrains the sparsity of the causal graph. These are acyclic constraints, ensuring that the causal graph is acyclic.
5. The method for joint enhancement of weak models in a digital twin hydraulic system according to claim 1, characterized in that: Step S4 embeds the causal graph structure and adapter parameters obtained in step S3 into the input of the large model along with the output of the pre-trained weak model of the subsystem. Weights are dynamically allocated through an attention mechanism to obtain the joint features. This achievement will enable information complementarity between large and weak models, thereby improving the prediction accuracy and generalization ability of large-scale digital twin models for water conservancy. The attention weight calculation steps for S4 are as follows: Define the internal features of the large model as... The weak model output is Weights are dynamically allocated through an attention mechanism: ⑤ In formula ⑤, A learnable query and key weight matrix; The scaling factor prevents the inner product from becoming too large; the Softmax function ensures that the attention weights are normalized.
6. The method for joint enhancement of weak models in a digital twin hydraulic system according to claim 3, characterized in that, The pseudo-label screening step in step S2 includes setting a confidence threshold. Only retain high-confidence pseudo-labels: The dataset expansion step included in S2 is as follows: [The original labeled dataset is then expanded]. Compared with the filtered pseudo-label dataset merge: .
7. The method for joint enhancement of weak models in a digital twin hydraulic system according to claim 4, characterized in that: The adapter parameter injection step included in S3 is as follows: inserting an adapter layer into the large model. Its parameters are derived from the features of the weak model. Dynamically generated: The output of the large model is corrected as follows: ⑥ In formula ⑥, This is the original output of the large model. These are prediction results from a weak model. The optimization steps included in S3 are: defining a prompt template for the water conservancy field. Fuse the output of the weak model as context: The large model generates causal logic based on prompts: .
8. The method for joint enhancement of weak models in a digital twin hydraulic system according to claim 5, characterized in that: The information joining step in S4 is as follows: passing the output of the weak model through the weight matrix. After mapping, it is weighted and fused with attention weights: ⑦ In formula ⑦, The combined features enable information complementarity between the large and weak models; The attention residual enhancement step in S4 is as follows: To mitigate information loss, residual connections are introduced: 。 9. A weak model joint enhancement system for digital twin hydraulic engineering models, characterized in that, The method for jointly enhancing weak models for a digital twin hydraulic system according to any one of claims 1-8 is used to execute the method for enhancing weak models for a digital twin hydraulic system. The enhancement system includes a weak model module, a data unification module, an enhancement training module, a knowledge management module, and a feedback and evaluation module. The weak model module, the data unification module, and the enhancement training module are used to execute steps S1 and S2, and the feedback and evaluation module and the knowledge management module are used to execute steps S3 and S4. The weak model module consists of multiple water conservancy subsystem modelers. Each water conservancy subsystem modeler is pre-trained on the corresponding water conservancy subsystem and constructed into a corresponding subsystem-level predictive weak model based on expert knowledge rules, statistical modeling, or through small-scale supervised learning. Different water conservancy subsystems include hydrology, hydraulic engineering, water regulation, and water quality. The data unification module includes input / output standardization, data preprocessing, and a label fusion unit, which is used to standardize the raw input data from different hydraulic subsystem modelers, remove noise and outliers, and fuse labels to generate an unlabeled dataset. The enhanced training module includes a pseudo-label generator, a prompt builder, and a fusion optimizer. The pseudo-label generator accepts an unlabeled dataset and predicts and generates pseudo-labels for the unlabeled dataset, which are then used to merge the pseudo-labeled data with the original labeled data to form an expanded dataset. The prompt builder generates prompt information based on a knowledge graph or causal graph structure. The fusion optimizer optimizes the absorption of weak model knowledge during the training of large models; The feedback and evaluation module consists of a simulator, an error monitor, and a reinforcement learning optimizer. The simulator simulates the operation of the water conservancy system, the error monitor monitors the error between the model output and the actual situation, and the reinforcement learning optimizer adjusts the model parameters based on the error feedback. The knowledge management module includes graph construction, causal relationship modeling, and domain knowledge base, which are used to build and maintain knowledge graphs and causal relationship models in the water conservancy field, and store relevant knowledge.