Workflow-based multi-modal water conservancy large model decision support method
Through the multimodal water conservancy large-scale decision-making support method based on workflow, the problems of insufficient multimodal data processing capabilities, contradiction between knowledge base updates and stability and poor model adaptability in the water conservancy monitoring system are solved, real-time acquisition and parallel processing of multimodal data are realized, and the emergency response speed and decision-making accuracy of water conservancy are improved.
Patent Information
- Application Number
- CN202510828573.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing water conservancy monitoring systems lack multimodal data processing capabilities, the knowledge base update is contradictory to stability, poor model adaptability and scalability, and fragmented decision-making processes, making it difficult to deal with dynamic changes and complex scenarios of multi-source data under extreme climate conditions.
The multimodal water conservancy big model decision-making support method is adopted based on workflow. Through a modular parallel processing architecture, dual-channel knowledge fusion, adversarial training and LoRA lightweight module, real-time acquisition and parallel processing of multimodal data are realized, and an efficient and real-time intelligent service platform is built.
The synchronous processing of multimodal data is realized, the water conservancy emergency response time is improved, the stability of offline knowledge base and the real-time update of networked data is ensured, the multimodal error rate is reduced, and the system's robustness and decision-making accuracy are improved.
Smart Images

Figure CN120353853A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water conservancy data processing, and particularly relates to a decision support method, system, device and storage medium for a multi-modal water conservancy large model based on a workflow. Background Art
[0002] Water conservancy projects play a crucial role in the safety monitoring and emergency management of key infrastructure such as reservoirs and dams. Existing water conservancy monitoring systems mainly rely on traditional single data collection and static data analysis methods, mostly using offline knowledge bases and single-modal data processing, and it is difficult to cope with the challenges brought by data diversity, real-time nature and complex working conditions in practical applications. Especially under extreme climate conditions (such as heavy rainfall, rainstorms), the dynamic changes of multi-source data and complex scenarios faced by water conservancy facilities urgently require more efficient and accurate monitoring and decision-making means.
[0003] The following technical bottlenecks exist in the prior art: (1) Insufficient multi-modal data processing ability: Existing systems mostly focus on single-type data (such as water level sensor values), lacking the ability to jointly analyze multi-modal data such as text (such as emergency reports), images (such as dam inspection photos), and voice (such as on-site personnel instructions), resulting in low information utilization rate; (2) Contradiction between knowledge base update and stability: Local knowledge bases (such as flood control plans, historical cases) are difficult to cope with real-time flood conditions due to lagging updates, while relying entirely on networked data (such as meteorological warnings) also faces the risk of network fluctuations, and the collaborative mechanism between the two is lacking; (3) Poor model adaptability and scalability: General artificial intelligence models lack adaptation to water conservancy domain terms and scenarios, and it is difficult for models (such as image recognition and text generation) to cooperate in parallel, restricting the processing efficiency of complex tasks; (4) Fragmented decision-making process: The link from data input, analysis to decision output depends on manual connection, lacking modular workflow support, resulting in slow response speed and poor traceability.
[0004] Therefore, this application specifically proposes a decision support method for a multi-modal water conservancy large model based on a workflow to solve the above technical problems, which can realize the real-time collection and parallel processing of multi-modal data (text, image, voice, etc.), and through the dual-channel fusion of an offline knowledge base and networked real-time data, build an efficient, real-time and intelligent "question - analysis - decision" one-stop service platform. Summary of the Invention
[0005] The main objective of the present invention is to provide a workflow-based decision support method for multi-modal water conservancy large models, which supports multi-modal data collection and processing (including text, images, voice, etc.), and realizes the dual-channel integration of local knowledge bases and real-time online knowledge, thereby constructing a one-stop "question - analysis - decision" intelligent service platform that can ensure the stability of offline data and the real-time update of network information. At the same time, by integrating embedding models, language models, re-ranking models, and voice models, multi-task parallel processing and accurate data analysis are achieved, providing efficient and accurate decision support for fields such as water conservancy project management, flood warning, and water resource scheduling, so as to solve the technical problems proposed in the background technology.
[0006] The present invention adopts the following technical solutions to solve the above technical problems: A workflow-based decision support method for multi-modal water conservancy large models, comprising: S1. Preprocess multi-modal input data based on a local model, and convert the data formats of files, images, and voice into JSON format; S2. Construct a modular workflow that supports parallel processing by specifying a DSL framework to ensure that different types of data after format conversion can be synchronously collected and processed; S3. Construct a feature model suitable for the water conservancy field through the synchronously collected and processed multi-modal data, build data dependency relationships based on the DSL framework, and perform domain adaptation and fine-tuning on the feature model. Embed the water conservancy mechanism into the model fine-tuning process after parameterizing the model through a three-layer architecture of implicit rules - explicit constraints - rule activation; S4. Based on the feature model after domain adaptation and fine-tuning, construct a specified workflow for dynamically scheduling multi-modal task resources to achieve a dynamic processing flow from data input, analysis to decision output; S5. Use a dynamic fusion algorithm to combine the local knowledge base and online data to achieve dual-channel knowledge fusion for parallel execution of the multi-modal tasks in S4 to ensure the stability of the offline knowledge base and the real-time update of network data.
[0007] Preferably, the specific operation process of the preprocessing in step S1 includes: S11. Based on the obtained multi-modal input data, perform preliminary classification according to the specified type criteria to ensure that they enter the corresponding preprocessing processes by type respectively; S12. For different types of multi-modal data, use specified models for preprocessing respectively: For text data, extract semantic features through an embedding model, and then use a language model for noise removal, format unification, word segmentation, and part-of-speech tagging; For image data, use an embedding model for image feature extraction, perform denoising, resolution adjustment, and feature point extraction; For voice data, a voice model is used for voice recognition and processing, including denoising, converting voice to text, and extracting voice features. Preferably, the specific construction process of the modular workflow in step S2 includes: S21. In a specified DSL framework, different data processing modules are defined, and a parallel processing architecture is adopted to ensure that specified data types perform data processing simultaneously in different processing channels; S22. Based on the specified DSL framework, by defining task dependencies and execution orders, the optimization and coordination of multi-task parallel execution are realized.
[0008] Preferably, the specific operation process of step S3 includes: S31. Use a feature extraction model adapted to the water conservancy field to extract specified features from different data types of multi-modal data in the water conservancy field to obtain a feature model; Based on the existing knowledge base in the water conservancy field, pre-training data is constructed, and the MLM task is used to perform domain adaptation operations on the feature model; Perform LoRA lightweight adjustment operations on the feature model, freeze the parameters of the low-rank matrix of the pre-trained data model, only adjust the newly added LoRA adaptation layer, and use the annotation data set in the water conservancy field for supervised adjustment, where the specified control indicators of water conservancy projects are incorporated into the model adjustment process as explicit constraint conditions to build a physical consistency verification mechanism for checking whether the output results violate hydrological physical laws during each model update, and a penalty term mechanism is used to correct predictions that violate physical laws to ensure the physical rationality of the model output.
[0009] Preferably, after step S3 is executed, model verification and dynamic optimization operations are performed, which specifically include: Construct a test set in the water conservancy field for domain accuracy evaluation, and use F1-score and BLEU metrics to verify the model performance; Construct a set of cross-modal consistency samples for inspection to ensure the logical consistency of outputs in different modalities; Deploy the model by integrating the A / B test module, and adopt a gray release strategy. Trigger incremental training of the model through user feedback data to perform online dynamic optimization and closed-loop improvement.
[0010] Preferably, during the inspection of the cross-modal consistency samples, a cross-modal alignment algorithm is used to ensure the semantic consistency of different modal data, and the usage method of the cross-modal alignment algorithm includes: L1. Perform knowledge graph embedding pre-training in the form of TransE on all triples Minimize the knowledge graph embedding pre-training loss The calculation formula is as follows:
[0011] Among them, is the margin hyperparameter of the model, which is used to ensure the effective distinction between positive and negative samples in the pre-training of knowledge graph embedding; after training, all and are fixed; , , are the vectors of the head entity, relation, and tail entity in the graph embedding space respectively, and represent the starting entity of the relation, the semantic relation between entities, and the target entity of the relation in the knowledge graph respectively. is the L2 norm, is the set of all triples in the knowledge graph; L2. For each positive sample , several negative entities are sampled from to form . Cross-modal feature mapping is performed, and the alignment loss is as follows:
[0012] Among them, is the set of negative samples or a space containing negative samples. Specifically, it is the source from which negative entities are sampled and is used to optimize by comparing with positive samples during the training process. represents the mapping network, which is used to map the visual feature and the text feature to the embedding space with the same dimension as the entity. is the hinge loss of the cross-modal alignment loss, which includes the visual alignment term and the text alignment term. represents the set of negative entities sampled for the positive entity . represents the margin hyperparameter of the alignment loss. represents the similarity function between A and B; L3. Construct the training objective, as follows:
[0013] Among them, is used to represent the set of weights and biases of the visual and text mapping networks. is the weight of the visual mapping network. is the bias of the visual mapping network. is the weight of the text mapping network. is the bias of the text mapping network. is the regularization term, and its value is the sum of the squares of all trainable parameters. is the L2-regularized weight.
[0014] Preferably, an adversarial training mechanism is further introduced in the step S32. A domain feature adversarial network is introduced, and a domain classifier is used to distinguish general data from water conservancy data, so as to improve the sensitivity of the feature model to professional terms, enable the feature model to implicitly learn the physical constraint relationships of water conservancy projects, and set a specified pre-training task to enable the feature model to predict results that conform to hydrological physical laws under given boundary conditions, so as to realize the implicit encoding and storage of mechanism knowledge in the model weights. The adversarial training mechanism specifically includes the following operation steps of the adversarial training improvement method: P1. Divide the water conservancy data set into a source domain and a target domain, and eliminate the data distribution difference through standardization processing; P2. Generate a feature extraction module of the adversarial network, combine the deep residual shrinkage network to compress the redundant feature dimensions, and introduce the long short-term memory network to capture the temporal dependence relationship to form a model architecture, and add a domain discriminator as the core of adversarial training in the model architecture, and force the feature encoder to generate domain-invariant features through the gradient reversal layer; P3. Based on the model architecture constructed in the step P2, perform adversarial training operations, including: Phase 1: Pre-train the feature encoder on the source domain data to initially extract the key features related to water conservancy parameters; Phase 2: Fix the feature encoder, train the domain discriminator to distinguish the source domain and target domain features, and optimize the feature encoder to confuse the discriminator by minimizing the domain discriminator loss function to achieve domain feature alignment; Phase 3: On the basis of adversarial training, jointly fine-tune the model using the source domain labeled data and a small amount of target domain labeled data, and adopt a weighted cross-entropy loss function to balance the problem of class imbalance; P4. Embed a dynamic threshold judgment module in the output layer of the model, which is used to adaptively adjust the classification threshold according to the statistical features of the historical data of the water conservancy project.
[0015] Preferably, the specific operation process of the step S4 includes: S41. Based on the fine-tuned feature model, construct a set of "question - analysis - decision" workflows, define the task transfer relationships between each sub-module, ensure that each module can start corresponding analysis and processing tasks according to the user's input, and smoothly advance to the decision output according to the predetermined logic; S42. Adopt a preset task priority scheduling strategy for dynamically optimizing the task execution efficiency in the workflow through resource scheduling. The scheduling judgment logic of the task priority scheduling strategy includes: Dynamically adjust the benchmark priority of various tasks according to the water conservancy operation status: During the flood season, the real-time sensor data task is automatically elevated to the highest priority; For flood control instructions and routine monitoring tasks, during the non-flood season, the setting is restored to the normal priority; The scheduler adopts preemptive scheduling. When a new task arrives, the priority of the new task is compared with that of the task being executed. If the former is higher, the execution will be immediately interrupted and switched; otherwise, the execution will continue; When the priorities are the same, the CPU is allocated according to the first-come-first-served or round-robin scheduling.
[0016] Preferably, the specific operation process of performing the dynamic fusion algorithm to combine data in the step S5 includes: S51. Obtain the offline knowledge base and real-time data through the acquisition interfaces of the local knowledge base and networked data respectively, and perform independent preprocessing operations on them; S52. Combine the local knowledge base and networked data. The offline knowledge base provides stable information, and the networked data is updated in real time and provides the latest decision support. The dynamic fusion algorithm is used to balance the stability of the local knowledge base and the real-time nature of the networked data.
[0017] Preferably, the usage method of the dynamic fusion algorithm includes; (1) Calculate the timeliness trust value, and the calculation formula is:
[0018]
[0019] Among them, and represent the timeliness trust values of the local knowledge base and networked data respectively. When the time since the local knowledge base was updated is longer, then is smaller. When the time since the networked data was updated is shorter, then is larger; represents the static credibility benchmark of the local knowledge base, represents the static credibility benchmark of the network search results, represents the time interval since the last full synchronization of the local knowledge base, represents the average delay of network data update crawling or the time since the article was published, represents the normalization factor for controlling the credibility decay of the local knowledge base, represents the normalization factor for controlling the credibility decay of network data; (2) Perform weight normalization operation, and there is:
[0020]
[0021] Among them, is the normalized weight of the local knowledge base, is the normalized weight of the network data, represents the network connectivity indication, indicates online availability, indicates switching from offline to pure local mode; (3) Perform retrieval result score fusion. For each candidate , there is:
[0022] Among them, represents the score fusion result for each candidate , and respectively represent the local knowledge base retrieval score and the network data retrieval score for each candidate ; (4) Finally output the fusion result list in descending order.
[0023] On the other hand, the present invention also discloses a workflow-based multi-modal water conservancy large model decision support system for constructing a corresponding model to execute the above-mentioned workflow-based multi-modal water conservancy large model decision support method, which respectively includes: A data input layer for data collection input and preprocessing; A model processing layer for constructing a feature model based on the preprocessed data; A water conservancy field adaptation layer for performing field adaptation and fine-tuning operations on the feature model; An application service layer for performing corresponding application service operations based on the model after the fine-tuning operation.
[0024] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the above method.
[0025] On yet another hand, the present invention also discloses a computer device including a memory and a processor, the memory storing a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the above method.
[0026] As can be seen from the above technical solutions, the present invention provides a workflow-based multi-modal water conservancy large model decision support method. Compared with the prior art, the present invention has the following advantages: 1. The present invention can synchronously process text, image, and voice data by setting a modular parallel processing architecture in the workflow, realizing deep feature extraction and multi-task collaborative execution operations to shorten the water conservancy emergency response time.
[0027] 2. The present invention constructs a dual-channel knowledge fusion mechanism, combines a large-scale local knowledge base with a real-time networked knowledge base, and effectively comprehensively processes information using a data fusion algorithm, enabling the stability of the offline knowledge base and the real-time update of network knowledge. At the same time, by setting a spatio-temporal exponential decay fusion algorithm in the knowledge management module, it can dynamically balance the stability of the local knowledge base and the real-time nature of networked data, achieving a smooth switching effect between offline / online modes, thereby shortening the flood season decision-making response time.
[0028] 3. The present invention can enhance the ability to recognize water conservancy professional terms by setting adversarial training and the LoRA lightweight module in the model fine-tuning stage, thereby further improving the domain adaptation accuracy and reducing the prediction error of hydrological parameters.
[0029] 4. The present invention can verify the logical consistency between image recognition and text reports by setting a cross-modal alignment algorithm based on a knowledge graph in the output layer, thereby reducing the multi-modal error rate and avoiding major misjudgments.
[0030] 5. The present invention can preferentially process real-time sensor data during heavy rainfall by setting a flood season dynamic priority strategy in the task scheduler to improve system throughput and facilitate automatically switching to the local library to ensure service continuity when the network is interrupted.
[0031] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Of course, any product implementing the present invention does not necessarily need to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The schematic diagrams in the specification drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 is a schematic diagram of the overall framework of the decision-making support method of the present invention; Figure 2 is a schematic diagram of the model data processing flow within the model processing layer of the present invention; Figure 3 is a schematic diagram of the model data adaptation iteration flow within the water conservancy domain adaptation layer of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0034] In the embodiments, for details, see Figures 1 to 3 .
[0035] A multimodal water conservancy large model decision support system and method based on a workflow proposed in the embodiments of the present invention.
[0036] Among them, as Figure 1 shown, the multimodal water conservancy large model decision support system is used to construct a corresponding model to execute the multimodal water conservancy large model decision support method. At this time, the multimodal water conservancy large model decision support system respectively includes: The data input layer is used for data collection input and preprocessing; The model processing layer is used to construct a feature model based on the preprocessed data; The water conservancy field adaptation layer is used to perform field adaptation and fine-tuning operations on the feature model; The application service layer is used to execute corresponding application service operations based on the model after the fine-tuning operation.
[0037] At this time, the process of the method executed by the system is as Figure 2 shown, including: First, realize the real-time input of text, image and voice data through the multimodal data collection interface, and use a unified preprocessing method to ensure the stability and accuracy of the data; Then, use the local embedding model bge-m3, language model, voice model and re-rank model deployed on the server to perform deep feature extraction and semantic parsing on the preprocessed data, and support multi-task parallel processing such as image recognition and text generation; Build a dual-channel knowledge fusion mechanism based on a large-scale local knowledge base and a real-time online knowledge base, and use a data fusion algorithm to comprehensively process information to realize the organic connection between the stability of the offline knowledge base and the real-time update of network knowledge; Further, for the needs of the water conservancy field, perform professional annotation and preprocessing on the collected data, and perform field fine-tuning on the model data to improve the system's accurate recognition and processing ability of hydrological information; Finally, automatically schedule each processing module by generating a workflow, develop exclusive API interfaces and user interfaces, and realize a one-stop intelligent service from "questioning" to "analysis" to "decision-making".
[0038] At this time, the local model data deployed in the server, combined with the full-link automation of "questioning - analysis - decision-making" in the water conservancy scenario, can achieve real-time collection and parallel processing of multi-modal data (text, images, voice, etc.). Through the dual-channel knowledge fusion of offline knowledge bases and online real-time data and the multi-modal task parallel technology, an efficient, real-time and intelligent "questioning - analysis - decision-making" one-stop service platform is constructed. This platform not only breaks through the limitations of traditional systems in data processing and real-time updates, but also realizes the intelligent processing and decision support of water conservancy information under complex working conditions through multi-task parallelism and modular design, thus significantly improving the accuracy and response speed of monitoring and early warning, and further significantly enhancing the offline environment stability and real-time response ability, providing an integrated solution for water conservancy project monitoring, emergency command and knowledge service.
[0039] Therefore, this system can perform deep feature extraction and semantic parsing on data, support parallel processing of multi-tasks such as image recognition and text generation, and greatly improve the system processing efficiency and accuracy.
[0040] Therefore, further referring to Figure 1 , the specific implementation steps of this multi-modal water conservancy large model decision support method include: S1. Preprocess the multi-modal input data (including text, images, and voice) in the water conservancy field through the local model, and convert various formats of files, images, and voice into JSON format using tags and specific Python code.
[0041] At this time, the specific operation process of preprocessing includes: S11. Obtain the multi-modal input data in the water conservancy field in the system, including text, images, and voice, etc. Based on the obtained multi-modal input data, conduct preliminary classification according to the specified type standard to ensure that they enter the corresponding preprocessing process by type respectively; S12. For different types of multi-modal data, use the specified model for preprocessing respectively: For text data, first perform semantic feature extraction through the embedding model, and then use the language model for noise removal, format unification, word segmentation, and part-of-speech tagging to ensure the quality and processability of the text data, and further ensure the stability and accuracy of its semantics; For image data, use the embedding model for image feature extraction, perform denoising, resolution adjustment, and feature point extraction to ensure the stability and accuracy of the image data; For voice data, use the voice model for voice recognition and processing, including denoising, converting voice to text, and voice feature extraction to ensure the quality of the voice data for subsequent processing and analysis.
[0042] S2. Through the Dify DSL framework, build a modular workflow that supports parallel processing to ensure that different types of data such as text, images, and voices after format conversion can be synchronously collected and processed.
[0043] The specific construction process of the modular workflow includes: S21. In the Dify DSL framework, perform modular design of the workflow, define different data processing modules, ensure that each data type (such as text, image, voice) has an independent processing channel, and adopt a parallel processing architecture to ensure that the specified data types perform data processing simultaneously within different processing channels. The data processing modules respectively include data acquisition, preprocessing, feature extractor, feature understanding, template conversion, etc.; S22. Based on the Dify DSL framework, adopt a parallel processing architecture to enable data types such as text, images, and voices to be executed simultaneously on multiple processing units. By defining task dependencies and execution sequences, optimize and coordinate the parallel execution of multiple tasks, ensure the coordination and synchronization between modules, avoid bottlenecks in data processing, and improve overall efficiency; At this time, through the dynamic resource scheduling and cross-modal synchronization mechanism of Dify DSL, optimize and coordinate the parallel execution of multiple tasks: configure dynamic resource allocation strategies (GPU gives priority to processing images, CPU processes text) in the modular workflow, and isolate task resources by combining lightweight containerization technology; ensure efficient use of resources and enhance the robustness of the system.
[0044] S3. Build a feature model adapted to the water conservancy field through the synchronously collected and processed multimodal data, build data dependency relationships based on the DSL framework, and perform domain adaptation and fine-tuning on the feature model to improve the system's accurate recognition and processing capabilities in aspects such as hydrological data and engineering parameters, so as to meet the high-precision requirements in the water conservancy decision-making process. Among them, the water conservancy mechanism is parameterized into the model fine-tuning process through a three-layer architecture of implicit rules - explicit constraints - rule activation. Specifically, by inputting early warning prediction rules and historical hydrological data into the water conservancy feature model, achieve the deep integration of mechanism knowledge and data-driven models, so that the water conservancy feature model can understand and apply hydrological physical laws for reasoning.
[0045] Reference Figure 3 , the specific operation process includes: S31. Multi-modal data analysis and feature extraction: First, conduct in-depth analysis and feature extraction on multi-modal data in the water conservancy field (including text, images, speech, etc.). Use a feature extraction model adapted to the water conservancy field to extract useful hydrological data, engineering parameters and other features from different data types, ensuring the domain relevance and high-precision expression of the data. These features will provide a basis for subsequent domain adaptation and model fine-tuning. Build a parametric encoder for the water conservancy mechanism model to convert hydrological physical laws (such as the Saint-Venant equation, Manning's formula, flood routing equation, etc.) into vector representation forms that can be understood by the water conservancy feature model. Through an adversarial training mechanism, enable the water conservancy feature model to implicitly learn the physical constraint relationships of water conservancy projects. Design a special pre-training task to let the model predict results that conform to hydrological physical laws under given boundary conditions, realizing the implicit encoding and storage of mechanism knowledge in the model weights; S32. Domain adaptation of model data: For the needs of the water conservancy field, first conduct domain adaptation on the model data by combining the knowledge and data of the water conservancy field, enabling it to better understand and process water conservancy-specific data, such as hydrometeorological data and engineering design parameters. To this end, construct pre-training data based on the water conservancy domain knowledge base and use the MLM (Masked Language Model) task to perform domain adaptation on the embedding model. In addition, introduce an adversarial training mechanism to distinguish general data from water conservancy data through a domain classifier, enhancing the model's sensitivity to professional terms; S33. Perform LoRA lightweight fine-tuning operations on the feature model, freeze the parameters of the low-rank matrix of the pre-trained data model, only fine-tune the newly added LoRA adaptation layer, and use the water conservancy domain annotation dataset for supervised adjustment, further reducing the video memory occupancy rate and improving the inference speed, thereby ensuring the accurate processing ability of the feature model in the water conservancy field, while optimizing resource consumption and processing efficiency; Among them, key control indicators of water conservancy projects (such as flood control limit water level, minimum ecological flow, structural safety threshold, etc.) are incorporated into the model optimization and fine-tuning process as explicit constraint conditions. By constructing a physical consistency verification mechanism, check whether the output results violate hydrological physical laws during each model update, and use a penalty term mechanism to correct predictions that violate physical laws, ensuring the physical rationality of the model output. Automatically identify applicable water conservancy project rules and mechanism models based on the characteristics of the input multi-modal data. By designing a dynamic rule weight allocation algorithm, dynamically call relevant physical laws and warning rules according to the current hydrological conditions, engineering status and task requirements during the model inference process, realizing an intelligent rule application process of "data feature → rule matching → weight allocation → inference activation", enabling the model to adaptively apply the most relevant water conservancy professional knowledge in different engineering scenarios; Specifically, through an innovative three - layer architecture of implicit rules - explicit constraints - rule activation, the parameters of the hydraulic mechanism model are parameterized and embedded in the fine - tuning process of the large model to achieve the deep integration of mechanism knowledge and data - driven models. The specific operation steps of this three - layer architecture are as follows: T1. Implicit rule encoding layer: Parameterize the hydraulic mechanism model (such as the Saint - Venant equation, hydrological model, etc.) into vector representations, and enable the large model to implicitly learn the physical constraint relationships of hydraulic engineering through specially designed pre - training tasks. These pre - training tasks include: Predict hydraulic results that conform to physical laws based on given hydrological boundary conditions; Establish a mapping relationship between numerical simulations and measured data to enable the model to learn the internal laws of physical models; Strengthen the implicit encoding of physical laws by the model through adversarial training, enabling it to form an ability similar to "physical intuition"; T2. Explicit constraint layer: Design a multi - objective constraint loss function based on the early warning standards and safety specifications of hydraulic engineering, and incorporate key control indicators (such as flood control limit water level, minimum ecological flow, structural safety threshold, etc.) as explicit constraint conditions into the model optimization process. This layer specifically includes: Construct a physical consistency verification mechanism to check whether the output results violate hydrological physical laws during each model update; Design a penalty term mechanism to automatically correct prediction results that violate physical laws; Introduce the statistical characteristics of historical hydrological data as prior knowledge to guide the model to generate prediction results within a physically reasonable range; T3. Rule activation mechanism layer: Establish an intelligent rule activation system for hydraulic engineering based on the attention mechanism to achieve an intelligent rule application process of "data feature → rule matching → weight assignment → reasoning activation". This layer specifically includes: Automatically identify applicable hydraulic engineering rules and mechanism models according to the input multi - modal data features; Design a dynamic rule weight assignment algorithm to dynamically call relevant rules according to the current hydrological conditions, engineering status, and task requirements; Construct an association mapping between the early warning rule base and the mechanism model to enable the model to call the corresponding early warning thresholds when predicting hydrological changes; Implement a scenario - adaptive mechanism to intelligently call the most relevant professional rules in different engineering scenarios (such as flood control, water supply, ecological regulation, etc.); T4. Co - training of the three - layer architecture: Conduct end - to - end training on the three - layer architecture, and make each layer coordinate with each other through gradient backpropagation. Specifically, it includes: Use distillation technology to use the prediction results of the mechanism model as soft labels to guide the learning of the large model; Design a multi-task learning framework to optimize both prediction accuracy and physical rationality simultaneously; Adopt a cyclic verification mechanism to form complementary verification between the model prediction results and the calculation results of the mechanism model; Introduce a knowledge graph in the field of water conservancy projects as an auxiliary information source to enhance the model's understanding of the conceptual relationships in water conservancy projects.
[0046] Through the above three-layer architecture, this method realizes the deep integration of water conservancy mechanism knowledge and deep learning models, enabling the large model to understand and apply hydrological physical laws for reasoning, providing more accurate and reliable technical support for water conservancy project risk prediction and decision-making.
[0047] S4. Execute model verification and dynamic optimization operations, specifically including: After the feature model is adapted and fine-tuned for the domain, conduct domain accuracy evaluation by constructing a test set in the water conservancy field (including text, image, and speech samples), and verify the model performance using indicators such as F1-score and BLEU to ensure that the term recognition accuracy rate reaches over 97%; Construct a set of cross-modal consistency samples for inspection to ensure the logical consistency of outputs in different modalities. For example, check whether the "crack width" identified in the image matches the "seepage pressure level" described in the text, thereby improving the credibility of multi-modal outputs; Execute the deployed model through the integrated A / B test module, and adopt a gray release strategy. Automatically trigger the incremental training of the model through user feedback data (such as reports corrected by experts) to perform online dynamic optimization and closed-loop improvement.
[0048] In summary, through these verification and optimization processes, the model will be continuously fine-tuned and optimized to ensure the data processing accuracy and decision support reliability in the water conservancy field, ultimately meeting the requirements for high-precision recognition and processing in the water conservancy decision-making process.
[0049] Furthermore, during the process of inspecting the cross-modal consistency samples, adopt a cross-modal alignment algorithm to ensure the semantic consistency of data in different modalities, so as to reduce the error rate after consistency verification and provide experimental data. The usage method of the cross-modal alignment algorithm includes: L1. Execute knowledge graph embedding pre-training in the form of TransE for all triples Minimize the knowledge graph embedding pre-training loss The calculation formula is:
[0050] Among them, is the margin hyperparameter of the model, used to ensure the effective distinction between positive and negative samples in knowledge graph embedding pre-training; after training, fix all and ; , , are the vectors of the head entity, relation, and tail entity in the graph embedding space, respectively, and represent the starting entity of the relation, the semantic relation between entities, and the target entity of the relation in the knowledge graph, is the L2 norm, is the set of all triples in the knowledge graph; L2. For each positive sample , sample several negative entities from to form , perform cross-modal feature mapping, and align the loss, we have:
[0051] where, is the set of negative samples or a space containing negative samples. Specifically, it is the source from which negative entities are sampled for optimizing by comparing with positive samples during the training process, represents the mapping network, which is used to map the visual feature and the text feature to the embedding space with the same dimension as the entity, is the hinge loss of the cross-modal alignment loss, which includes the visual alignment term and the text alignment term, represents the set of negative entities sampled for the positive entity , represents the margin hyperparameter of the alignment loss, represents the similarity function between A and B; L3. Build the training objective, we have:
[0052] where, , which is used to represent the set of weights and biases of the visual and text mapping networks, is the weight of the visual mapping network, is the bias of the visual mapping network, is the weight of the text mapping network, is the bias of the text mapping network, is the regularization term, and its value is the sum of the squares of all trainable parameters, is the L2 regularization weight.
[0053] This cross-modal alignment algorithm first pre-trains the entities and relations in the water conservancy engineering knowledge graph using TransE embedding to obtain a unified vector space representation; then uses two linear mapping networks—one to map image features to the embedding space, and the other to map text features into the same space—to enable both visual and text representations to be directly compared with the knowledge graph entity vectors. By constructing a hinge-like alignment loss for positive and negative samples, the network parameters can be fine-tuned while keeping the "crack width" image feature and its correct "crack width segment" entity vector, and the text description and its "permeability level" entity vector close to each other. At the same time, the wrong entities are pulled away. During reasoning, the most similar entities of the image and text features in the knowledge graph entity space are selected respectively, and the two are checked for matching based on the predefined "semantic consistency" relationship in the knowledge graph, thereby completing the cross-modal consistency check. This method can significantly improve the algorithm alignment accuracy and reduce the consistency error rate in multimodal water conservancy detection scenarios.
[0054] In summary, adversarial training, LoRA fine-tuning and dynamic incremental optimization are combined to solve the problem of professional terminology recognition in the field of water conservancy using general models.
[0055] In addition, it should be noted that in step S32, an adversarial training mechanism is introduced, a domain feature adversarial network is introduced, and the performance differences of the general model are compared. The domain classifier is used to distinguish general data from water conservancy data to improve the sensitivity of the feature model to professional terms. The adversarial training mechanism specifically includes the following adversarial training improvement method operation steps: P1. Divide the water conservancy dataset into source domain (labeled data, such as typical watershed monitoring data) and target domain (unlabeled or small amount of labeled data, such as real-time data of newly built facilities), and eliminate data distribution differences through standardization; P2. Construct a feature extraction module based on an improved conditional generative adversarial network, combine a deep residual shrinkage network to compress redundant feature dimensions, and introduce a long short-term memory network to capture temporal dependencies to form a model architecture. In the model architecture, add a domain discriminator as the core of adversarial training, and force the feature encoder to generate domain-invariant features through a gradient reversal layer. P3. Based on the model architecture built in step P2, perform adversarial training operations, including: Phase 1: Pre-train feature encoders on source domain data to preliminarily extract key features related to hydraulic parameters (such as flow, pressure, and vibration frequency); Stage 2: Fix the feature encoder, train the domain discriminator to distinguish the source domain and target domain features, and optimize the feature encoder to confuse the discriminator by minimizing the domain discriminator loss function to achieve domain feature alignment; Phase 3: Based on adversarial training, the model is jointly fine-tuned using source domain labeled data and a small amount of target domain labeled data, and a weighted cross-entropy loss function is adopted to balance the class imbalance problem; P4. Embed a dynamic threshold judgment module in the model output layer, and adaptively adjust the classification threshold according to the statistical characteristics of historical data of water conservancy projects (such as the flow fluctuation range under extreme weather) to improve the robustness of the model.
[0056] In a specific embodiment, through actual experimental measurements, the proposed adversarial fine-tuning model DFAN in this application is significantly better than the general model and the ordinary fine-tuning model on the target domain water conservancy project data: in terms of the RMSE (root mean square error) index, DFAN is reduced by 13.4% compared with the ordinary fine-tuning model (1.032→0.894), and the MAE (mean absolute error) synchronously decreases by 14.9% (0.824→0.701), and the coefficient of determination R² is increased to 0.837. It is worth noting that by introducing the domain feature adversarial network, DFAN effectively eliminates the data distribution difference between the general domain and the target domain in the adversarial training stage. This mechanism not only improves the prediction accuracy (R² approaches the ideal value of 1), but also significantly enhances the robustness of the model - the extraction of cross-domain invariant features during the adversarial training process enables the model to maintain a stable output when facing data fluctuations or local outliers in the target domain (such as the sudden sensor noise interference scenario in water conservancy projects). In contrast, the ordinary fine-tuning model has limitations in both accuracy and stability due to the failure to solve the domain shift problem.
[0057] S4. Based on the domain adaptation and fine-tuned feature model, construct a "question - analysis - decision" workflow for automatically and dynamically scheduling the multimodal task resources of each sub-module to achieve a fully automated dynamic processing flow from data input, analysis to decision output.
[0058] The specific operation process includes: S41. Based on the fine-tuned feature model, construct a set of "question - analysis - decision" workflows to clarify each link of data input, analysis process and decision output. Through the automated scheduling system, define the task transfer relationship between each sub-module, including data collection, model inference, result analysis and other links. Ensure that each module can automatically start the corresponding analysis and processing tasks according to the user's input and smoothly advance to the decision output according to the predetermined logic; S42. The full automation process is realized and optimized through the workflow automation system: establish a full automation process from data input to decision output to ensure that the system can automatically process various requirements in the water conservancy field according to real-time data and preset rules, ensure that the system can execute various tasks efficiently and stably, and provide intelligent decision support.
[0059] At this time, a modular workflow is designed for water conservancy scenarios (such as full-link automation of "questioning-analysis-decision-making"). A preset task priority scheduling strategy for water conservancy is adopted (such as giving priority to real-time sensor data during flood season). Its improvement on system throughput is verified. At the same time, the unique fault-tolerant mechanism in the workflow is described (such as automatically switching to the local knowledge base when the network is interrupted). The stability difference of the traditional system is compared to optimize the task execution efficiency in the workflow for dynamic resource scheduling, including: A. Network status monitoring: (a) The system periodically (e.g., every 5 seconds) initiates a lightweight heartbeat check on the cloud retrieval interface or gateway in the background; (b) If the heartbeat detection fails three times in a row or the request times out, the network is considered unavailable.
[0060] B. Automatic switching: (a) Trigger condition: triggered immediately after the network is detected to be unavailable; (b) Switching logic: All tasks in the "questioning" phase that originally required calling external retrieval or online model prediction are changed to calling the local offline knowledge base and local model; at the same time, user requests and system feedback are written to the local persistent queue (to ensure that they are not lost); (c) Rollback recovery: After the network is restored, the local queue is immediately cleared, and all requests accumulated during the offline period are resent to the cloud service in the original order in an idempotent manner, and the latest index or model update is synchronized; after the local queue is cleared, the system returns to online mode.
[0061] A. Subtask retry and backoff: (a) For critical sub-processes (such as real-time sensor pull and image recognition), a maximum of three retries are performed in the event of a single failure, with an exponential backoff interval (2s, 4s, 8s) between each retry; (b) If all retries fail, the system immediately switches to local mode and does not continue to block the main process.
[0062] Furthermore, in the scheduling judgment logic of the task priority scheduling strategy, the baseline priority of each task is dynamically adjusted according to the water conservancy operation status (such as flood season and non-flood season): During flood season, real-time sensor data tasks are automatically elevated to the highest priority; For flood control instructions and routine monitoring tasks, during the non-flood season, they will be restored to the normal priority settings; The scheduler uses preemptive scheduling. When a new task arrives, the priority of the new task is compared with the currently executing task. If the former is higher, the new task is immediately interrupted and switched, otherwise it continues to execute. For the same priority, the CPU is allocated on a first-come-first-served basis or in a round-robin manner.
[0063] At this time, through simulation experiments, compared with the static priority and non-priority strategies, the system throughput of this scheme is significantly improved in the flood season scenario, verifying its effectiveness in ensuring real-time data processing capabilities at critical moments.
[0064] S5. Use the dynamic fusion algorithm to combine the local knowledge base and networked data to achieve dual-channel knowledge fusion for parallel execution of the multimodal tasks in S4, so as to ensure the stability of the offline knowledge base and the real-time update of network data. At this time, through dual-channel knowledge fusion and multimodal task parallel technology, the stability of the offline environment and the real-time response ability are significantly improved, providing an integrated solution for water conservancy project monitoring, emergency command and knowledge service, and realizing accurate and efficient intelligent decision-making support.
[0065] Among them, the specific operation process of executing the dynamic fusion algorithm in combination with data includes: S51. Obtain the offline knowledge base and real-time data respectively through the acquisition interfaces of the local knowledge base and networked data, and perform independent preprocessing operations respectively to ensure the uniformity of data format and the guarantee of quality. For networked real-time data, format conversion, denoising and verification need to be carried out in a timely manner to ensure its timeliness and accuracy, while the local knowledge base focuses on the stability and consistency of offline data to ensure that high-quality reference information can be provided when there is no network connection; S52. Use the fusion strategy to effectively combine the local knowledge base and networked real-time data, dynamically integrate the advantages of the two types of data, and ensure that in different application scenarios, the offline knowledge base provides stable information, while the networked data can be updated in real time and provide the latest decision-making support. This fusion process ensures that the system can balance stability and real-time performance in various working environments and provide accurate and comprehensive data support.
[0066] At this time, a new type of fusion algorithm is proposed based on the spatio-temporal weight allocation strategy for water conservancy scenarios. By using the dynamic fusion algorithm to balance the stability of the local knowledge base and the real-time performance of networked data, the decision response time after fusion can be shortened by 30%, which is used to solve the contradiction between the offline environment and real-time update in traditional systems.
[0067] At this time, the dynamic fusion algorithm adopts a dynamic fusion formula of spatio-temporal exponential decay + normalization, which can smoothly transition between the offline / online two modes and effectively shorten the decision response time after fusion in water conservancy scenarios. Its specific usage method adopts a multi-dimensional credibility evaluation algorithm based on the unique time series characteristics of water conservancy projects. This algorithm comprehensively considers factors such as data timeliness, hydrological periodicity, and the impact of emergencies, and constructs a dynamic weight adaptive integration mechanism, including: (1) Calculate the timeliness trust value, and the calculation formula is:
[0068]
[0069] Among them, and represent the timeliness trust values of the local knowledge base and the networked data respectively. The longer the time since the local knowledge base was updated, the smaller it is. The shorter the time since the networked data was updated, the larger it is; represents the static credibility benchmark of the local knowledge base, represents the static credibility benchmark of the network search results, represents the time interval since the last full synchronization of the local knowledge base, represents the average latency of network data update crawling or the time since the article was published, represents the normalization factor for controlling the credibility decay of the local knowledge base ("time half-life"), represents the normalization factor for controlling the credibility decay of network data ("time half-life"); (2) Perform weight normalization operation, and we have:
[0070]
[0071] Among them, is the normalized weight of the local knowledge base, is the normalized weight of the network data, represents the network connectivity indication, indicates online availability, indicates switching to the pure local mode when offline; (3) Perform retrieval result score fusion. For each candidate , we have:
[0072] Among them, represents the score fusion result for each candidate , and respectively represent the local knowledge base retrieval score and the network data retrieval score for each candidate ; (4) Finally output the fusion result list in descending order.
[0073] The spatio-temporal exponential decay + normalization fusion algorithm first converts two types of data sources into decay trust scores according to the update time delay and spatial distance between the local knowledge base and the networked data, using adjustable spatio-temporal sensitivity and normalization factors. Then, these scores are normalized so that their sum is one, enabling automatic fallback to the pure local mode when offline or when local data is outdated, and quickly leaning towards online results when the network is unobstructed and the data is fresh. Finally, these normalized weights are used to weight-fuse the decision vectors of the local and the network, and the final decision is output. This method not only achieves smooth switching between offline / online but also significantly shortens the response time and improves spatio-temporal accuracy, meeting the dual requirements of real-time and stability in the water conservancy scenario.
[0074] In a further embodiment, the method also constructs a dual-channel knowledge fusion mechanism: First, a dual-channel knowledge fusion mechanism is constructed, which combines the local offline knowledge base and the real-time networked knowledge base, and effectively comprehensively processes information using a data fusion algorithm, enabling the stability of the offline knowledge base and the real-time update of network knowledge to achieve an organic connection between the two.
[0075] The dual-channel knowledge fusion mechanism also performs professional annotation and preprocessing on the collected data according to the special requirements of the water conservancy field, and fine-tunes the feature model in the field, thus significantly improving the system's ability to accurately identify and process hydrological information.
[0076] The dual-channel knowledge fusion mechanism also integrates the local knowledge base and the networked data through a data fusion algorithm, ensuring that the system can provide stable and reliable knowledge support in an offline environment, while guaranteeing the real-time update of networked data and enhancing the system's adaptability in a changing environment.
[0077] In a further embodiment, the method also optimizes the multi-modal task parallel processing technology: Based on the dual-channel knowledge fusion, the multi-modal task parallel processing technology is optimized to ensure that tasks such as image recognition, text generation, and speech recognition can be executed simultaneously without interference. Through efficient task scheduling and resource management, the execution efficiency of each task is improved, and it is ensured that the system can achieve seamless collaborative processing between different tasks, thus accelerating the response speed of decision support.
[0078] In a further embodiment, the method can also implement an integrated solution for an intelligent decision support system: Based on the dual-channel knowledge fusion and multi-modal task parallel processing technology, an integrated intelligent decision support solution is provided and applied to fields such as water conservancy project monitoring, emergency command, and knowledge services. By integrating different data sources and processing modules, it is ensured that the system can provide accurate and efficient decision support in various environments, providing comprehensive intelligent services for the water conservancy industry.
[0079] In a further embodiment, the method automatically schedules each processing module by generating a workflow, and develops a dedicated API interface and user interface to ensure a one-stop intelligent service from "asking questions" to "analysis" to "decision-making". For the first time, the system realizes the full-link automation of the local model data deployed in the server combined with the water conservancy scenario, significantly improving the stability and real-time response capabilities in the offline environment, and providing an integrated solution for water conservancy project monitoring, emergency command and knowledge services.
[0080] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0081] On the other hand, the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0082] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any of the workflow-based multi-modal water conservancy large model decision support methods in the above embodiments.
[0083] It is understandable that the system provided by the embodiment of the present invention corresponds to the method provided by the embodiment of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the above method.
[0084] The embodiment of the present application also provides an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus. Memory, used to store computer programs; The processor is used to implement the above-mentioned workflow-based multi-modal water conservancy large model decision support method when executing the program stored in the memory.
[0085] The communication bus mentioned in the above electronic device can be a peripheral component interconnect standard bus or an extended industrial standard architecture bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0086] The communication interface is used for communication between the above electronic device and other devices.
[0087] The memory may include a random access memory, or may include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0088] The above-mentioned processor may be a general-purpose processor, including a central processing unit, a network processor, etc.; it may also be a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0089] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part.
[0090] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0091] In addition, it should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.
[0092] In addition, if there are descriptions such as "first", "second", etc. involved in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the meaning of "and / or" appearing throughout the text includes three parallel scenarios. Taking "A and / or B" as an example, it includes scenario A, scenario B, or the scenario where both A and B are satisfied simultaneously. In addition, in the embodiments of the present invention, "a plurality" means two or more. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
Claims
1. A decision-making support method for a multi-modal water conservancy large model based on a workflow, characterized in that Including: S1. Preprocess the multi-modal input data based on the local model, and convert the data formats of files, images, and voices into JSON format; S2. By specifying the DSL framework, construct a modular workflow that supports parallel processing to ensure that different types of data after format conversion can be synchronously collected and processed; S3. Construct a feature model suitable for the water conservancy field through the synchronously collected and processed multi-modal data, build data dependency relationships based on the DSL framework, and perform domain adaptation and fine-tuning on the feature model. Embed the water conservancy mechanism into the model fine-tuning process after parameterizing the model through a three-layer architecture of implicit rules - explicit constraints - rule activation; S4. Based on the feature model after domain adaptation and fine-tuning, construct a specified workflow for dynamically scheduling multi-modal task resources to achieve a dynamic processing process from data input, analysis to decision output; S5. Use the dynamic fusion algorithm to combine the local knowledge base and network data to achieve dual-channel knowledge fusion for parallel execution of the multi-modal tasks in S4 to ensure the stability of the offline knowledge base and the real-time update of network data.
2. The multi-modal water conservancy large model decision support method based on workflow according to claim 1, wherein, The specific operation process of the preprocessing in step S1 includes: S11. Based on the obtained multi-modal input data, perform preliminary classification according to the specified type standard to ensure that they enter the corresponding preprocessing process by type respectively; S12. For different types of multi-modal data, use specified models for preprocessing respectively: For text data, extract semantic features through an embedding model, and then use a language model for noise removal, format unification, word segmentation, and part-of-speech tagging; For image data, use an embedding model for image feature extraction, perform denoising, resolution adjustment, and feature point extraction; For voice data, use a voice model for voice recognition and processing, including noise removal, voice conversion to text, and voice feature extraction.
3. The decision-making support method for the multimodal water conservancy large model based on workflow according to claim 1, characterized in that The specific construction process of the modular workflow in step S2 includes: S21. In the specified DSL framework, define different data processing modules and adopt a parallel processing architecture to ensure that the specified data types perform data processing simultaneously in different processing channels; S22. Based on the specified DSL framework, achieve optimization and coordination of multi-task parallel execution by defining task dependency relationships and execution orders.
4. The decision-making support method for the multimodal water conservancy large model based on workflow according to claim 1, wherein, The specific operation process of step S3 includes: S31. Use a feature extraction model suitable for the water conservancy field to extract specified features from different data types of multi-modal data in the water conservancy field to obtain a feature model; S32. Construct pre-training data based on the existing water conservancy field knowledge base, and perform domain adaptation operations on the feature model using the MLM task; S33. Perform LoRA lightweight adjustment operation on the feature model, freeze the parameters of the low-rank matrix of the pre-trained data model, only adjust the newly added LoRA adaptation layer, and use the labeled dataset in the water conservancy field for supervised adjustment. Incorporate the specified control indicators of water conservancy projects as explicit constraint conditions into the model adjustment process to build a physical consistency verification mechanism for checking whether the output results violate the hydrological physical laws during each model update, and adopt a penalty term mechanism to correct the predictions that violate the physical laws to ensure the physical rationality of the model output.
5. The decision-making support method for the multimodal water conservancy large model based on the workflow according to claim 4, characterized in that After the execution of the S3 step, perform model verification and dynamic optimization operations, specifically including: Construct a test set in the water conservancy field for domain accuracy evaluation, and use F1-score and BLEU metrics to verify the model performance; Construct a set of cross-modal consistency samples for inspection to ensure the logical consistency of outputs in different modalities; Deploy the model by integrating the A / B test module, and adopt a gray release strategy to trigger incremental training of the model through user feedback data to perform online dynamic optimization and closed-loop improvement.
6. The workflow-based multi-modal water conservancy large model decision support method according to claim 5, wherein, During the inspection of the cross-modal consistency samples, adopt a cross-modal alignment algorithm to ensure the semantic consistency of data in different modalities. The usage method of the cross-modal alignment algorithm includes: L1. Perform knowledge graph embedding pre-training in the form of TransE for all triples Minimize the knowledge graph embedding pre-training loss The calculation formula is as follows: Among them, is the margin hyperparameter of the model, which is used to ensure the effective distinction between positive and negative samples in the pre-training of knowledge graph embedding; after training, all and are fixed; , , are the vectors of the head entity, relation, and tail entity in the graph embedding space respectively, and represent the starting entity of the relation, the semantic relation between entities, and the target entity of the relation in the knowledge graph respectively. is the L2 norm, is the set of all triples in the knowledge graph; L2. For each positive sample , sample several negative entities from to form . Perform cross-modal feature mapping and alignment loss, and there is Among them, is a set of negative samples or a space containing negative samples. Specifically, it is the source from which negative entities are sampled and used to optimize by comparing with positive samples during the training process. represents a mapping network for mapping visual features and text features into an embedding space with the same dimension as the entity. is the hinge loss of the cross-modal alignment loss, which includes visual alignment terms and text alignment terms. represents the set of negative entities sampled for positive entities represents the margin hyperparameter of the alignment loss. represents the similarity function between A and B; L3. Construct training objectives, including: Among them, , which is used to represent the set of weights and biases of the visual-to-text mapping network, is the weight of the visual mapping network, is the bias of the visual mapping network, is the weight of the text mapping network, is the bias of the text mapping network, is the regularization term, and its value is the sum of the squares of all trainable parameters, is the L2 regularization weight.
7. The decision-making support method for the multimodal water conservancy large model based on workflow according to claim 4, characterized in that In the S32 step, an adversarial training mechanism is also introduced. Introduce a domain feature adversarial network, use a domain classifier to distinguish general data from water conservancy data to enhance the sensitivity of the feature model to professional terms, enable the feature model to implicitly learn the physical constraint relationships of water conservancy projects, and set specified pre-training tasks to make the feature model predict results that conform to hydrological physical laws under given boundary conditions to achieve implicit encoding and storage of mechanism knowledge in the model weights. The specific operation steps of the adversarial training mechanism include the following adversarial training improvement method: P1. Divide the water conservancy dataset into a source domain and a target domain, and eliminate data distribution differences through standardization processing; P2. Generate the feature extraction module of the adversarial network, combine the deep residual shrinkage network to compress redundant feature dimensions, and introduce a long short-term memory network to capture temporal dependencies to form a model architecture. Add a domain discriminator as the core of adversarial training to the model architecture, and force the feature encoder to generate domain-invariant features through a gradient reversal layer; P3. Based on the model architecture constructed in the P2 step, perform adversarial training operations, including: Stage 1: Pre-train the feature encoder on the source domain data to initially extract key features related to water conservancy parameters; Stage 2: Fix the feature encoder, train the domain discriminator to distinguish source domain and target domain features, and optimize the feature encoder to confuse the discriminator by minimizing the domain discriminator loss function to achieve domain feature alignment; Stage 3: On the basis of adversarial training, jointly fine-tune the model using source domain labeled data and a small amount of target domain labeled data, and adopt a weighted cross-entropy loss function to balance the problem of class imbalance; P4. Embed a dynamic threshold judgment module in the output layer of the model to adaptively adjust the classification threshold according to the statistical features of water conservancy project historical data.
8. The decision-making support method for the multimodal water conservancy large model based on workflow according to claim 1, characterized in that, The specific operation process of the S4 step includes: S41. Based on the fine-tuned feature model, construct a set of "question - analysis - decision" workflows, define the task transfer relationships between each sub-module, ensure that each module can start corresponding analysis and processing tasks according to the user's input, and smoothly advance to the decision output according to the predetermined logic; S42. Adopt a preset task priority scheduling strategy to optimize the task execution efficiency in the workflow for dynamic resource scheduling. The scheduling judgment logic of the task priority scheduling strategy includes: Dynamically adjust the baseline priorities of various tasks according to the water conservancy operation status: During the flood season, automatically elevate the real-time sensor data task to the highest priority; For flood control command and routine monitoring tasks, during the non-flood season, restore to the routine priority setting; The scheduler adopts preemptive scheduling. When a new task arrives, compare the priority of the new task with the task being executed. If the former is higher, immediately interrupt and switch, otherwise continue to execute; When the priorities are the same, allocate the CPU according to the first-come-first-served or time slice rotation; 9. The decision-making support method for the multimodal water conservancy large model based on workflow according to claim 1, wherein, The specific operation process of performing the dynamic fusion algorithm in step S5 in combination with the data includes: S51. Obtain the offline knowledge base and real-time data through the acquisition interfaces of the local knowledge base and networked data respectively, and perform independent preprocessing operations on them; S52. Combine the local knowledge base and networked data. The offline knowledge base provides stable information, the networked data is updated in real time and provides the latest decision support, and a dynamic fusion algorithm is used to balance the stability of the local knowledge base and the real-time nature of the networked data.
10. The decision-making support method for the multimodal water conservancy large model based on the workflow according to claim 9, wherein The usage method of the dynamic fusion algorithm includes; (1) Calculate the timeliness trust value, and the calculation formula is: Among them, and represent the timeliness trust values of the local knowledge base and the networked data respectively. The longer the time since the local knowledge base was updated, the smaller it is. The shorter the time since the networked data was updated, the larger it is; Represents the static credibility benchmark for the local knowledge base, Represents the static credibility benchmark for the online search results, Represents the time interval since the last full synchronization of the local knowledge base, Represents the average latency of network data update crawling or the time since the article was published, Represents the normalization factor for controlling the credibility decay of the local knowledge base, Represents the normalization factor for controlling the credibility decay of network data; (2) Perform weight normalization operation, and there is: Among them, is the normalized weight of the local knowledge base, is the normalized weight of the network data, represents the network connection indication, indicates online availability, indicates the offline switch to the pure local mode; (3)Perform retrieval result score fusion. For each candidate , there is: Among them, represents the score fusion result for each candidate , and respectively represent the local knowledge base retrieval score and the network data retrieval score for each candidate . (4)Finally Output the fusion result list in descending order.
Citation Information
Patent Citations
Self-adaptive extension method and system for knowledge base
CN113268604A
Construction method of pre-training model, computer device and computer readable storage medium
CN117473310A
Multi-modal reasoning method and device based on large language model and knowledge graph
CN118193684A
Multi-modal large model training method and system fusing time series data of Internet of Things
CN118296462A
Domain large model multi-modal knowledge base construction method based on feature representation
CN118779469A
Cited By
Knowledge base-based reasoning comparison method
CN121119142A
Hydrodynamic analysis scheduling method, system and equipment based on large language model
CN121213295A
Water conservancy design decision support system based on knowledge graph
CN121328351A
Water conservancy design decision support system based on knowledge graph
CN121328351B
Equipment inspection method and device
CN121788959A