A water affair decision system based on multi-modal model fusion

Through the water conservancy decision-making system that integrates multimodal models, virtual digital human interaction and intelligent business middle platform, the cross-modal data fusion and real-time reasoning problems of the water conservancy emergency decision-making system are solved, efficient and intelligent human-computer collaboration is achieved, and the accuracy of flood forecasting and decision-making is improved.

CN120634051BActive Publication Date: 2025-10-24ZHEJIANG KEEPSOFT INFORMATIONTECHNOLOGY CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511106009.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-10-24
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

The existing water conservancy emergency decision-making system has significant technical bottlenecks in cross-modal data fusion, real-time reasoning in complex business scenarios, and natural human-computer interaction, and cannot meet the time limit requirements. Traditional system analysis and solution generation require 6 to 8 hours.

Method used

A water affairs decision-making system based on multimodal model fusion is adopted, including a virtual digital human interaction module, an intelligent business middle platform and a large model technology base. It integrates a water conservancy professional terminology intention recognition model, a water science mechanism sub-model and a dynamic knowledge distillation module, and achieves efficient decision-making through multimodal sensors and emotion recognition models.

Benefits of technology

It has achieved efficient, intelligent and humanized human-machine collaboration in water conservancy emergency management, improved flood forecasting accuracy and decision-making accuracy, shortened decision-making time, and provided reliable water conservancy emergency response support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634051B_ABST
    Figure CN120634051B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computers, in particular to a water affair decision system based on multi-modal model fusion, which comprises a virtual digital human interaction module, an intelligent agent business platform and a large model technology base, the virtual digital human interaction module is used for listening to voice instructions of a user, acquiring biological feature data of the user by using multi-modal sensors, identifying the relevance features between the voice instructions and the biological feature data by using a water conservancy emergency scene emotion recognition model, acquiring indication information, and sending the indication information to the intelligent agent business platform; the intelligent agent business platform analyzes the indication information based on a water conservancy professional term intention recognition model to obtain different task requirements, calls the large model technology base based on the different task requirements to make decisions, and obtains decision information; the large model technology base is integrated with a water conservancy large model trained based on a domain mechanism model library. The application can provide more efficient and intelligent services.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer technology, and in particular to a water affair decision system based on multi-modal model fusion. BACKGROUND

[0002] With the development of artificial intelligence (AI) technology, especially the breakthroughs in large language models (LLM) and multi-modal models, traditional software service models have been unable to meet the increasingly complex business needs. Existing platforms have significant technical bottlenecks in cross-modal data fusion (such as the spatio-temporal alignment of hydrological monitoring data and remote sensing images), real-time inference in complex business scenarios (such as flood prediction and emergency plan generation), and natural human-computer interaction (such as semantic understanding of water conservancy professional terms). In the context of water conservancy emergency decision-making, traditional systems take an average of 6-8 hours to complete multi-source data analysis and plan generation, which cannot meet the time limit requirements.

[0003] Therefore, there is an urgent need for a water affair decision system based on multi-modal model fusion. SUMMARY

[0004] (I) Technical solution

[0005] To achieve the above purpose, the main technical solution adopted by the present application includes:

[0006] The present application provides a water affair decision system based on multi-modal model fusion, which includes a virtual digital human interaction module, an intelligent agent business platform, and a large model technology base, wherein,

[0007] The virtual digital human interaction module is used to listen to the user's voice instructions, acquire the user's biological feature data using multi-modal sensors, and identify the association features between the voice instructions and the biological feature data through a water conservancy emergency scenario emotion recognition model to obtain indication information, and send the indication information to the intelligent agent business platform.

[0008] The intelligent agent business platform analyzes the task chain of the indication information based on a water conservancy professional term intent recognition model, obtains different task requirements, and calls the large model technology base based on different task requirements to make decisions and obtain decision information.

[0009] The large model technology base is integrated with a water conservancy large model trained based on a domain mechanism model library, the domain mechanism model library includes a plurality of water science mechanism sub-models, and the water conservancy large model couples the knowledge bases of the plurality of water science mechanism sub-models during the training process.

[0010] Optionally, the virtual digital human interaction module is further configured to receive decision information sent by the intelligent agent business platform, and render a VR scene based on the decision information.

[0011] Optionally, the intelligent agent business platform is configured to perform data matching on sub-scenes of various field scenes expressed by the virtual digital human interaction module based on different task requirements when calling the large model technology base.

[0012] Optionally, the large model technology base comprises:

[0013] a dynamic knowledge distillation module, a digital twin base, and a feedback optimization module,

[0014] The dynamic knowledge distillation module is configured to collect water conservancy data, input the water conservancy data into various water science mechanism sub-models to obtain flood evolution key parameters, encode the flood evolution key parameters into a structured feature vector, fuse the spatial and temporal distribution characteristics in the decision intermediate result with the semantic representation of the water conservancy large model through a cross-modal attention mechanism, and obtain the spatial and temporal distribution characteristics in the decision intermediate result.

[0015] The dynamic knowledge distillation module is further configured to drive each water science mechanism sub-model to perform Monte Carlo simulation based on the spatial and temporal distribution characteristics in the decision intermediate result, to generate an enhanced training data set covering the rainstorm intensity-flooded area mapping relationship, and to update the weight parameters of the water conservancy large model using the ADMM algorithm with the enhanced training data set as input, to obtain a trained water conservancy large model.

[0016] The feedback optimization module is configured to collect decision feedback data of the intelligent agent business platform in real time, and dynamically adjust the attention weight of the water conservancy large model according to the decision feedback data, to update and optimize the trained water conservancy large model.

[0017] Optionally, the water science mechanism sub-models include at least two of a distributed hydrological model, a hydrodynamic model, a water quality model, a pipe network model, and a hydraulic model.

[0018] Optionally, the digital twin base is a virtual network model that fuses satellite remote sensing data and ground observation data, wherein the virtual network model is a virtual network model that uses a visualization modeling tool to establish a target area involving all water elements and processes, from grid precipitation in the sky to surface runoff, river flood evolution, reservoir water conservancy scheduling, and underground pipe network modeling.

[0019] Optionally, the intelligent agent business platform comprises a plurality of scene intelligent agents and a task chain analysis engine.

[0020] The task chain analysis engine decomposes the task requirement into an atomic operation chain, and calls different scene intelligent agents according to the atomic operation chain, so as to generate the decision information by calling the water conservancy large model trained on the large model technology base.

[0021] Optionally, the water conservancy emergency scene emotion recognition model comprises a multi-modal feature extraction layer and a scene-based emotion classification layer.

[0022] The multi-modal feature extraction layer adopts a Mel spectrum convolution network to extract voice instructions and voice emotion spectrum features of the voice instructions, and adopts a lightweight ResNet-18 network to extract facial action unit intensity features.

[0023] The scene-based emotion classification layer adopts a graph attention network to fuse the voice instructions, the voice emotion spectrum features and the facial action unit intensity features, and outputs the instruction information.

[0024] Optionally, the water conservancy emergency scene emotion recognition model is trained by using an adversarial training mechanism.

[0025] Optionally, the virtual digital human interaction module comprises:

[0026] a multi-modal sensor array and a VR scene dynamic mapping unit.

[0027] The multi-modal sensor array comprises a high-definition camera, an infrared thermal imager and an array microphone, so as to collect facial expressions, body temperature changes and voices of a user.

[0028] The VR scene dynamic mapping unit is used to convert the decision information into a three-dimensional particle system for rendering.

[0029] (II) Beneficial effects

[0030] The water management decision system based on multi-modal model fusion has the beneficial effects that: the dynamic knowledge injection mechanism is adopted, the parameter bidirectional optimization of a water science mechanism submodel and a general large model is realized, and therefore more efficient and intelligent software services are provided; meanwhile, the virtual digital human interaction module and the VR digital twin scene are integrated, the human-computer interaction paradigm of water management is redefined, and a reusable technical framework is provided for digital transformation of smart water infrastructure. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 FIG. 1 is an architecture diagram of a water management decision system based on multi-modal model fusion according to an embodiment of the present application;

[0032] Figure 2 FIG. 2 is an internal function block diagram of the water management decision system based on multi-modal model fusion according to the embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to better explain the present application, so as to be understood, the present application is described in detail below by specific embodiments in combination with the drawings.

[0034] The water affair decision system based on multi-modal model fusion provided by the embodiment of the present application proposes full-coupling digital twinning, which is different from the single condition of considering only rainfall input in the traditional model. The present application establishes a model from grid rainfall in the sky to surface runoff, river flood evolution, reservoir water conservancy dispatching, and underground pipe network in the target area in a full-factor and full-process way, thereby significantly improving the prediction accuracy of flood. At the same time, the present application adopts a dynamic knowledge injection mechanism to realize the double breakthrough of improving flood control efficiency and optimizing flood control cost. Further, the present application also integrates a water conservancy emergency scene emotion recognition model, which can monitor 7 types of emotional states such as anxiety and urgency of users in real time, and associate the emotion recognition result with biological feature data, thereby effectively improving the accuracy of decision-making.

[0035] In order to better understand the above technical solutions, the exemplary embodiments of the present application will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided so that the present application can be more clearly, thoroughly understood, and the scope of the present application can be completely conveyed to those skilled in the art.

[0036] Embodiment 1

[0037] Referring to Figure 1 The water affair decision system based on multi-modal model fusion of the present embodiment comprises:

[0038] a virtual digital human interaction module, an agent business platform, and a large model technology base, wherein,

[0039] The virtual digital human interaction module is used to listen to the voice instructions of the user, acquire the biological feature data of the user by using multi-modal sensors, and identify the association feature between the voice instructions and the biological feature data by using the water conservancy emergency scene emotion recognition model, acquire the indication information, and send the indication information to the agent business platform.

[0040] In the specific implementation process, the virtual digital human interaction module is the core component of natural interaction between the entire system and the user. The virtual digital human interaction module has multi-modal perception ability, can listen to the voice instructions of the user in real time, and acquire the biological feature data of the user by using integrated various sensors such as a camera, a heart rate detector, a galvanic skin response sensor, and the like, including facial expressions, eye movement trajectories, heart rate changes, micro expressions, and the like physiological signals.

[0041] After obtaining the voice instruction and the biometric data, the virtual digital human interaction module utilizes a water conservancy emergency scene emotion recognition model to perform deep analysis and fusion recognition on the correlation features between the multi-source information. The water conservancy emergency scene emotion recognition model is trained based on a large amount of user behavior and emotion data in water conservancy industry related scenes, and can recognize the emotional state of the user when facing a water conservancy emergency, so as to more accurately extract the actual intention of the user.

[0042] The intelligent agent business middle platform performs task chain analysis on the indication information based on a water conservancy professional term intention recognition model, obtains different task requirements, and calls a large model technology base based on different task requirements to make decisions and obtain decision information.

[0043] The intelligent agent business middle platform is a bridge connecting the user interaction layer and the bottom layer of the large model decision layer, that is, connecting the virtual digital human interaction module and the large model technology base, and undertakes the core functions of task analysis, logical reasoning and resource scheduling. Further, the intelligent agent business middle platform has a built-in water conservancy professional term intention recognition model, which is trained based on a water conservancy domain knowledge graph and a large amount of historical instruction corpus, and has the ability to understand complex water conservancy terms and ambiguous expressions. Through semantic analysis of the indication information from the virtual digital human interaction module, the intelligent agent business middle platform can accurately identify the operation intention of the user, such as "viewing the rainfall trend of a certain river basin" and "evaluating the downstream flood discharge impact range".

[0044] Further, the intelligent agent business middle platform will perform task chain decomposition and arrangement according to the analysis results, decompose the complex user request into multiple executable sub-tasks, call the large model technology base based on different task requirements to make decisions, and obtain decision information.

[0045] The large model technology base has integrated a water conservancy large model trained based on a domain mechanism model library. The domain mechanism library includes multiple water science mechanism sub-models, and the water conservancy large model couples the knowledge bases of multiple water science mechanism sub-models during the training process.

[0046] The large model technology base is the intelligent core of the entire system, and carries the capabilities of complex reasoning, simulation prediction and knowledge fusion. The base integrates a water conservancy large model specially for the water conservancy field, and fully integrates professional knowledge of hydrology, hydraulics, water resources management, flood control and disaster reduction and other disciplines during the construction process.

[0047] The water affair decision system based on multi-modal model fusion of the embodiment realizes an efficient, intelligent and humanized man-machine cooperation mechanism in the water conservancy emergency management scene by constructing a "virtual digital human interaction module-intelligent agent business middle platform-large model technology base" three-in-one system architecture.

[0048] Embodiment 2

[0049] One embodiment of the water affair decision system based on multi-modal model fusion, as shown in the figure, comprises: Figure 2

[0050] a virtual digital human interaction module, an agent business middle platform and a large model technology base, wherein,

[0051] The virtual digital human interaction module, i.e. the intelligent interaction service layer, is used to listen to the voice instructions of the user, acquire the biological feature data of the user by using the multi-modal sensor, and identify the correlation features between the voice instructions and the biological feature data by using the water conservancy emergency scene emotion recognition model, acquire the indication information, and send the indication information to the agent business middle platform.

[0052] The agent business middle platform analyzes the task chain of the indication information based on the water conservancy professional term intention recognition model, obtains different task requirements, and calls the large model technology base based on different task requirements to make decisions and obtain decision information.

[0053] The large model technology base is integrated with a water conservancy large model trained based on a domain mechanism model library, the domain mechanism library comprises a plurality of water science mechanism sub-models, and the water conservancy large model couples the knowledge bases of the plurality of water science mechanism sub-models in the training process.

[0054] The virtual digital human interaction module is also used to receive the decision information sent by the agent business middle platform, and render the decision information into a VR scene for expression based on the decision information.

[0055] In the embodiment, the agent business middle platform performs data matching on the sub-scenes of various domain scenes expressed by the virtual digital human interaction module based on different task requirements when calling the large model technology base.

[0056] Further, in the specific implementation process, the large model technology base comprises:

[0057] a dynamic knowledge distillation module, a digital twin base and a feedback optimization module,

[0058] The dynamic knowledge distillation module is used to collect water conservancy data, input the water conservancy data into each water science mechanism sub-model to obtain flood evolution key parameters, encode the flood evolution key parameters into a structured feature vector, fuse the time and space distribution features in the decision intermediate results by using the cross-modal attention mechanism and the semantic representation of the water conservancy large model, and obtain the time and space distribution features in the decision intermediate results.

[0059] ​The dynamic knowledge distillation module is also used to drive each water science mechanism sub-model to perform Monte Carlo simulation based on the spatio-temporal distribution features in the decision intermediate result, to generate an enhanced training data set covering a storm intensity-flooded area mapping relationship, and to update the weight parameters of the water conservancy large model using an ADMM algorithm with the enhanced training data set as input, to obtain a trained water conservancy large model.

[0060] Specifically, the dynamic knowledge distillation module comprises:

[0061] The water conservancy data acquisition and processing unit is used to collect historical water conservancy data from multiple sources such as meteorological stations, hydrological monitoring points, remote sensing platforms, GIS databases, etc., including but not limited to rainfall, water level, flow rate, terrain elevation, land use type, etc.

[0062] The water science mechanism sub-model calling unit inputs the collected water conservancy data into multiple water science mechanism sub-models such as flood routing models, rainfall-runoff models, etc., simulates different hydrological processes, and outputs key parameters such as maximum flooded depth, flood peak arrival time, flow variation curve, etc.

[0063] The structured feature encoding unit uses graph neural networks, spatio-temporal convolution networks or multi-head attention mechanisms to encode the above key parameters into structured feature vectors of a unified dimension, facilitating subsequent fusion processing.

[0064] The cross-modal attention fusion unit constructs a cross-modal attention mechanism to fuse the structured feature vectors with the semantic representations of the water conservancy large model to obtain the spatio-temporal distribution features in the decision intermediate result.

[0065] The enhanced data generation and model updating unit is used to drive each water science mechanism sub-model to perform Monte Carlo simulation based on the spatio-temporal distribution features in the decision intermediate result, to simulate flood routing processes under different storm intensities, underlying surface conditions and dispatching strategies, and to generate an enhanced training data set containing a "storm intensity-flooded area" mapping relationship. Subsequently, the weight parameters of the water conservancy large model are updated in a distributed manner using the ADMM algorithm (Alternating Direction Method of Multipliers) to ensure that the model maintains physical consistency while having stronger data fitting capability and generalization performance, thereby obtaining a trained water conservancy large model.

[0066] In the specific implementation process, the Monte Carlo simulation specifically includes the following steps:

[0067] A large number of training samples with physical consistency are generated by simulating flood routing processes under different combinations of storm intensities, underlying surface conditions and water conservancy engineering dispatching strategies under uncertainty conditions through high-dimensional random sampling methods.

[0068] Then, the initial hydrological state and boundary conditions are set according to the spatio-temporal distribution characteristics, and the Latin hypercube sampling is used to generate parameter combinations; the parameter combinations are input into the mechanism sub-model of water science, the runoff process is simulated by the distributed hydrological model, the flood propagation is deduced by the hydrodynamic model, the urban drainage capacity is evaluated by the pipe network model, and the safety of the hydraulic structure is judged by the hydraulic model.

[0069] The simulation results of each mechanism sub-model of water science are output, and the input parameters and output results of each simulation are structured and packaged to form sample records in a unified format.

[0070] For example, after a historical rainstorm event, the system extracts the spatio-temporal distribution characteristics of the event (such as: severe water accumulation in the upstream hilly area, overflow in the middle reaches of the river, and blockage of the downstream urban pipe network) through the dynamic knowledge distillation module. Subsequently: Monte Carlo simulation simulates the flood evolution under different rainstorm intensities (100mm vs 200mm) and different scheduling strategies (early flood release vs normal operation); generates thousands of "rainstorm intensity-flooded area" mapping samples; inputs the samples into the water conservancy big model for retraining; thereby significantly improving the recognition accuracy of "urban waterlogging points" and "weak links of dikes" in subsequent similar events.

[0071] Further, the general big model in the embodiment includes BERT, GPT, CLIP, DEEPSEEK, and other big models.

[0072] The feedback optimization module is used to collect decision feedback data of the agent business platform in real time, and dynamically adjust the attention weights of the water conservancy big model according to the decision feedback data, so as to update and optimize the trained water conservancy big model.

[0073] The feedback optimization module is a key component for the system to realize closed-loop evolution, aiming to dynamically adjust the behavior strategy and internal parameters of the water conservancy big model by collecting real-time decision feedback data of users, so as to continuously improve its intelligent level.

[0074] In the embodiment, the mechanism sub-model of water science includes at least two of the distributed hydrological model, the hydrodynamic model, the water quality model, the pipe network model, and the hydraulic model.

[0075] The distributed hydrological model is used for simulating spatial heterogeneity distribution of hydrological processes such as precipitation-evapotranspiration-runoff-surface water recharge in a basin scale, can provide upstream inflow boundary conditions for flood routing, and support analysis of basin response under multiple scenarios of rainfall; the hydrodynamic model can simulate water flow movement process in a river channel and a plain area, realize dynamic deduction of flood inundation range, depth, flow velocity and other elements, and can be used for flood risk assessment and auxiliary decision-making for emergency dispatch under extreme events such as urban waterlogging, river dam break and typhoon storm; the water quality model simulates migration, transformation and fate of pollutants in water, considers factors such as biochemical reaction, sediment adsorption and water temperature influence, and can support tasks such as water resource protection, emergency disposal of pollution accidents and effect evaluation of ecological restoration engineering. The pipe network model can simulate the whole process of rainwater collection, transportation, storage and discharge, and provide technical support for urban flood control and drainage planning, sponge city construction and underground pipe network reconstruction. The hydraulic structure model simulates stress, stability and operation state of key water conservancy engineering facilities such as reservoir dam, gate, pump station, spillway and dike, and can be used for safety monitoring, health diagnosis, dispatch optimization and disaster warning of major water conservancy projects.

[0076] The above-mentioned water science mechanism sub-model is fused with a general large model to obtain a water conservancy large model, so that the system can effectively fuse physical laws while maintaining data-driven modeling capability, and improve prediction accuracy and decision reliability.

[0077] In addition, the digital twin base is a virtual network model fused with satellite remote sensing data and ground observation data, wherein the virtual network model is a target area full-factor and full-process model from grid precipitation in the sky to surface runoff, river flood evolution, reservoir water conservancy dispatch and underground pipe network, which is established by using a visual modeling tool.

[0078] In the system, the digital twin base establishes a virtual network model covering the whole process from grid precipitation in the sky to surface runoff, river flood evolution, reservoir water conservancy dispatch and underground pipe network by integrating satellite remote sensing data and ground observation data, so as to realize comprehensive perception, dynamic simulation and intelligent deduction of water-related processes in the target area.

[0079] The satellite remote sensing data includes meteorological satellites such as Fengyun series, GPM precipitation inversion data, optical remote sensing images for land use classification and water body identification, and SAR radar images for flood inundation range detection; the ground observation data includes data obtained by meteorological stations, hydrological stations, urban waterlogging monitoring points and underground pipe network sensors.

[0080] Further, the virtual network model with real geographical space attributes is constructed by using a visual modeling tool based on BIM (Building Information Modeling), GIS, ENVI, CAD and professional water conservancy modeling software.

[0081] By constructing a virtual network model that fuses satellite remote sensing and ground observation data, the digital twin base realizes fine modeling and dynamic deduction of "all elements involved in water" and "the whole process" in the target area, providing strong technical support for key businesses such as flood control and drought relief, urban drainage, and water resources scheduling. It not only enhances the visualization capabilities and interactive experience of the system, but also provides a high-quality physical simulation environment for the training and optimization of large water models, promoting the development of smart water conservancy to a higher level.

[0082] Specifically, the intelligent agent business platform in the embodiment includes a plurality of scene intelligent agents and a task chain analysis engine.

[0083] The task chain analysis engine decomposes the task demand into an atomic operation chain and calls different scene intelligent agents according to the atomic operation chain, so as to generate decision information by calling the water conservancy large model trained on the large model technology base.

[0084] The task chain analysis engine is the "brain" of the intelligent agent business platform, and its main functions include:

[0085] Natural language understanding and intent recognition of the instruction information from the virtual digital human interaction module; decompose complex tasks into executable atomic operation sequences; select appropriate scene intelligent agents according to task type and context environment; build a task execution plan and schedule each intelligent agent to cooperate in sequence to complete the task.

[0086] The scene intelligent agent is a professional intelligent agent designed for a specific water conservancy business scenario. Each intelligent agent is responsible for the execution of a typical task, for example, scene intelligent agents include flood warning intelligent agents, reservoir scheduling intelligent agents, urban waterlogging intelligent agents, and water resource allocation intelligent agents.

[0087] Among them, the flood warning intelligent agent monitors water regime changes in real time, predicts flood peak arrival time, and identifies high-risk areas; the reservoir scheduling intelligent agent develops an optimal water release strategy based on water inflow forecasts and downstream safety requirements; the urban waterlogging intelligent agent analyzes pipe network full-flow conditions and recommends pump station scheduling and water accumulation point disposal measures; the water resource allocation intelligent agent balances agricultural water and ecological water demand and optimizes water intake scheduling schemes.

[0088] By constructing an intelligent agent business platform that includes a task chain analysis engine and multiple scene intelligent agents, accurate analysis and efficient execution of instruction information are achieved. This platform not only has good scalability and flexibility, but also can quickly switch and deploy corresponding scene intelligent agents according to different water conservancy business scenarios, significantly improving the adaptability and decision-making efficiency of the system.

[0089] In this embodiment, the water emergency scene emotion recognition model includes a multi-modal feature extraction layer and a scene-based emotion classification layer, wherein

[0090] The multi-modal feature extraction layer extracts the speech instruction and the speech emotion spectrum feature by using a Mel spectrum convolution network, and extracts the facial action unit intensity feature by using a lightweight ResNet-18 network.

[0091] The scene-based emotion classification layer fuses the speech instruction, the speech emotion spectrum feature and the facial action unit intensity feature by using a graph attention network, and outputs the instruction information.

[0092] Further, the water emergency scene emotion recognition model is trained by using an adversarial training mechanism. By introducing an adversarial sample to enhance the training process, the model can still maintain stable emotion recognition and intention understanding ability when facing actual challenges such as noise interference, accent difference or light change. This design effectively improves the perception accuracy of the system to the real needs of the user, and provides a high-quality input basis for subsequent task analysis and intelligent decision-making.

[0093] Specifically, the multi-modal feature extraction layer extracts the speech instruction information and the corresponding speech emotion spectrum feature from the user's speech instruction by using a Mel spectrum convolution network, and analyzes the user's facial image by using a lightweight ResNet-18 network to extract the intensity feature of the facial action coding unit, so as to comprehensively capture the language content and emotional state of the user. Subsequently, the scene-based emotion classification layer uses a graph attention network to cross-modally fuse the speech instruction, the speech emotion feature and the facial action intensity feature, combines the semantic logic under the water emergency scene, and outputs the structured instruction information containing the emotional state of the user.

[0094] In addition, the water professional term intention recognition model in this embodiment is a deep learning model, which is trained based on a large amount of user behavior and emotion data under water industry related scenes, and can recognize the emotional state of the user when facing water emergency events, so as to more accurately extract the actual intention of the user.

[0095] In this embodiment, the virtual digital human interaction module includes:

[0096] A multi-modal sensor array and a VR scene dynamic mapping unit.

[0097] The multi-modal sensor array includes a high-definition camera, an infrared thermal imager and an array microphone to collect facial expressions, body temperature changes and speech of the user.

[0098] The VR scene dynamic mapping unit is used to convert the decision information into a three-dimensional particle system for rendering.

[0099] Among them, the multi-modal sensor array is used to collect the facial expressions, voice instructions, body temperature changes and other biometric data of the user, to construct a comprehensive user state portrait, and to provide high-quality input for subsequent emotion recognition and intent analysis. The VR scene dynamic mapping unit is responsible for converting the decision information generated by the system into three-dimensional visual content with immersion, enhancing the user's understanding and corresponding efficiency of water conservancy emergency events. By constructing a virtual reality environment and combining particle system rendering technology, dynamic display of key information such as flood evolution, dispatching recommendations, and risk areas is realized.

[0100] For example, when the system determines that a certain reservoir is facing dam collapse risk, the VR scene dynamic mapping unit can render the following three-dimensional visual content in real time:

[0101] Animation demonstration of rising water level; red highlight flashing prompt of embankment weak point; particle cloud diffusion simulation of downstream inundation area; comparative demonstration of recommended dispatching strategies (such as early flood release vs. no action);

[0102] By setting up a multi-modal sensor array and a VR scene dynamic mapping unit, the flood evolution path, risk area distribution, and dispatching strategy comparison can be intuitively and vividly presented, improving the user's cognitive efficiency and decision confidence in complex water conservancy events.

[0103] The water affair decision system based on multi-modal model fusion of the embodiment, by constructing a smart water conservancy interactive system integrating "multi-modal perception-scenario-based emotion recognition-task chain analysis-water conservancy big model decision-three-dimensional visual feedback", significantly improves the human-computer cooperation efficiency and intelligent level in water conservancy emergency management.

[0104] In the description of the present application, it should be understood that the terms "first", "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can be explicitly or implicitly included one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.

[0105] In the present application, unless otherwise specifically defined and limited, the terms "installation", "connection", "connection", "fixation" and other terms should be understood broadly, for example, it can be fixed connection, or detachable connection, or integrated; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through intermediate medium; it can be the internal communication of two elements or the interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0106] In the present application, unless otherwise explicitly specified and limited, a first feature is "on" or "under" a second feature can mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, a first feature is "over", "above" and "on top of" a second feature can mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is horizontally higher than the second feature. A first feature is "under", "below" and "underneath" a second feature can mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is horizontally lower than the second feature.

[0107] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are contained in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0108] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and the person skilled in the art can make changes, modifications, replacements and variations to the above-described embodiments within the scope of the present application.

Claims

1. A water affair decision system based on multi-modal model fusion, characterized in that, The system comprises a virtual digital human interaction module, an agent business middle platform, and a large model technology base. The virtual digital human interaction module is configured to listen to voice instructions of a user, acquire biological feature data of the user by using a multi-modal sensor, identify association features between the voice instructions and the biological feature data by using a water conservancy emergency scene emotion recognition model, acquire indication information, and send the indication information to the agent business middle platform. The agent business middle platform is configured to analyze the indication information based on a water conservancy professional term intention recognition model, obtain different task requirements, call the large model technology base based on the different task requirements to make decisions, and obtain decision information. The large model technology base is integrated with a water conservancy large model trained based on a domain mechanism model library. The large model technology base comprises a dynamic knowledge distillation module, a digital twin base, and a feedback optimization module. The dynamic knowledge distillation module is configured to acquire water conservancy data, input the water conservancy data into each water science mechanism sub-model to acquire key parameters of flood evolution, encode the key parameters of flood evolution into a structured feature vector, fuse the spatial and temporal distribution features in the decision intermediate result with semantic representations of the water conservancy large model by using a cross-modal attention mechanism, and acquire time and space distribution features in the decision intermediate result. The dynamic knowledge distillation module is further configured to drive each water science mechanism sub-model to perform Monte Carlo simulation based on the time and space distribution features in the decision intermediate result, generate an enhanced training data set covering a rainstorm intensity-flooded area mapping relationship, input the enhanced training data set, update weight parameters of the water conservancy large model by using an ADMM algorithm, and obtain a trained water conservancy large model. The feedback optimization module is configured to acquire decision feedback data of the agent business middle platform in real time, dynamically adjust attention weights of the water conservancy large model according to the decision feedback data, and update and optimize the trained water conservancy large model.

2. The water affair decision system based on multi-modal model fusion according to claim 1, wherein The virtual digital human interaction module is further configured to receive decision information sent by the agent business middle platform, and render a VR scene based on the decision information. The agent business middle platform is configured to perform data matching on sub-scenes of various domain scenes expressed by the virtual digital human interaction module based on different task requirements when calling the large model technology base.

3. The water affair decision system based on multi-modal model fusion according to claim 1, characterized in that, The water science mechanism sub-models comprise at least two of a distributed hydrological model, a water dynamics model, a water quality model, a pipe network model, and a hydraulic model.

4. The water affair decision system based on multi-modal model fusion according to claim 1, characterized in that, 5. The water affair decision system based on multi-modal model fusion according to claim 1, wherein ​ The digital twin base is a virtual network model integrating satellite remote sensing data and ground observation data, wherein the virtual network model is a model of all elements and processes involved in water from precipitation in the sky to surface runoff, river flood evolution, reservoir water engineering scheduling, and underground pipe network in the target area, which is established by using a visual modeling tool.

6. The water affair decision system based on multi-modal model fusion according to claim 1, characterized in that, The intelligent agent service middle platform includes multiple scene intelligent agents and a task chain analysis engine. The task chain analysis engine decomposes the task demand into an atomic operation chain and calls different scene intelligent agents according to the atomic operation chain, so as to generate the decision information by calling the water conservancy large model trained on the large model technology base.

7. The water affair decision system based on multi-modal model fusion according to claim 1, characterized in that The water conservancy emergency scene emotion recognition model includes a multi-modal feature extraction layer and a scene-based emotion classification layer, wherein The multi-modal feature extraction layer extracts voice instructions and voice emotion spectrum features of the voice instructions by using a Mel spectrum convolution network, and extracts facial action unit intensity features by using a lightweight ResNrt-18 network; The scene-based emotion classification layer fuses the voice instructions, the voice emotion spectrum features, and the facial action unit intensity features by using a graph attention network, and outputs the instruction information.

8. The water affair decision system based on multi-modal model fusion according to claim 7, characterized in that, The water conservancy emergency scene emotion recognition model is trained by using an adversarial training mechanism.

9. The water affair decision system based on multi-modal model fusion according to claim 2, characterized in that, The virtual digital human interaction module includes: a multi-modal sensor array and a VR scene dynamic mapping unit; The multi-modal sensor array includes a high-definition camera, an infrared thermal imager, and an array microphone to collect facial expressions, body temperature changes, and voices of the user; The VR scene dynamic mapping unit is used to convert the decision information into a three-dimensional particle system for rendering.

Citation Information

Patent Citations

  • Flood control four-pre-platform application system based on digital twinning

    CN118052047A

  • Multi-modal large model construction method, system, equipment and medium applied to water conservancy field

    CN119721233A