Multi-mode intelligent monitoring method and device in water affair field
By acquiring entity status through video and combining it with large language models and dependency parsing, the problem of low monitoring accuracy caused by the diversity and complexity of data in water monitoring systems has been solved, enabling efficient identification and handling of water-related incidents.
Patent Information
- Application Number
- CN202511352386.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-01-16
AI Technical Summary
Existing water monitoring systems suffer from low accuracy in identifying water-related incidents due to the diversity and complexity of data sources and incident scenarios, which in turn affects the effectiveness of response.
By acquiring entity states through video and combining the matching relationship between entity states and water affairs events, a state recognition model based on a large language model and dependency parsing are used to extract entity states and relation triples of the target water surface. Spatial topological relationships, event triggering conditions, and handling rules are used to identify water affairs events and determine handling plans.
It has improved the accuracy of monitoring and identifying water-related incidents, enhanced the effectiveness of handling water-related incidents, and reduced the harm caused by water-related incidents.
Smart Images

Figure CN121354007A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent monitoring technology for water-related incidents, and in particular to a multimodal intelligent monitoring method and device for the water-related field. Background Technology
[0002] Rapid identification and handling of water-related incidents can reduce the harm caused by them.
[0003] However, due to the diverse and complex data sources and scenarios in which water incidents occur in existing water monitoring systems, the accuracy of water incident monitoring and identification is low, resulting in poor handling of water incidents. Summary of the Invention
[0004] This invention proposes a multimodal intelligent monitoring method and device for the water sector. It acquires the status of entities through video and identifies water events by combining the matching relationship between the entity status and water events. It can adapt to diverse and complex data sources and water event scenarios, thereby improving the accuracy of water event monitoring and identification, and thus improving the handling effect of water events and reducing the harm caused by water events.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] In a first aspect, the present invention provides a multimodal intelligent monitoring method for the water resources field, comprising: acquiring multiple frames of video footage of a target water surface over a period of time; determining the physical state of the target water surface based on the multiple frames of video footage; the physical state indicating the physical action information contained in the multiple frames of video footage; and further determining, based on the physical state of the target water surface and water resources event matching rules, water resources events occurring on the target water surface over a period of time and corresponding handling plans for the water resources events.
[0007] This invention provides a multimodal intelligent monitoring method for the water resources field. First, it extracts the entity state from multiple frames of video footage of a target water surface over a period of time. Then, it matches the entity state with water resources event matching rules to determine the water resources event occurring on the target water surface and the corresponding handling plan. This process obtains the entity state solely through video and combines the matching relationship between the entity state and the water resources event to achieve water resources event identification. It can adapt to diverse and complex data sources and water resources event scenarios, improving the accuracy of water resources event monitoring and identification, thereby enhancing the handling effect of water resources events and reducing the harm caused by them.
[0008] In one implementation of the first aspect, determining the entity state of the target water surface includes: extracting the entity state of the target water surface from the multi-frame video images using a state recognition model based on a large language model, based on multi-frame video images and entity state recognition prompts; the entity state recognition prompts are used to guide the state recognition model to extract the entity state of the target water surface.
[0009] In one implementation of the first aspect, the method further includes: verifying the consistency between the movement direction of entities in a water-related event and the water flow direction of the target water surface, and obtaining a verification result. When the verification result is inconsistent, the entity state recognition prompt is modified, and the process returns to the step of extracting the entity state of the target water surface from the multi-frame video footage using a state recognition model based on a large language model, based on the multi-frame video footage and the entity state recognition prompt.
[0010] In one implementation of the first aspect, the water event matching rules include spatial topology relationship matching rules, event triggering condition matching rules, and disposal rules. The spatial topology relationship matching rules indicate the matching relationship between a relation triple and a spatial topology relationship. A relation triple includes a subject, a predicate, and an object. The event triggering condition matching rules indicate the matching relationship between a spatial topology relationship and a water event. The disposal rules indicate the matching relationship between a water event and a disposal plan.
[0011] In one implementation of the first aspect, determining water-related events occurring on a target water surface within a certain period and the corresponding handling schemes for these events includes: using dependency parsing to extract relation triples of the target water surface from its entity state; matching these relation triples with spatial topology relation matching rules to obtain spatial topology relations in the entity state; matching these spatial topology relations with event triggering condition matching rules to obtain water-related events occurring on the target water surface within a certain period; and matching these water-related events with handling rules to obtain the handling schemes for these events.
[0012] In one implementation of the first aspect, acquiring multiple frames of video footage of the target water surface over a period of time includes: acquiring monitoring videos of the target water surface from a monitoring network of the target water surface; the monitoring network of the target water surface includes multiple drones and multiple fixed cameras. Keyframe extraction is then performed on the monitoring videos of the target water surface to obtain multiple frames of video footage of the target water surface over a period of time.
[0013] Secondly, this invention provides a multimodal intelligent monitoring device for the water resources field, including a video image acquisition module, an entity state determination module, and a water resources event determination module. The video image acquisition module is used to acquire multiple frames of video images of a target water surface over a period of time. The entity state determination module is used to determine the entity state of the target water surface based on the multiple video frames; the entity state indicates the entity action information contained in the multiple video frames. The water resources event determination module is used to determine water resources events occurring on the target water surface within a period of time and corresponding handling plans for those events, based on the entity state of the target water surface and water resources event matching rules.
[0014] In one implementation of the second aspect, the entity state determination module is specifically used to extract the entity state of the target water surface from the multi-frame video images based on the multi-frame video images and entity state recognition prompts, using a state recognition model based on a large language model; the entity state recognition prompts are used to guide the state recognition model to extract the entity state of the target water surface.
[0015] In one implementation of the second aspect, the water event determination module specifically uses dependency parsing to extract relation triples of the target water surface from the entity state of the target water surface. It then matches these relation triples with spatial topology relation matching rules to obtain the spatial topology relations in the entity state. Next, it matches these spatial topology relations with event triggering condition matching rules to obtain the water events that occurred on the target water surface within a certain period. Finally, it matches these water events with handling rules to obtain the handling plan for the water events.
[0016] Thirdly, the present invention provides an electronic device including a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the method described in the first aspect above or any implementation thereof.
[0017] Fourthly, the present invention provides a computer-readable storage medium including computer program instructions that, when executed by a computer, cause the computer to perform the method described in the first aspect above or any implementation thereof.
[0018] Fifthly, the present invention provides a computer program product, including computer program instructions, which, when executed on a computer, cause the computer to perform the method described in the first aspect above or any implementation thereof.
[0019] The technical effects corresponding to the second to fifth aspects and their possible implementations can be referred to the above description of the technical effects of the first aspect and its possible implementations, and will not be repeated here. Attached Figure Description
[0020] Figure 1 This is one of the schematic diagrams of a multimodal intelligent monitoring method in the water resources field provided in the embodiments of this application;
[0021] Figure 2 This is the second schematic diagram of a multimodal intelligent monitoring method in the water sector provided in this application embodiment;
[0022] Figure 3 This is the third schematic diagram of a multimodal intelligent monitoring method in the water sector provided in this application embodiment;
[0023] Figure 4 This is the fourth schematic diagram of a multimodal intelligent monitoring method in the water sector provided in this application embodiment;
[0024] Figure 5 This is the fifth schematic diagram of a multimodal intelligent monitoring method in the water sector provided in this application embodiment;
[0025] Figure 6 This is a schematic diagram of a multimodal intelligent monitoring system in the water sector provided in an embodiment of this application. Detailed Implementation
[0026] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0027] In the description of this invention, unless otherwise stated, "multiple frames" means two or more frames. For example, a multi-frame video frame refers to a video frame with two or more frames.
[0028] The method and apparatus provided in this application embodiment relate to a multimodal intelligent monitoring method and apparatus in the field of water affairs, which can be used to identify water affairs events on a target water surface and provide a solution for handling water affairs events.
[0029] To address the issue of low accuracy in monitoring and identifying water-related incidents due to the diverse and complex data sources and scenarios in existing water monitoring systems, which leads to ineffective handling of water-related incidents, this application provides a multimodal intelligent monitoring method and device for the water sector. This method acquires entity states through video and identifies water-related incidents by matching these states with the events they occur in. By adapting to diverse and complex data sources and scenarios, this method improves the accuracy of water-related incident monitoring and identification, thereby enhancing the effectiveness of incident handling and reducing the harm caused by water-related incidents.
[0030] For example, the multimodal intelligent monitoring method for the water resources field provided in this embodiment of the invention can be executed by an electronic device with processing capabilities, such as a computer or server. Taking a computer as an example, the hardware components of the computer may include: a processor, memory, a network interface, a user interface, a communication bus, etc.
[0031] The processor controls the electronic equipment to perform related processing and calculation tasks, such as acquiring multiple frames of video footage, determining the status of entities, and identifying water-related incidents and their handling plans. The processor may include a central processing unit (CPU) or other processors, and can be single-core or multi-core; for example, a processor may include multiple CPUs.
[0032] Memory is used to store computer instructions and related data, such as storing multiple frames of video footage, physical states, water-related incidents, and their handling plans. Memory can be random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical storage, magnetic disk storage media, or other magnetic storage devices, or any other medium capable of storing program code or data accessible by a computer. Optionally, memory can be integrated into the processor, or it can be independent of the processor.
[0033] A network interface is used for communication between a computer and other devices or communication networks. A network interface can be a transceiver with transmit and receive capabilities. Optionally, a network interface may include standard wired interfaces or wireless interfaces (such as Wi-Fi interfaces, Bluetooth interfaces, and 5G interfaces).
[0034] The communication bus is used to enable communication between different components. For example, the processor, memory, network interface and user interface mentioned above can be interconnected through the communication bus.
[0035] The user interface may include a display screen and an input unit (such as a keyboard). Optionally, the user interface may also include a standard wired interface or a wireless interface.
[0036] Those skilled in the art will understand that the computer described above may include more or fewer components, or combine certain components, or have different component arrangements; the embodiments of this application do not limit this.
[0037] To better understand the technical solutions of the embodiments of this application, the technical terms involved in the embodiments of this application will be explained below.
[0038] 1. Intelligent agent (or artificial intelligence agent)
[0039] Artificial intelligence (AI) is a technological system that simulates human intelligence through algorithms and data. The four core technologies of AI are perception, reasoning, learning, and action. An AI agent, as a specific application of AI, is an intelligent entity capable of perceiving its environment, making autonomous decisions, and executing actions. Its goal is to complete specific tasks through interaction with the outside world.
[0040] Specifically, an intelligent agent typically comprises four core modules: a perception module, a decision-making module (also known as a technology module), an action module, and a memory module. The perception module is used to acquire information through data input (such as APIs) and transform that information into a understandable format. The decision-making module (e.g., a large language model) performs logical reasoning, task planning, or strategy generation based on the information acquired by the perception module. The action module executes the decision results (e.g., generating text, calling APIs). The memory module (e.g., a database, a knowledge base) is used to store data, supporting long-term reasoning and contextual understanding.
[0041] It should be noted that, for the intelligent agent in the embodiments of this application, the perception module of the intelligent agent can be either a large language model or other deep learning models, and the embodiments of this application do not limit it.
[0042] 2. Prompt words
[0043] Prompts are instructions or questions that users input into large language models (such as GPT-4, Claude, LLaMA, etc.) to guide the large language model to output specific types of response results.
[0044] As the core commands for user-model interaction, prompts directly influence the output quality and direction of large language models. To ensure a satisfactory output, prompts need to include task instructions, contextual information, input data, output format, and examples. If the output of a large language model is unsatisfactory after prompts are input, adjustments to the prompts can improve its performance.
[0045] 3. Rule base
[0046] A rule base is a specific form of knowledge base; it's a structured collection used to store and manage business rules. These rules capture domain expertise, business strategies, regulatory constraints, and operational processes in an "if-then" logical form (i.e., IF-THEN statements). The rule base is a core component of the rule engine. The rule engine is responsible for interpreting and executing the rules in the rule base, thereby automating the processing of complex business logic.
[0047] like Figure 1 As shown in the embodiment of this application, a multimodal intelligent monitoring method in the water affairs field includes S101-S103.
[0048] S101. Acquire multiple frames of video footage of the target water surface over a period of time.
[0049] Optionally, combined Figure 1 ,like Figure 2 As shown, S101 includes S1011-S1012.
[0050] S1011. Obtain monitoring videos of the target water surface from the monitoring network of the target water surface.
[0051] The monitoring network for the aforementioned target water surface includes multiple drones and several fixed cameras.
[0052] S1012. Extract keyframes from the monitoring video of the target water surface to obtain multiple video frames of the target water surface over a period of time.
[0053] In one application scenario, a monitoring network consisting of drones and fixed cameras is used to monitor the target water surface. The video stream is parsed using the GB28181 network video protocol to form a visualized video image. The video stream is captured at a frequency of 5 minutes / base frequency and 1 second / burst frequency, and multiple key frames (i.e., multiple frames of video images of the target water surface) are visualized in JPEG format.
[0054] S102. Determine the physical state of the target water surface based on multiple frames of video footage.
[0055] The entity status indicator contains entity action information across multiple video frames. Entity action information includes the entity name, entity description, entity action, and the time the action occurred.
[0056] Specifically, in combination Figure 2 ,like Figure 3 As shown, S102 includes S1021.
[0057] S1021. Based on multiple video frames and entity state recognition prompts, an entity state recognition model based on a large language model is used to extract the entity state of the target water surface from the multiple video frames.
[0058] It should be noted that the aforementioned state recognition model based on a large language model refers to a state recognition model trained on a large language model (also known as a general semantic model). Therefore, the aforementioned state recognition model can be considered a type of large language model. Specifically, the aforementioned state recognition model is obtained by fine-tuning a pre-trained large language model using knowledge from this technical field. The aforementioned large language model is a transformer-based large language model, which can be the Tongyi Qianwen-Omni-Turbo model, the GPT series of large language models (GPT-3, GPT-4, and GPT-5, etc.), or the LLaMA series of large language models (LLaMA-1 and LLaMA-2, etc.). This application does not limit the specific type of the aforementioned large language model in its embodiments.
[0059] The aforementioned entity state recognition prompts are used to guide the state recognition model in extracting the entity state of the target water surface.
[0060] For example, the entity status recognition prompts are as follows.
[0061] You are a water management expert specializing in analyzing water-related images. Please carefully analyze the given image and, based on the context of water management and emergency operations, output the following:
[0062] 1. Key Geographical Features: List the main water-related objects, facilities, or locations in the image (e.g., water pipes, reservoirs, valves, pumping stations, rivers, dams, etc.). Describe their type and location.
[0063] 2. Water Event Time: Indicate the time information of the event, action, or state shown in the image (e.g., the time the event occurred, the timestamp of the operation, or the time inferred from image cues). If the time information is not obvious, infer it reasonably from the context or state "cannot be determined".
[0064] 3. Actions of key objects: Describe the actions or states that key objects in the image are performing (e.g., water flow, valve opening and closing, pump station operation, flooding, etc.). Include the direction, intensity, or impact of the action.
[0065] Please output in JSON format, ensuring the content is accurate and concise. An example of the output format is shown below.
[0066] {
[0067] "key_entities":["Entity 1: Type / Description","Entity 2: Type / Description",...],
[0068] "water_time": "Time description or inference",
[0069] "key_actions":["action1: description","action2: description",...]
[0070] }
[0071] In one application scenario, the aforementioned state recognition model is constructed based on the Qianwen model. Guided by the aforementioned entity state recognition prompts, the aforementioned state recognition model first identifies physical entities (such as "river / plastic bag / unwearing personnel"), then extracts action relationships (such as "floating in / going into the water"), and finally generates a structured semantic description.
[0072] It should be understood that, unlike traditional CV (Computer Vision) models, the Qianwen model utilizes its pre-trained world knowledge to transform pixels into natural language (such as "Black plastic bags are gathered in the river, and two people without life jackets are entering the water at the embankment"), achieving a breakthrough in zero-sample water scene transfer. This process simultaneously integrates infrared thermal images and visible light data to ensure the accuracy of semantic descriptions in complex scenes such as nighttime water surface reflections, providing a low-threshold solution for event decision-making without labeled data.
[0073] S103. Based on the physical state of the target water surface and the water affairs event matching rules, determine the water affairs events that occur on the target water surface within a certain period of time and the handling plan for the water affairs events.
[0074] The aforementioned water incident matching rules include spatial topology relationship matching rules, event triggering condition matching rules, and disposal rules. Spatial topology relationship matching rules indicate the matching relationship between relation triples and spatial topology relationships; relation triples include a subject, a predicate, and an object. Event triggering condition matching rules indicate the matching relationship between spatial topology relationships and water incidents. Disposal rules indicate the matching relationship between water incidents and disposal plans.
[0075] In one implementation, the aforementioned water event matching rules are implemented through a business rule base. The construction process of this business rule base is described below.
[0076] First, by combining water sector standards and regulations, industry experience, and regionally specific data (such as policy documents, management methods, and emergency response plans), the physical entities, subject-object relationships, and handling plans involved in water-related incidents are recorded in a standardized manner to form a rule system. This rule system is then transformed into a structured knowledge system of rules that enables machine-executable decision-making logic, resulting in the aforementioned business rule base.
[0077] The aforementioned business rule base includes three layers of rules: spatial topology layer, event logic layer, and decision layer, as detailed in Table 1 below.
[0078] Table 1. Schematic diagram of business rule base
[0079]
[0080] In Table 1 above, “∩” means “and”.
[0081] Specifically, the aforementioned spatial topology layer primarily extracts physical entities involved in the event, such as waterways, rivers, ice floes, algae, and the number of people not wearing police (security guards, firefighters) uniforms, as well as the location of the event, including the latitude and longitude of the drone footage and the contextual location of the physical entity in the image (center of the river, on the river, halfway through the river, left bank, etc.). The aforementioned event logic layer primarily establishes a mapping relationship between physical entities and water-related events over time, such as (personnel ∩ no life jacket ∩ river) → illegal swimming, (hidden pipe outlet ∩ abnormal water flow ∩ sudden change in water quality) → suspected illegal discharge, (vessel stationary ∩ trawl equipment ∩ coordinates of prohibited fishing area) → illegal fishing, etc. The aforementioned decision-making layer primarily establishes a mapping relationship between water-related events and response plans, such as illegal swimming → drone warning, suspected illegal discharge → joint on-site inspection, illegal fishing → evidence collection and punishment, etc.
[0082] Optionally, the spatial topological relationships contained in the above spatial topological layer can be represented using structured data. For example, I(a), B(a), and E(a) represent the interior, boundary, and exterior of a, respectively; I(b), B(b), and E(b) represent the interior, boundary, and exterior of b, respectively. Then, the intersection matrix of a and b is shown in Table 2 below. Here, a and b represent entities. Taking the UAV inspection of river ice floes as an example, the river channel can be vectorized into a rectangular surface in a certain frame of image, and the ice floes can be vectorized into polygonal surfaces.
[0083] Table 2. Intersection Matrix Diagram of Spatial Topological Relationships
[0084]
[0085] In Table 2 above, dim(·) refers to the command to obtain the dimension. The value range of dim(·) is {-1, 0, 1, 2, T, F}. 0 indicates that there is an intersection, intersecting at a point; 1 indicates that there is an intersection at a line; 2 indicates that there is an intersection at a surface; T indicates that there is an intersection, intersecting at a point, a line, and a surface respectively; -1 and F indicate that there is no intersection, i.e., an empty set.
[0086] Taking dim(I(a)I(b)) as an example, dim(I(a)I(b)) = 0 indicates that the interiors of a and b intersect at a point; dim(I(a)I(b)) = 1 indicates that the interiors of a and b intersect at a line; dim(I(a)I(b)) = 2 indicates that the interiors of a and b intersect at a surface; dim(I(a)I(b)) = T indicates that the interiors of a and b intersect at a point, a line, and a surface respectively; dim(I(a)I(b)) = -1 or dim(I(a)I(b)) = F indicates that the interiors of a and b do not intersect, i.e., the empty set. If the ice crystals are all within the river channel, they can be represented in a structured manner using the matrix “T*****FF*” (where “*” is a wildcard, indicating any result is acceptable).
[0087] For example, in combination Figure 3 ,like Figure 4 As shown, S103 includes S1031-S1034.
[0088] S1031. Using dependency parsing, extract the relation triples of the target water surface from the entity state of the target water surface.
[0089] As a natural language processing (NLP) task, dependency parsing can be performed using NLP tools (such as Stanford CoreNLP, SpaCy, HanLP, etc.) to analyze sentences and generate <subject, predicate, object> triples (corresponding to the relation triples of the target water surface mentioned above). Since dependency parsing is a commonly used technique in this field, the extraction process of the relation triples of the target water surface will not be elaborated in this embodiment.
[0090] S1032. Match the relation triples with the spatial topology relation matching rules to obtain the spatial topology relation in the entity state.
[0091] S1033. Match the spatial topology relationship with the event triggering condition matching rules to obtain the water affairs events that occur on the target water surface within a certain period of time.
[0092] S1034. Match water incidents with handling rules to obtain water incident handling solutions.
[0093] In some embodiments, a water affairs rule engine can be used to match the aforementioned relation triples with the spatial topology relation matching rules in the spatial topology layer of the aforementioned business rule base to obtain the spatial topology relation in the entity state; and the water affairs rule engine can be used to match the aforementioned spatial topology relation with the event triggering condition matching rules in the event logic layer of the aforementioned business rule base to obtain the water affairs events that occur on the target water surface within a certain period of time; and the water affairs rule engine can then be used to match the aforementioned water affairs events that occur on the target water surface within a certain period of time with the handling rules in the decision layer of the aforementioned business rule base to obtain the handling plan for the water affairs events.
[0094] In one application scenario, based on the semantic description (i.e., the entity state of the target water surface) output by the state recognition model, the system initiates a water affairs rule engine for deep analysis: first, it extracts <subject, predicate, object> triples (e.g., <multiple people without life jackets, floating in, river channel>) through dependency parsing, and then matches them with the business rule base. At the spatial topology layer, it verifies geographical relationships (e.g., whether the people without life jackets belong to the river channel); at the event logic layer, it matches combined conditions (people without life jackets + floating action + river area triggering a public safety incident of illegal water entry); and at the decision-making layer, it associates emergency response plans (drone warnings, emergency rescue, etc.). This process innovatively introduces dynamic confidence calibration. When the visual description contains ambiguous semantics such as "floating" or "struggling," it automatically retrieves historical work order data for weighted verification, enabling the large language model and water affairs-specific knowledge to form a closed-loop collaboration.
[0095] In one implementation, combined with Figure 4 ,like Figure 5 As shown, the above method also includes S104.
[0096] S104. Verify the consistency between the movement direction of entities in the water event and the water flow direction of the target water surface, and obtain the verification result; when the verification result is inconsistent, modify the entity state recognition prompt word and return to S1021; when the verification result is that the water event is consistent, end the execution.
[0097] In one embodiment, the water-related event output in S1033 above forms an event trigger judgment. After the event is triggered, multi-source cross-validation (corresponding to the consistency verification mentioned above) is performed: three consecutive frames of images are extracted for identification based on the time of the keyframe images and the continuous time interval (configurable) (to prevent instantaneous misjudgment). The consistency between the movement direction of entities in the event and the water flow vector is recorded using UAV images. It is necessary to match the spatial location of the keyframe with the GIS data of the river to obtain the water flow direction and spatial location of the area. For abnormal ice floe scenarios, the freezing and floating time of ice floes in the river are identified in multiple frames, and the spatial location information is obtained. For the weather temperature data of the time, the confidence of the time is judged to reduce the interference of abnormal water surface reflection and floating white garbage in the video image. If the final verification is successful, a structured event report is generated, which includes the GIS coordinates of the entity, the confidence level, and the disposal suggestions. The system establishes a continuous evolution mechanism. Misjudgment cases reported by humans will be transformed into business rule base parameter optimization (such as adjusting the "blue algae" color threshold) and entity status recognition prompt words enhancement (adding professional expressions such as "piping and seepage").
[0098] In summary, the multimodal intelligent monitoring method for the water resources field provided in this application first extracts the entity state from multiple frames of video footage of the target water surface over a period of time. Then, it matches the entity state with water resources event matching rules to determine the water resources event occurring on the target water surface and the corresponding handling plan. This process obtains the entity state solely through video and combines the matching relationship between the entity state and the water resources event to identify the event. It can adapt to diverse and complex data sources and various scenarios in which water resources events occur, improving the accuracy of water resources event monitoring and identification, thereby enhancing the handling effect of water resources events and reducing the harm caused by them.
[0099] Accordingly, embodiments of this application provide a multimodal intelligent monitoring device for the water resources field, such as... Figure 6 As shown, it includes a video image acquisition module 501, an entity status determination module 502, and a water affairs event determination module 503.
[0100] The video image acquisition module 501 is used to acquire multiple frames of video images of the target water surface over a period of time. For example, the video image acquisition module 501 is used to implement step S101 of the above method.
[0101] The entity state determination module 502 is used to determine the entity state of the target water surface based on multiple video frames; the entity state indicates the entity action information contained in the multiple video frames. For example, the entity state determination module 502 is used to implement S102 of the above method.
[0102] The water event determination module 503 is used to determine, based on the entity status of the target water surface and the water event matching rules, the water events occurring on the target water surface within a certain period of time and the corresponding handling plans. For example, the water event determination module 503 is used to implement S103 of the above method.
[0103] Optionally, the entity state determination module 502 is specifically used to: extract the entity state of the target water surface from the multi-frame video images using a state recognition model based on a large language model, based on the multi-frame video images and entity state recognition prompts. For example, the entity state determination module 502 is specifically used to implement S1021 of the above method.
[0104] Optionally, the water event determination module 503 is specifically used for: using dependency parsing to extract relation triples of the target water surface from the entity state of the target water surface; matching the relation triples with spatial topology relation matching rules to obtain spatial topology relations in the entity state; matching the spatial topology relations with event triggering condition matching rules to obtain water events occurring on the target water surface within a certain period of time; and matching the water events with handling rules to obtain handling plans for the water events. For example, the water event determination module 503 is specifically used to implement S1031-S1034 of the above method.
[0105] The modules of the aforementioned multimodal intelligent monitoring device in the water sector can also be used to execute other steps in the above method embodiments. All relevant content involved in the above method embodiments can be referenced from the functional descriptions of the corresponding functional modules, and will not be repeated here.
[0106] This application also provides an electronic device, including: a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the methods in the above embodiments. The processor can implement the video image acquisition module 501, the entity state determination module 502, and the water event determination module 503; the memory can also be used to store multiple frames of video images, entity states, water events, and their handling plans, etc.
[0107] This application also provides a computer-readable storage medium including a computer program that, when run on a computer, performs the methods described in the above embodiments.
[0108] This application also provides a computer program product, which includes computer program instructions that, when run on a computer, execute the methods described in the above embodiments.
[0109] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A multi-modal intelligent monitoring method in the field of water affairs, characterized by, The method comprises the following steps: acquiring a plurality of frames of video pictures of a target water surface in a period of time; determining an entity state of the target water surface based on the plurality of frames of video pictures; the entity state indicates entity action information contained in the plurality of frames of video pictures; determining a water affair event occurring in the target water surface in the period of time and a disposal scheme of the water affair event according to the entity state of the target water surface and a water affair event matching rule.
2. The method of claim 1, wherein, The determination of the entity state of the target water surface comprises the following steps: extracting the entity state of the target water surface from the plurality of frames of video pictures by using a state recognition model based on a large language model according to the plurality of frames of video pictures and an entity state recognition prompt; the entity state recognition prompt is used to guide the state recognition model to extract the entity state of the target water surface.
3. The method of claim 2, wherein, The method further comprises the following steps: verifying consistency between a moving direction of an entity in the water affair event and a water flow direction of the target water surface to obtain a verification result; when the verification result is inconsistent, modifying the entity state recognition prompt, and returning to the step of extracting the entity state of the target water surface from the plurality of frames of video pictures by using the state recognition model based on the large language model according to the plurality of frames of video pictures and the entity state recognition prompt.
4. The method of claim 1, wherein the water affair event matching rule comprises a spatial topological relation matching rule, an event trigger condition matching rule, and a disposal rule; the spatial topological relation matching rule indicates a matching relationship between a relation triple and a spatial topological relation; the relation triple comprises a subject, a predicate, and an object; the event trigger condition matching rule indicates a matching relationship between a spatial topological relation and a water affair event; the disposal rule indicates a matching relationship between a water affair event and a disposal scheme.
5. The method of claim 4, wherein, The determination of the water affair event occurring in the target water surface in the period of time and the disposal scheme of the water affair event comprises the following steps: extracting a relation triple of the target water surface from the entity state of the target water surface by using a dependency syntax analysis; matching the relation triple with the spatial topological relation matching rule to obtain a spatial topological relation in the entity state; matching the spatial topological relation with the event trigger condition matching rule to obtain the water affair event occurring in the target water surface in the period of time; matching the water affair event with the disposal rule to obtain the disposal scheme of the water affair event.
6. The method of claim 1, wherein, The acquisition of the plurality of frames of video pictures of the target water surface in the period of time comprises the following steps: acquiring a monitoring video of the target water surface from a monitoring network of the target water surface; the monitoring network of the target water surface comprises a plurality of unmanned aerial vehicles and a plurality of fixed cameras; performing key frame extraction on the monitoring video of the target water surface to obtain the plurality of frames of video pictures of the target water surface in the period of time.
7. A multi-modal intelligent monitoring device in the field of water affairs, characterized by The system comprises a video picture acquisition module, an entity state determination module, and a water affair determination module; the video picture acquisition module is configured to acquire a plurality of frames of video pictures of a target water surface in a period of time; The entity state determination module is configured to determine an entity state of the target water surface based on the multiple frames of video pictures; the entity state indicates entity action information contained in the multiple frames of video pictures; The water affair event determination module is configured to determine a water affair event occurring on the target water surface in the period of time and a disposal scheme of the water affair event according to the entity state of the target water surface and a water affair event matching rule.
8. An electronic device, comprising: An electronic device includes a processor and a memory coupled to the processor; the memory is configured to store computer instructions; when the electronic device is running, the processor executes the computer instructions stored in the memory, so that the electronic device executes the method in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, A computer program product includes computer program instructions; when the computer program instructions are executed by a computer, the computer executes the method in any one of claims 1 to 6.
10. A computer program product, characterised in that, A computer program product includes computer program instructions; when the computer program instructions are executed by a computer, the computer executes the method in any one of claims 1 to 6.
Citation Information
Patent Citations
River and lake supervision method and system based on unmanned aerial vehicle
CN115690628A
Water event risk prediction method and device, computer equipment and storage medium
CN118821928A
Flood intelligent scheduling decision-making method and device based on scheduling service knowledge graph
CN119904124A
Oilfield station unattended intelligent operation and maintenance centralized control method and system based on intelligent agent
CN120070090A
Urban inland inundation monitoring method and device, electronic equipment and storage medium
CN120088943A