Highway flood damage intelligent identification method and system based on multi-modal large model
By using an intelligent identification method based on a multimodal large model, combined with a structured knowledge base and spatiotemporal causal reasoning, the problem of accuracy and efficiency in identifying potential road flood damage hazards has been solved. This has enabled precise identification and risk warning of flood damage disasters, reducing reliance on manual labor and safety risks.
Patent Information
- Application Number
- CN202511130118.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-25
AI Technical Summary
Existing methods for identifying potential road damage from flooding suffer from low identification efficiency, heavy reliance on manual labor, poor identification accuracy, and an inability to conduct comprehensive causal analysis, making it difficult to meet the needs for precise and intelligent identification and response.
We employ a multimodal large model-based intelligent recognition method, combining a structured knowledge base with a spatiotemporal-causal reasoning mechanism. By collecting multimodal data (images, latitude and longitude, and shooting time), we utilize a large language model and image recognition engine for intelligent recognition, constructing an intelligent recognition system that includes data collection, knowledge retrieval, prompt word templates, and weather/map API linkage.
It enables accurate identification of potential road flood damage and intelligent analysis of delayed disaster causes, improving identification accuracy and adaptability, reducing labor costs and risks, and supporting efficient response by traffic management departments.
Smart Images

Figure CN121010907A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and highway flood damage identification technology, and in particular to a method and system for intelligent identification of highway flood damage based on a multimodal large model. Background Technology
[0002] Highway flood damage refers to the severe damage to highway subgrade, pavement, bridges, drainage systems, and ancillary facilities caused by natural phenomena such as torrential rain, floods, and freeze-thaw cycles. Typical manifestations include subgrade erosion, bridge slope damage, pavement siltation, and failure of protective works. Severe flood damage can even cause traffic disruptions or paralysis of regional road networks, making it one of the most destructive types of disasters.
[0003] Highway flood damage is the result of the coupling effect of multiple factors, including natural forces and engineering defects. Its causative mechanism exhibits a typical chain-like evolution: under continuous hydrological processes, roadbed slopes first undergo a gradual cumulative damage stage, manifested as surface erosion, crack initiation, and soil weakening. As hydrodynamic conditions exceed critical thresholds, sudden destructive events such as slope instability and roadbed collapse are induced. This disaster evolution pattern, from quantitative to qualitative change, dictates that risk management of flood-prone road sections must construct a full-chain prevention and control system encompassing "monitoring-early warning-response." Therefore, conducting systematic hazard investigation is a crucial link in breaking the chain of disaster transmission. By establishing dynamic risk files, implementing precise monitoring and early warning, and strengthening preventative maintenance and treatment, the risk of major casualties and property losses can be effectively reduced, thereby enhancing the resilience of highway infrastructure and shifting the disaster prevention and control focus forward.
[0004] Currently, the field of highway flood damage prevention and control is undergoing a technological paradigm shift, but the integration of traditional inspection models with emerging monitoring technologies still faces multiple challenges. The manual inspection system, long relied upon by road administration, public security, and traffic management departments, is essentially an experience-driven, subjective judgment model with four core drawbacks: first, inspection efficiency is limited by personnel experience; second, the average daily inspection mileage is less than 80 kilometers, more than 60% lower than mechanized inspections; third, the personal safety risk to inspection personnel increases significantly under extreme weather conditions; and fourth, traditional evidence collection methods are insufficient to comprehensively record the characteristics of flood-damaged road sections. With the empowerment of artificial intelligence, technologies such as satellite remote sensing monitoring, lidar scanning, and video AI detection are becoming increasingly mature. However, limited by high costs and signal obstruction, the contradiction between this "high-precision" technological equipment and the need for "lightweight" applications has become a core bottleneck restricting the large-scale promotion of monitoring and early warning systems.
[0005] With the development of general artificial intelligence models, the potential of Large Language Models (LLM) in the field of traffic management is gradually emerging. It possesses capabilities such as cross-modal understanding, complex reasoning, and multi-turn question answering, providing a new technological path for intelligent traffic management and decision support. National and local traffic management departments have also successively introduced policies to encourage the introduction of intelligent equipment (such as drones) and AI recognition algorithms into highway inspections to improve inspection efficiency and safety.
[0006] While existing research has introduced large language models into the transportation field, their application in identifying and investigating highway flood damage remains a gap. Current LLMs lack native image-based structural disaster recognition capabilities and suffer from issues such as "illusion" and "instability" when processing visual information. Therefore, how to build a structured highway flood damage knowledge base and integrate techniques such as prompting engineering, knowledge retrieval enhancement, and localized large model deployment to construct an intelligent flood damage identification system with professional understanding, knowledge-driven reasoning, and practical output capabilities has become a core challenge urgently needing to be overcome in the industry.
[0007] In summary, existing methods for identifying potential road flood damage hazards suffer from problems such as low identification efficiency, heavy reliance on manual labor, poor identification accuracy, and lack of comprehensive causal analysis capabilities, making it difficult to meet the demand for "precise and intelligent" identification and response to potential road flood damage hazards.
[0008] On the one hand, traditional manual inspections are time-consuming, labor-intensive, and pose high safety risks, and are difficult to cover high-risk structural areas. On the other hand, image-based small-model recognition systems have limitations such as limited recognition categories and the inability to perform structural reasoning and historical comparison, leading to the easy omission or misjudgment of potential road flood damage hazards. In addition, existing systems lack a comprehensive analysis mechanism for the relationship between shooting time, geographical location, and weather, making it difficult to identify delayed or hidden causes of flood damage and limiting the improvement of flood damage risk early warning capabilities. Summary of the Invention
[0009] The purpose of this invention is to propose an intelligent identification method and system for highway flood damage based on a multimodal large model, in order to solve the problems existing in the prior art. This invention proposes an intelligent identification method that integrates a multimodal large model, a structured knowledge base, and a spatiotemporal-causal reasoning mechanism to improve the accuracy, adaptability, and practical efficiency of identifying highway flood damage hazards. It enables efficient classification and accurate identification of various types of flood damage events, meeting the practical needs of highway on-site inspection scenarios for "high-precision and intelligent" inspection capabilities.
[0010] To achieve the above objectives, the present invention provides the following solution:
[0011] A method for intelligent identification of highway flood damage based on a multimodal large model includes:
[0012] Collect multimodal data from the highway site; wherein, the multimodal data includes: highway site images, latitude and longitude coordinates, and shooting time;
[0013] Using a pre-built structured knowledge base of highway flood damage, knowledge retrieval and reasoning are performed on the multimodal data;
[0014] Obtain the prompt word template that matches the recognition task;
[0015] Based on the multimodal data, knowledge retrieval and reasoning results, and prompt word templates, the input content of the large language model is constructed;
[0016] Based on the input content, a large language model and image recognition engine are invoked, and combined with latitude and longitude and shooting time, to intelligently identify road damage caused by floods.
[0017] Optionally, the structured highway flood damage knowledge base includes structural features, image representations, causal paths, and risk levels for several types of flood damage.
[0018] Optionally, the types of water damage include: roadbed water damage, pavement water damage, slope water damage, bridge water damage, tunnel water damage, retaining wall water damage, drainage facility water damage, and traffic safety facility water damage.
[0019] Optionally, obtaining the prompt word templates that match the recognition task includes:
[0020] A prompt word library is built based on preset prompt words;
[0021] Call up the prompt word template that matches the recognition task from the prompt word library.
[0022] Optionally, the prompt words include:
[0023] Intelligent agent prompts are used to define the recognition behavior style, inference boundaries, and output specifications of large language models;
[0024] Image recognition prompts are used to guide the model to identify structural states and flood damage features in images;
[0025] Structured output prompts are used to standardize the format of output fields, including type, cause, and risk.
[0026] Results review and causal reasoning prompts are used to combine time, space, and weather information to complete disaster cause analysis and confidence verification.
[0027] Optionally, intelligent identification of road damage caused by floods includes:
[0028] By utilizing large language models and image recognition engines, the system can identify the types of water damage hazards in images, analyze their causes, and infer their risk levels.
[0029] By combining the shooting time and latitude and longitude information, and by calling weather API and map information services, it is possible to determine whether there are any delayed causes of water damage.
[0030] By combining the identification of water damage hazard types, causal analysis and risk level reasoning, as well as whether there are delayed causes of water damage, a structured identification result is output.
[0031] Optionally, the structured recognition results include: water damage type, image features, structural location, disaster mechanism, risk level, associated weather, recognition confidence, and matching cases.
[0032] A multimodal large model-based intelligent identification system for highway flood damage, used to implement the multimodal large model-based intelligent identification method for highway flood damage as described above, the system comprising:
[0033] The data acquisition and input module is used to acquire multimodal data from the highway site; wherein, the multimodal data includes: highway site images, latitude and longitude coordinates, and shooting time;
[0034] The multimodal recognition model execution module is used to call the locally deployed large language model and image recognition model, and combine latitude and longitude and shooting time to intelligently identify road flood damage;
[0035] The knowledge base management module is used to perform knowledge retrieval and reasoning on the multimodal data using a pre-built structured knowledge base of highway flood damage.
[0036] The prompt word scheduling and dialogue control module is used to obtain prompt word templates that match the recognition task;
[0037] The weather / map API linkage module automatically calls external weather services and map information interfaces based on image metadata (capture date, time, latitude and longitude) to obtain historical weather records for the corresponding time period at that location. This module can assist models in determining whether flood damage is related to previous weather conditions and other factors.
[0038] The front-end display and structured result output module is used to visualize the recognition results. The structured results can be synchronized to the operation and maintenance management platform or traffic safety early warning system to improve the linkage response capability.
[0039] The beneficial effects of this invention are as follows:
[0040] This invention provides a method and system for intelligent identification of highway flood damage based on a multimodal large model. It overcomes the limitations of existing flood damage investigation methods from three levels: technical architecture, information processing logic, and identification capabilities. This enables accurate identification of structural disasters and intelligent analysis of delayed disaster causes.
[0041] This invention constructs a knowledge base of eight typical water damage structures and combines a large language model and an image recognition model for multi-round reasoning, thereby achieving accurate classification, feature analysis, and causal explanation of various types of water damage in complex road environments.
[0042] By linking weather APIs and map information, causal reasoning can be performed on rainfall, strong winds, and geological conditions for the next 1-7 days. Combining image content with a knowledge base, it can determine whether flood damage is related to "abnormal weather." This addresses the problem of existing models being unable to identify intangible risks (such as internal erosion or drainage blockage), supporting traffic management departments in proactive handling and tiered responses.
[0043] This invention supports immediate recognition upon front-end image input, combined with an automatic prompt word scheduling mechanism and structured result output. It also avoids manual on-site work in high-risk areas (such as steep slopes and bridge piers), significantly reducing personal risk and labor costs. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is a schematic diagram of a method for intelligent identification of road water damage based on a multimodal large model according to an embodiment of the present invention;
[0046] Figure 2 Images taken by a drone in an embodiment of the present invention;
[0047] Figure 3 This is an example of the recognition result output in an embodiment of the present invention;
[0048] Figure 4 This is a schematic diagram of a highway flood damage intelligent identification system based on a multimodal large model, according to an embodiment of the present invention. Detailed Implementation
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0051] like Figure 1 As shown in the figure, this embodiment proposes a method for intelligent identification of highway water damage based on a multimodal large model, including:
[0052] Collect multimodal data from the highway site; wherein, the multimodal data includes: highway site images, latitude and longitude coordinates, and shooting time;
[0053] Using a pre-built structured knowledge base of highway flood damage, data knowledge retrieval and reasoning are performed on the multimodal data;
[0054] Obtain the prompt word template that matches the recognition task;
[0055] Based on the multimodal data, knowledge retrieval and reasoning results, and prompt word templates, the input content of the large language model is constructed;
[0056] Based on the input content, a large language model and image recognition engine are invoked, and combined with latitude and longitude and shooting time, to intelligently identify road damage caused by floods.
[0057] Specifically, in this embodiment, the proposed intelligent identification method for highway flood damage based on a multimodal large model includes: acquiring multimodal input information containing highway scene images, latitude and longitude, and shooting time; constructing a structured highway flood damage knowledge base, integrating information such as flood damage definitions, classification systems, typical cases, structural features, risk levels, and causal mechanisms to support knowledge retrieval and reasoning during the identification process; setting intelligent agent prompt words and designing multiple types of task prompt word templates, including image recognition instructions, structured output templates, causal analysis strategies, and review dialogue control strategies, with prompt words categorized and stored in a prompt word library and dynamically invoked during the identification process; matching task prompt word templates using a prompt word scheduling module based on image input information, and providing prompts related to engineering and knowledge. The Retrieval Enhancement (RAG) mechanism jointly constructs the input content of a large model; it calls upon the deployed large language model and image recognition engine, combining structural information and prompt word logic from the knowledge base, to perform multi-round reasoning and recognition on the input image, determining whether water damage exists in the image, and analyzing its type, location, structural damage form, and causes; it links with weather and map service interfaces, querying weather records such as rainfall, strong winds, and blizzards for the target location within the next 1-7 days based on the image's shooting time and latitude and longitude information, and combining this with structural conditions to determine whether there are delayed disaster triggers; it outputs structured results, including water damage type, image features, structural location, disaster mechanism, risk level, associated weather, and recognition confidence, and uses the recognition results to assist in risk management and report generation. This embodiment proposes a "highway water damage intelligent recognition method" that combines a large language model, multimodal perception, knowledge base embedding, and spatiotemporal causal reasoning, improving the accuracy and timeliness of highway water damage identification, while possessing high practicality and promotional value, adapting to complex road conditions, reducing personnel input, assisting decision-making and early warning, and reducing the risk of traffic accidents caused by water damage.
[0058] Furthermore, the structured highway flood damage knowledge base includes structural features, image representations, causal paths, and risk levels for several types of flood damage.
[0059] Specifically, in this embodiment, the knowledge base construction includes:
[0060] We collected and organized standards, specifications, professional books, references, and engineering data to extract the structural features, image representations, causal paths, and risk levels of eight types of water damage, including roadbeds, pavements, bridges, tunnels, slopes, retaining walls, drainage facilities, and traffic safety facilities. We then constructed a multi-level knowledge base structure with semantic embedding capabilities for large models to use for semantic retrieval and comparison.
[0061] Furthermore, obtaining the prompt word templates that match the recognition task includes:
[0062] A prompt word library is built based on preset prompt words;
[0063] Call up the prompt word template that matches the recognition task from the prompt word library.
[0064] Specifically, in this embodiment, the prompt word design includes the following four categories:
[0065] Agent prompts are used to define the recognition behavior style, inference boundaries, and output specifications of large language models;
[0066] Image recognition prompts are used to guide the model in recognizing structural states and water damage features in images;
[0067] Structured output prompts are used to standardize the format of output fields, including type, cause, risk, etc.
[0068] Results review and causal reasoning prompts are used to combine time, space, and weather information to complete disaster cause analysis and confidence verification.
[0069] All prompts are standardized and templated before deployment and stored in a prompt word library, which users can freely call upon based on image information during the recognition process.
[0070] Furthermore, intelligent identification of highway flood damage includes:
[0071] By utilizing large language models and image recognition engines, the system can identify the types of water damage hazards in images, analyze their causes, and infer their risk levels.
[0072] By combining the shooting time and latitude and longitude information, and by calling weather API and map information services, it is possible to determine whether there are any delayed causes of water damage.
[0073] By combining the identification of water damage hazard types, causal analysis, and risk level reasoning, as well as the existence of delayed water damage causes, a structured identification result is output. An example of the output result is shown below. Figure 3 As shown.
[0074] Specifically, in this embodiment, by parsing the timestamp and latitude and longitude coordinates embedded in the image, the weather query interface is called to obtain the cumulative rainfall and extreme weather events in the next 1-7 days, and the presence of typical delayed disaster causes is determined by combining the water damage type and structural characteristics.
[0075] In this embodiment, the final output includes: hazard type, sub-type, image features, structural analysis, disaster mechanism, risk level, associated weather, remediation suggestions, identification confidence level, and matching cases. The output can be used for report generation, early warning distribution, or integration with traffic maintenance platforms, supporting engineering deployment and frontline practical applications.
[0076] This embodiment also proposes a highway flood damage intelligent identification system based on a multimodal large model, including:
[0077] The data acquisition module is used to collect multimodal data from the highway site; wherein, the multimodal data includes: highway site images, latitude and longitude, and shooting time;
[0078] The knowledge base retrieval module is used to perform data knowledge retrieval and reasoning on the multimodal data using a pre-built structured knowledge base of highway flood damage.
[0079] The prompt word acquisition module is used to acquire prompt word templates that match the recognition task;
[0080] The input content construction module is used to construct the input content of the large language model based on the multimodal data, knowledge retrieval and reasoning results, and prompt word templates.
[0081] The intelligent recognition module is used to intelligently identify road damage caused by floods by calling a large language model and an image recognition engine based on the input content, combined with latitude and longitude and shooting time.
[0082] Specifically, such as Figure 4 As shown, in this embodiment, data acquisition uses a data acquisition and input module: acquiring highway patrol images and completing input processing in a unified format. The images can be acquired by drones and contain metadata such as timestamps, latitude and longitude information, shooting angle, and device identification.
[0083] The knowledge base retrieval utilizes a knowledge base management module: this module constructs, updates, and retrieves the "Highway Flood Damage Disaster Knowledge Base." This knowledge base covers information such as definitions, classifications, structural characteristics, causative mechanisms, and historical cases of various flood damage disasters. This module supports semantic retrieval based on text embedding, allowing the recognition module to access relevant knowledge and enhance its image-text linkage recognition and reasoning capabilities.
[0084] The prompt word acquisition utilizes a prompt word scheduling and dialogue control module: it invokes preset prompt words and controls the interaction strategy between the prompt word and the recognition model. This module comprises two levels: system-level prompt words and task-level prompt words. System-level prompt words are used to define the agent's behavioral style and recognition boundaries, while task-level prompt words are stored according to water damage type, recognition purpose, and output target.
[0085] The intelligent recognition module employs a multimodal recognition model: it calls upon locally deployed large language and image recognition models to process the input images. This module integrates visual semantic understanding, image classification, and multi-turn dialogue capabilities, enabling it to identify potential road flood damage hazards in images, determine their specific types (such as roadbed flood damage, slope flood damage, and pavement flood damage), and generate preliminary structured description results.
[0086] The intelligent recognition also employs a weather / map API linkage module: based on image metadata (capture date, time, latitude and longitude), it automatically calls external weather services and map information interfaces to obtain historical weather records for the corresponding time period at that location. This module can assist the model in determining whether flood damage is related to previous weather conditions and other factors.
[0087] This device also includes a front-end display and structured result output module: visualizing the recognition results. The structured results can be synchronized to the operation and maintenance management platform or the traffic safety early warning system to improve the linkage response capability.
[0088] This invention provides a method and system for intelligent identification of highway flood damage based on a multimodal large model. It overcomes the limitations of existing flood damage investigation methods from three levels: technical architecture, information processing logic, and identification capability. It achieves accurate identification of structural disasters and intelligent analysis of delayed disaster causes, and has the following significant technical effects:
[0089] Significantly improved recognition accuracy and adaptability: This invention constructs a knowledge base of eight typical water damage structures and combines a large language model and an image recognition model for multi-round reasoning, thereby achieving accurate classification, feature analysis, and causal explanation of multiple types of water damage in complex road environments.
[0090] Enhanced risk warning capabilities: By linking weather APIs and map information, this system can perform causal reasoning on recent rainfall, strong winds, and geological conditions, and combine image content with a knowledge base to determine whether flood damage is related to "abnormal weather." This addresses the problem of existing models being unable to identify intangible risks (such as internal erosion or drainage blockage), supporting traffic management departments in proactive handling and tiered responses.
[0091] Improved recognition efficiency and significantly reduced reliance on manual labor: The system supports immediate recognition upon front-end image input, combined with an automatic prompt word scheduling mechanism and structured result output. It also avoids manual on-site work in high-risk areas (such as steep slopes and bridge piers), significantly reducing personal risk and labor costs.
[0092] It has strong system portability and deployment flexibility: The method of this invention can be adapted to various deployment forms: it supports online calling of large models and local deployment in traffic operation and maintenance units; the knowledge base and prompt words can be customized and updated, which facilitates its application in different regional highway networks.
[0093] This invention is applicable to practical work scenarios such as routine highway patrols, maintenance response, and disaster early warning. After patrol personnel collect images using drones or handheld cameras, they can identify, classify, deduce the causes of potential water damage hazards on roads, and conduct risk assessments based on a multimodal recognition model and a knowledge-based question-and-answer mechanism. The images collected by the drone include, for example, images from drones. Figure 2 As shown.
[0094] The system supports input including multimodal information such as images, time, latitude and longitude, and combines a knowledge base retrieval module and a prompt word scheduling module to automate the entire process of intelligent recognition and decision support. Its core process can be represented as follows:
[0095] D = R(Q, K)
[0096] A = MLLM(P(Q,D),D)
[0097] in:
[0098] Q represents the input image and its metadata (such as time, latitude and longitude, etc.);
[0099] K represents a pre-built knowledge base of water damage, which includes water damage types, structural features, etc.
[0100] R(Q, K) represents the retrieval of related content from knowledge base K based on input information Q, resulting in a set of candidate knowledge items D;
[0101] P(Q, D) represents combining the input Q with the search result D and filling the prompt word template to form a structured prompt input;
[0102] MLLM stands for Locally Deployed Multimodal Large Model, used to perform question answering and causal reasoning.
[0103] A represents the final structured identification result, which includes the type of water damage, location, risk level, analysis basis, and recommended measures.
[0104] In this embodiment, using an inspection image of a section of a highway showing road surface potholes and damage (flood damage) after rain in June 2025 as input, the specific application process of the multimodal large-model highway flood damage identification method proposed in this invention is demonstrated, including input composition, module calls, output results, and its technical effects. The process of this embodiment is as follows:
[0105] The orthophoto image, collected by the patrol drone, was input into the system. The image clearly shows a suspected structural pothole in the center of the lane. The image metadata includes the shooting time (2025-06-11 16:35), GPS location (112.5369 E, 30.5432 N), and additional aircraft number and weather tag (rain followed by clear skies). This module completes image format conversion, compression, and transmission to the recognition engine.
[0106] Load the pre-built "Highway Flood Damage Knowledge Base," which is specifically designed for different types of road surface flood damage. The content regarding road surface flood damage is shown in Table 1 below:
[0107] Definition: Road surface damage caused by water accumulation, seepage, or water pressure after rain, including cracks, subsidence, potholes, frost heave, and peeling, which affects vehicle driving safety.
[0108] Table 1 Image Recognition Features
[0109]
[0110] Prompt word design and usage:
[0111] Intelligent Agent Prompt: You are an intelligent expert assistant for road flood damage identification, specializing in image content understanding and professional identification of flood events. You should possess knowledge of civil engineering structures, hydrodynamic erosion, geological mechanisms, and disaster latency, and have cross-modal reasoning capabilities. Users upload road patrol images, which may include the shooting time, geographical location, watermarks, etc. Your task is to analyze the image content, identify the type of flood damage, and combine external knowledge to infer the causes and risk levels of the disaster.
[0112] Using Amap's MCP (Multi-Card Strategy), first extract the latitude and longitude from the image and convert it to Amap's latitude and longitude format: (106.79415, 39.655048). Then, convert the converted Amap latitude and longitude into an administrative division address. Extract the date from the image and, based on the city name in the obtained administrative division address, query the weather for that city for the next 1-7 days…
[0113] Output results and audit prompts:
[0114] Type of hazard: Road surface damaged by water;
[0115] Subcategories: Potholes and cracks;
[0116] Image features: road surface collapse, cavities;
[0117] Location analysis: Yangqu County, Taiyuan City, Shanxi Province;
[0118] Disaster mechanism: Rainwater accumulation and prolonged soaking of the road surface reduce its strength, leading to cracks and potholes;
[0119] Related weather: Continuous rainfall after the shooting date may further exacerbate road damage;
[0120] Risk level: Medium to high; timely repairs and accident prevention are required.
[0121] Repair recommendations: Check the drainage system for blockages and clean it promptly to prevent further water accumulation. For potholes, repair road surface cracks immediately by filling them with asphalt or concrete to ensure a smooth road surface.
[0122] Recognition confidence level: 0.90;
[0123] Matching case: Highway pavement water damage case (similarity 88%).
[0124] The locally deployed multimodal recognition engine works in collaboration between an image vision model (e.g., xx) and a large language model (xx):
[0125] The image input is encoded and then used for structural recognition.
[0126] Text prompts and knowledge base retrieval results are input into the language model;
[0127] The model completes image-knowledge fusion reasoning and generates structured recognition results.
[0128] As can be seen from this embodiment, the present invention can efficiently and accurately identify the type and risk level of road surface defects in actual road images, and can make a delayed judgment on the cause of water damage by combining spatiotemporal information, which significantly improves the scientific nature and automation of the identification.
[0129] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for intelligent identification of highway flood damage based on a multimodal large model, characterized in that, include: Collect multimodal data from the highway site; wherein, the multimodal data includes: highway site images, latitude and longitude coordinates, and shooting time; Using a pre-built structured knowledge base of highway flood damage, data knowledge retrieval and reasoning are performed on the multimodal data; Obtain the prompt word template that matches the recognition task; Based on the multimodal data, knowledge retrieval and reasoning results, and prompt word templates, the input content of the large language model is constructed; Based on the input content, a large language model and image recognition engine are invoked, and combined with latitude and longitude and shooting time, to intelligently identify road damage caused by floods.
2. The intelligent identification method for highway flood damage based on a multimodal large model according to claim 1, characterized in that, The structured highway flood damage knowledge base includes the structural features, image representations, causal paths, and risk levels of several types of flood damage.
3. The intelligent identification method for highway flood damage based on a multimodal large model according to claim 2, characterized in that, The types of water damage include: roadbed water damage, pavement water damage, slope water damage, bridge water damage, tunnel water damage, retaining wall water damage, drainage facility water damage, and traffic safety facility water damage.
4. The intelligent identification method for highway flood damage based on a multimodal large model according to claim 1, characterized in that, The prompt word templates that match the recognition task include: A prompt word library is built based on preset prompt words; Call up the prompt word template that matches the recognition task from the prompt word library.
5. The intelligent identification method for highway water damage based on a multimodal large model according to claim 4, characterized in that, The prompt words include: Intelligent agent prompts are used to define the recognition behavior style, inference boundaries, and output specifications of large language models; Image recognition prompts are used to guide the model to identify structural states and flood damage features in images; Structured output prompts are used to standardize the format of output fields, including type, cause, and risk. Results review and causal reasoning prompts are used to combine time, space, and weather information to complete disaster cause analysis and confidence verification.
6. The intelligent identification method for highway flood damage based on a multimodal large model according to claim 1, characterized in that, Intelligent identification of highway flood damage includes: By utilizing large language models and image recognition engines, the system can identify the types of water damage hazards in images, analyze their causes, and infer their risk levels. By combining the shooting time and latitude and longitude information, and by calling weather API and map information services, it is possible to determine whether there are any delayed causes of water damage. By combining the identification of water damage hazard types, causal analysis and risk level reasoning, as well as whether there are delayed causes of water damage, a structured identification result is output.
7. The intelligent identification method for highway flood damage based on a multimodal large model according to claim 6, characterized in that, The structured recognition results include: water damage type, image features, structural location, disaster mechanism, risk level, associated weather, recognition confidence, and matching cases.
8. A highway flood damage intelligent identification system based on a multimodal large model, characterized in that, For implementing the intelligent identification method for highway flood damage based on a multimodal large model as described in any one of claims 1-7, the system comprises: The data acquisition and input module is used to acquire multimodal data from the highway site; wherein, the multimodal data includes: highway site images, latitude and longitude coordinates, and shooting time; The multimodal recognition model execution module is used to call the locally deployed large language model and image recognition model, and combine latitude and longitude and shooting time to intelligently identify road flood damage; The knowledge base management module is used to perform knowledge retrieval and reasoning on the multimodal data using a pre-built structured knowledge base of highway flood damage. The prompt word scheduling and dialogue control module is used to obtain prompt word templates that match the recognition task; The weather / map API linkage module is used to automatically call external weather services and map information interfaces based on image metadata to obtain historical weather records for the corresponding time period at that location; the weather / map API linkage module is also used to assist the model in determining whether flood damage is related to factors such as previous weather conditions; wherein, the image metadata includes: shooting date, time, latitude and longitude; The front-end display and structured result output module is used to visualize the recognition results and synchronize the structured results to the operation and maintenance management platform or traffic safety early warning system.
Citation Information
Patent Citations
Highway subgrade water damage assessment method and system based on knowledge graph
CN117520796A
Intelligent road patrol system based on machine vision
CN117993541A
Multi-source and multi-mode fused knowledge reasoning method, system and device and medium
CN119005340A
Multi-source pavement disease identification method based on language and image large model
CN119107447A
Safety production supervision system and method based on multi-modal large model
CN119274142A