Reservoir safety intelligent inspection method and system based on YOLO and VLM fusion

By integrating YOLO and VLM, the entire process of intelligent reservoir safety inspection has been made intelligent, solving the problems of insufficient spatiotemporal coverage and insufficient model generalization ability in traditional reservoir inspection, and generating high-quality structured reports.

CN120823533AActive Publication Date: 2025-10-21JIANGXI SHUITOUJIANG INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511329017.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-10-21
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Traditional reservoir safety inspection methods rely on manual foot patrols, fixed-point observations, and periodic aerial photography, which suffer from insufficient spatiotemporal coverage, delayed hazard identification, and high labor costs. Furthermore, the YOLO algorithm cannot perform scenario-based judgments in drone inspection data processing, and traditional detection models lack generalization capabilities, making it difficult to generate structured reports with in-depth analysis.

Method used

A smart reservoir safety inspection method that integrates YOLO and Visual Language Model (VLM) is adopted. Through multi-source data acquisition and preprocessing, improved YOLO target detection, VLM post-link analysis and report generation, the entire process of target detection-semantic analysis-report generation is made intelligent.

Benefits of technology

It enhances the comprehensiveness of the water conservancy project safety monitoring system, solves the problems of insufficient scenario-based judgment and model generalization ability, and generates structured reports containing in-depth analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823533A_ABST
    Figure CN120823533A_ABST
Patent Text Reader

Abstract

The invention discloses a reservoir safety intelligent inspection method and system based on YOLO and VLM fusion, and the method comprises the following steps: S1, multi-source data collection and preprocessing: employing an unmanned plane and ground equipment to collect image / video data, and carrying out the noise reduction, enhancement and space-time alignment processing of the data; s2, improving YOLO target detection: optimizing a network structure and a training strategy; s3, link analysis after VLM: target / scene association judgment is realized by adopting the VLM; and S4, report generation. The invention provides a reservoir safety intelligent inspection method and system fusing YOLO and VLM. Cross-modal semantic understanding, zero sample reasoning and video global analysis capabilities of a visual language large model are utilized, the visual language large model is used as a post-processing tool of YOLO and is fused with the post-processing tool to work in parallel, full-process intelligentization of target detection-semantic analysis-report generation can be realized, the existing technical problems are effectively solved, and the visual language large model has a wide application prospect. And the comprehensiveness of the hydraulic engineering safety monitoring system is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water conservancy project safety monitoring. More specifically, the present invention relates to a reservoir safety intelligent inspection method and system based on the fusion of YOLO and VLM. Background Art

[0002] As a core hub for water resource regulation, the safe operation of reservoirs is directly linked to flood control, water resource supply, and ecological stability in the river basin. Traditional reservoir safety inspections rely on manual patrols, fixed-point observations, and periodic aerial photography. These inspections suffer from inherent flaws such as insufficient temporal and spatial coverage, delayed identification of potential hazards, and high labor costs. With the deep integration of drone technology and computer vision, drone inspections, with their wide coverage and flexible maneuverability, have become a crucial tool for reservoir safety monitoring.

[0003] In existing technologies, the processing of drone inspection data still has obvious limitations: First, although relying solely on the YOLO algorithm can achieve rapid target detection, it cannot complete scenario-based judgments such as "whether floating objects are close to the spillway"; second, the types of hidden dangers in reservoir scenarios are diverse, and some new targets (such as special breeding equipment and new geological disaster precursors) are difficult to construct complete data sets due to scarcity of samples, resulting in insufficient generalization capabilities of traditional detection models; third, drone inspection data lacks semantic level correlation with ground monitoring data, making it difficult to generate structured reports containing in-depth analysis. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a reservoir safety intelligent inspection method and system based on the fusion of YOLO and VLM. By setting up a reservoir safety intelligent inspection method and system that integrates YOLO and VLM; utilizing the cross-modal semantic understanding, zero-sample reasoning and video global analysis capabilities of the visual language model (VLM), using it as a post-processing tool of YOLO, and integrating it with it to work in parallel, the full process of "target detection-semantic analysis-report generation" can be realized. Intelligentization, effectively solving the problems of the existing technology, enhancing the comprehensiveness of the water conservancy project safety monitoring system, and solving the problems raised in the above-mentioned background technology.

[0005] To achieve the above objectives, the present invention provides the following technical solutions: a reservoir safety intelligent inspection method and system based on the fusion of YOLO and VLM, comprising the following steps: S1. Multi-source data collection and preprocessing: Use drones and ground equipment to collect image / video data, and perform noise reduction, enhancement, and spatiotemporal alignment on the data; S2. Improve YOLO target detection: Optimize network structure and training strategies to achieve high-precision detection of dam cracks and floating objects in drone scenarios; S3, VLM post-link analysis: Extract the visual feature vector of the target in step S2, and at the same time extract the global visual feature vector of the entire image, build a text dictionary containing the core semantic concepts of reservoir inspection, and use VLM to achieve target / scene association judgment, zero-shot target recognition and video time sequence understanding; S4. Report generation: Based on the detection results and semantic analysis, a structured report containing hidden danger information, risk assessment and disposal suggestions is generated.

[0006] In a preferred embodiment, step S1 specifically includes: (1) Data collection operation: A multi-rotor drone equipped with an industrial-grade camera is used to collect images / videos of the reservoir dam, water surface, spillway, and surrounding dangerous areas. Data on the dam bottom and gate area are supplemented by ground-based high-definition fixed cameras. The timestamp and GPS coordinates of each frame of data are recorded during collection. (2) Noise reduction: Gaussian filtering algorithm is used to remove high-frequency noise from the original image / video, and the video data is additionally subjected to inter-frame difference method to eliminate dynamic noise; (3) Enhancement processing: For complex lighting scenes, the Retinex algorithm is used to adjust the image brightness and contrast to highlight the key target features of dam cracks and floating objects; (4) Spatiotemporal alignment processing: Based on timestamps and GPS coordinates, the RANSAC algorithm is used to achieve spatial alignment of UAV and ground equipment data, and the linear interpolation method is used to complete the time synchronization of video frames to form a standardized multi-source fusion data set.

[0007] In a preferred embodiment, step S2 specifically includes: (1) Network structure optimization: Based on the YOLOv8 architecture, the backbone convolution is replaced with depthwise separable convolution to reduce the number of parameters, and an attention module is added to the neck part to enhance the feature extraction capability of small-sized dam crack targets; (2) Training strategy optimization: A special dataset for reservoir inspection containing multiple types of targets such as dam cracks and floating objects was constructed, and the model was trained using a cosine annealing learning rate strategy combined with a Focal Loss function until the accuracy of the validation set was stable. (3) Target detection execution: The preprocessed dataset is input into the improved YOLO model, reasonable confidence and IOU thresholds are set, and duplicate detection frames are removed through non-maximum suppression to achieve target detection of dam cracks, floating objects, and illegal personnel, and output the target category, bounding box coordinates, and confidence.

[0008] In a preferred embodiment, in step S3, the ROI region corresponding to the target detection bounding box is input into the pre-trained ResNet model, and then the visual feature vector is extracted to extract the global visual feature vector of the entire image; In step S3, a text dictionary containing the core semantic concepts of reservoir inspection is constructed, and the text concepts are input into the pre-trained BERT model to generate text embedding vectors, which are unified in dimension through linear transformation and matched with the visual feature vectors.

[0009] In a preferred embodiment, step S3 specifically includes: (1) Target / scene association judgment: Calculate the cosine similarity between the target ROI visual feature vector and the text embedding vector of the scene key area. When the similarity reaches the set threshold, the target and scene are determined to be associated; (2) Zero-sample target recognition: For unknown targets that are not detected, extract their ROI visual feature vectors, calculate the distance with the text embedding vectors of candidate categories in the text dictionary, and select the category with the smallest distance as the recognition result without additional training; (3) Video timing understanding: Analyze the target motion trajectory and state changes of continuous multi-frame video detection results, and output target timing correlation information.

[0010] In a preferred embodiment, step S4 specifically includes: (1) Data integration: Associating and integrating target detection results, VLM analysis results, and original spatiotemporal data to form structured data containing target information, scene association information, and spatiotemporal information; (2) Risk assessment: Using the fuzzy comprehensive evaluation method, the risk level of hidden dangers is calculated from three dimensions: severity, impact scope, and development trend; (3) Generation of disposal suggestions: Based on the risk level, the preset disposal suggestion knowledge base is called and targeted disposal suggestions are generated in combination with the temporal and spatial information of hidden dangers; (4) Structured report output: Generate a standardized report containing basic inspection information, a list of hidden dangers, risk assessment, and disposal suggestions, and store it synchronously in the database to support subsequent queries.

[0011] In a preferred embodiment, in the VLM post-link analysis of step S3, the cosine similarity calculation between the visual features and the text embedding specifically includes: Extract visual feature vectors from the target ROI area and generate text embedding vectors for text concepts in key areas of the scene; Unify the dimensions of text embedding vector and visual feature vector through linear transformation matrix; Calculate the vector cosine similarity. When the similarity reaches the set threshold, determine that the target is associated with the key area of ​​the scene, and realize the scene association judgment of floating objects / spilling outlets and personnel / dangerous areas to meet the actual inspection accuracy requirements.

[0012] In a preferred embodiment, in the VLM post-link analysis of step S3, the cosine similarity calculation between the visual features and the text embedding specifically includes: Extract visual feature vectors from the target ROI area and generate text embedding vectors for text concepts in key areas of the scene; Unify the dimensions of text embedding vector and visual feature vector through linear transformation matrix; Calculate the vector cosine similarity. When the similarity reaches the set threshold, determine that the target is associated with the key area of ​​the scene, and realize the scene association judgment of floating objects / spilling outlets and personnel / dangerous areas to meet the actual inspection accuracy requirements.

[0013] In a preferred embodiment, the reservoir safety intelligent inspection system is as follows, including the following functional layers, and each layer works in sequence: Data collection layer: It consists of a six-rotor drone with camera, GPS, and timestamp recording functions and a ground-based high-definition fixed camera to collect images / videos of key areas of the reservoir and transmit them to the pre-processing module; Preprocessing sublayer: Integrates Gaussian filtering, Retinex enhancement, and spatiotemporal alignment modules, performs the preprocessing operations of step S1, and outputs a standardized multi-source fusion dataset; YOLO detection layer: Contains a model training module and a real-time detection module, receives preprocessed data and performs the detection operation in step S2, and outputs the target detection result; VLM analysis layer: Contains visual feature extraction, text embedding generation, association judgment, zero-shot recognition, and time series understanding modules. It receives the detection results and performs the analysis operation in step S3 to output association, recognition, and time series information. Report generation layer: It consists of data integration, risk assessment, disposal suggestion and report output modules. It receives analysis results and raw data, executes step S4, and outputs standardized reports. Storage layer: Adopts an architecture that combines relational database and object storage to store structured data, raw data, model files, and reports, supporting data storage and expansion. The collaborative logic of each layer: data acquisition layer → preprocessing sublayer → YOLO detection layer → VLM analysis layer → report generation layer. Finally, all data is stored in the storage layer, forming a closed loop of data flow.

[0014] Technical effects and advantages of the present invention: The present invention provides an intelligent reservoir safety inspection method and system by integrating YOLO and VLM; utilizing the cross-modal semantic understanding, zero-sample reasoning and video global analysis capabilities of the visual language model (VLM), using it as a post-processing tool of YOLO, and integrating it with it to work in parallel, which can realize the full process intelligence of "target detection-semantic analysis-report generation", effectively solving multiple problems such as the inability of existing technologies to complete scenario judgment, insufficient generalization ability of traditional detection models, and difficulty in generating structured reports containing in-depth analysis, and further enhancing the comprehensiveness of the work of the water conservancy project safety monitoring system. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flowchart of preprocessing image data in step S1 of the present invention; Figure 2 Flowchart of post-link analysis and video understanding of the visual language large model of the present invention. DETAILED DESCRIPTION

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0017] As attached Figure 1 To the attached Figure 2 The method and system for intelligent inspection of reservoir safety based on the fusion of YOLO and VLM are shown, including the following steps: S1: Multi-source data collection and preprocessing Data collection operation: A multi-rotor drone equipped with an industrial-grade camera is used to collect images / videos of the reservoir dam, water surface, spillway, and surrounding dangerous areas. At the same time, a ground-based high-definition fixed camera is used to supplement the data collected from the bottom of the dam and the gate area. The timestamp and GPS coordinates of each frame of data are recorded during the collection.

[0018] Noise reduction: Gaussian filtering is used to remove high-frequency noise from the original image / video. The video data is additionally processed through inter-frame differencing to eliminate dynamic noise, ensuring that the image clarity meets subsequent detection requirements.

[0019] Enhancement processing: For complex lighting scenes, the Retinex algorithm is used to adjust image brightness and contrast to highlight key target features such as dam cracks and floating objects.

[0020] Spatiotemporal alignment processing: Based on timestamps and GPS coordinates, the RANSAC algorithm is used to achieve spatial alignment of drone and ground equipment data, and linear interpolation is used to complete video frame time synchronization to form a standardized multi-source fusion dataset.

[0021] S2: Improved YOLO object detection Network structure optimization: Based on the YOLOv8 architecture, the backbone convolution is replaced with depthwise separable convolution to reduce the number of parameters. An attention module is added to the neck part to enhance the feature extraction capability of targets such as small-sized dam cracks.

[0022] Training strategy optimization: A dedicated dataset for reservoir inspections containing multiple types of targets, such as dam cracks and floating objects, was constructed. The model was trained using a cosine annealing learning rate strategy combined with a Focal Loss loss function until the accuracy of the validation set stabilized.

[0023] Target detection execution: The preprocessed dataset is input into the improved YOLO model, reasonable confidence and IOU thresholds are set, and duplicate detection frames are removed through non-maximum suppression to detect targets such as dam cracks, floating objects, and illegal personnel. The target category, bounding box coordinates, and confidence level are output.

[0024] S3: VLM post-link analysis Visual feature extraction: The ROI area corresponding to the target detection bounding box is input into the pre-trained ResNet model to extract the visual feature vector, and the global visual feature vector of the entire image is also extracted at the same time.

[0025] Text embedding generation: Build a text dictionary containing core semantic concepts for reservoir inspections (such as dam cracks and spillways), input the text concepts into a pre-trained BERT model to generate text embedding vectors, unify the dimensions through linear transformation, and match them with the visual feature vectors.

[0026] Target-scene association judgment: Calculate the cosine similarity between the target ROI visual feature vector and the text embedding vector of the key area of ​​the scene. When the similarity reaches the set threshold, the target and scene are determined to be associated (such as floating objects approaching a spillway). The accuracy of association judgment meets actual inspection needs.

[0027] Zero-sample target recognition: For undetected unknown targets, its ROI visual feature vector is extracted, and the distance between it and the text embedding vector of the candidate categories in the text dictionary is calculated. The category with the smallest distance is selected as the recognition result. No additional training is required and it is suitable for new hidden danger scenarios.

[0028] Video timing understanding: Based on the detection results of multiple consecutive video frames, the system analyzes the target motion trajectory (such as the direction and speed of floating objects) and state changes (such as whether the cracks are expanding), and outputs the target timing correlation information.

[0029] S4: Report Generation Data integration: Associating and integrating target detection results, VLM analysis results and original spatiotemporal data to form structured data containing target information, scene-related information, and spatiotemporal information.

[0030] Risk assessment: Using the fuzzy comprehensive evaluation method, the hidden danger risk level (low, medium, high, and extremely high) is calculated based on the three dimensions of hidden danger severity, impact scope, and development trend.

[0031] Generation of disposal suggestions: Call the preset disposal suggestion knowledge base according to the risk level, and generate targeted disposal suggestions based on the temporal and spatial information of hidden dangers.

[0032] Structured report output: Generates standardized reports containing basic inspection information, hidden danger lists, risk assessments, and disposal recommendations, and stores them synchronously in the database to support subsequent queries.

[0033] In the VLM post-link analysis of step S3, the cosine similarity calculation between the visual features and the text embedding specifically includes: Visual feature vectors are extracted from the target ROI area, and text embedding vectors are generated for text concepts in key areas of the scene.

[0034] The text embedding vector and the visual feature vector dimensions are unified through a linear transformation matrix (the parameters are obtained by transfer learning of the pre-trained CLIP model).

[0035] Calculate the vector cosine similarity. When the similarity reaches the set threshold, determine that the target is associated with the key area of ​​the scene, and realize the association judgment of scenes such as floating objects-spilling outlets and personnel-dangerous areas, so as to meet the actual inspection accuracy requirements.

[0036] The zero-sample target recognition in step S3 specifically includes: for undetected unknown targets, extracting the complete ROI area by improving the YOLO candidate box generation module, extracting the visual feature vector of the unknown target ROI area, selecting candidate categories from the text dictionary, generating text embedding vectors for each category and unifying the dimensions, calculating the distance between the unknown target feature vector and the candidate category vector, and selecting the category with the smallest distance as the recognition result. This process does not require additional training and is suitable for new hidden danger scenarios where data sets are difficult to construct. It meets actual recognition needs and significantly increases its practicality.

[0037] An intelligent inspection system for reservoir safety based on the fusion of YOLO and a large visual language model includes the following functional layers, each of which works in sequence: Data acquisition layer: It consists of a six-rotor drone with camera, GPS, and timestamp recording functions and a ground-based high-definition fixed camera to collect images / videos of key areas of the reservoir and transmit them to the preprocessing module.

[0038] Preprocessing sublayer: Integrates Gaussian filtering, Retinex enhancement, and spatiotemporal alignment modules, performs the preprocessing operations in step 1, and outputs a standardized multi-source fusion dataset.

[0039] YOLO detection layer: Contains a model training module (supporting dataset training and network optimization) and a real-time detection module. It receives preprocessed data, performs the detection operation in step 2, and outputs the target detection results.

[0040] VLM analysis layer: contains visual feature extraction, text embedding generation, association judgment, zero-sample recognition and time series understanding modules. It receives the detection results and performs the analysis operation in step 3, outputting association, recognition and time series information.

[0041] Report generation layer: It consists of data integration, risk assessment, disposal suggestions and report output modules. It receives analysis results and original data, performs step 4 operations, and outputs standardized reports.

[0042] Storage layer: It adopts an architecture that combines relational database and object storage to store structured data and raw data, model files, and reports respectively, supporting data storage and expansion.

[0043] The collaborative logic of each layer: data acquisition layer → preprocessing sublayer → YOLO detection layer → VLM analysis layer → report generation layer. Finally, all data is stored in the storage layer, forming a closed loop of data flow.

[0044] System architecture workflow 1. Data collection layer: UAV inspection units (5K cameras + lidar) and ground monitoring equipment form a three-dimensional network, collecting ≥8 hours of high-definition video daily; 2. Processing layer: This includes the YOLO target detection module and the VLM analysis module. The former outputs target coordinates and categories, while the latter implements association judgment, zero-shot recognition, and video understanding. 3. Application layer: Generates drone inspection reports, including time series comparison of hidden dangers, slope stability assessment (based on three-week deformation analysis of the VLM), and air-ground data correlation cases; Storage layer: NAS arrays are used to store raw videos, PostgreSQL stores detection results and reports, and HDFS stores model parameters, supporting retrieval by time, region, and target type.

[0045] Multi-source heterogeneous data collection and preprocessing Constructing a three-dimensional acquisition network of "UAV + ground monitoring": UAV inspection unit: Equipped with a multi-rotor UAV (endurance ≥ 50 minutes) equipped with a 5K high-definition optical camera, thermal imager, and lidar, it conducts two patrols daily over the dam crest, slopes, and reservoir center, collecting oblique photography data (heading overlap ≥ 85%, lateral overlap ≥ 75%), thermal infrared images, and 3D point cloud data. Ground monitoring equipment: High-definition cameras (one every 50 meters, 4K resolution) are deployed along the dam body, and 360° spherical cameras are installed around the reservoir area to achieve all-weather monitoring of the nearshore area; Data preprocessing: UAV data undergoes distortion correction, image stitching (SIFT feature matching, error ≤ 2 pixels), and point cloud denoising (statistical filtering). Ground data uses BM3D denoising and Retinex enhancement. GPS / IMU fusion positioning is used to unify air-ground data into the UTM coordinate system, with a spatiotemporal alignment error of ≤ 0.5 meters.

[0046] Optimization of drone target detection based on improved YOLO YOLO, a classic algorithm in the field of real-time object detection, innovatively transforms the object detection task into a single regression problem. Its core is to divide the input image into an S×S grid, with each grid cell responsible for predicting multiple bounding boxes and their corresponding category confidences. Taking the common YOLOv8 model as an example, after an input image (e.g., 640×640 pixels) enters the network, it first passes through a backbone network consisting of a series of convolutional and pooling layers to extract image features. The convolutional layers in the backbone network use convolution kernels of varying sizes to capture rich image features such as texture and shape. The pooling layers are used to reduce the resolution of the feature maps, minimizing computational effort while preserving key features.

[0047] The YOLO algorithm is used to achieve efficient detection of reservoir hidden danger targets. The optimization for drone aerial photography scenarios is as follows: 1. Network Adaptation and Optimization: The backbone network adopts a hybrid architecture of CSPDarknet and MobileNetV3, supporting 2K-8K dynamic resolution input. The neck introduces an improved BiFPN structure, which uses attention-weighted fusion of multi-scale features to enhance the detection capability of small high-altitude targets (such as floating objects ≤30 pixels and fine cracks). 2. Dataset Construction: 150,000 drone inspection images were collected and annotated for 12 typical target categories, including dam cracks, slope hazards, floating objects, and illegal personnel. 60% of the samples were from oblique perspectives, addressing the scarcity of target annotations from drone perspectives and increasing coverage of the target collection area. 3. Training Strategy: Mosaic + CutMix data augmentation is used to simulate aerial photography occlusion and scale change scenarios. The loss function combines CIoU (positioning accuracy) and Focal Loss (category balance), with a 2.5x weighting for small objects. Transfer learning based on the COCO pre-trained model improves convergence speed by 40%. Real-time inference: Deploy a quantized model on a drone's onboard edge unit (computing power ≥ 20TOPS). After TensorRT optimization, single-frame 5K image detection takes ≤ 30ms and outputs the target category, bounding box, and coordinates (confidence ≥ 0.7).

[0048] Post-link analysis and video understanding of large visual language models Vision-Language Models (VLMs) combine technologies from computer vision and natural language processing, aiming to establish semantic connections between visual information (such as images and videos) and natural language. Their architecture typically consists of two main components: a visual encoder and a language encoder. The visual encoder extracts important visual attributes such as color, shape, and texture from image or video input and converts them into vector embeddings that can be processed by machine learning models. Early VLMs often used convolutional neural networks (CNNs) for feature extraction, but more advanced models today often utilize visual transformers (ViTs). ViTs segment images into tiles and treat these tiles as sequences, similar to tokens in language transformers. A self-attention mechanism is then applied to the tiles to create a transformer-based representation of the input image. The language encoder captures the semantic and contextual connections between words and phrases and converts them into text embeddings. Most VLMs use transformer models (such as Google's BERT and OpenAI's GPT series) as the language encoder. During the training process, VLM aligns and fuses information from visual and language encoders through strategies such as contrastive learning, masking, and generative model training, and learns how to associate images with text, thereby gaining the ability to complete a variety of visual language tasks and improving its work effectiveness.

[0049] Leveraging VLM's cross-modal understanding capabilities, we conduct in-depth analysis of YOLO detection results to achieve global video understanding and semantic reasoning: 1. Target-Scene Association Judgment: The VLM receives the target coordinates and category information output by YOLO, combines it with text descriptions of key reservoir areas (e.g., "the spillway coordinate range is (X0, Y0) - (X1, Y1)" and "the restricted area of ​​the dam is within 5 meters of the edge"), and calculates the cosine similarity between visual features and text embeddings to determine scenario-based issues such as whether floating objects are close to the spillway, whether abnormal personnel are in a dangerous area, and whether illegal vessels have entered a restricted navigation zone. The association judgment accuracy rate is ≥85%; 2. Zero-shot object recognition: For unlabeled new targets (such as new surveying and mapping equipment and illegal breeding facilities), VLM achieves zero-shot classification (confidence ≥ 0.7) without additional training by comparing the distance between the unknown target's visual features and the text embeddings of candidate categories, solving the pain point of building special hidden danger datasets. 3. Video Time Sequence Understanding: VLM performs inter-frame correlation analysis on drone cruise videos, capturing target dynamic trends through temporal attention. For example, it can track the drift trajectory of floating objects within 2 hours to predict whether they will enter the discharge zone; it can also analyze the morphological changes of slope cracks in consecutive frames to determine whether there is a risk of expansion. 4. Global semantic integration: Fusion of multiple detection frames builds a dynamic scene map, integrating target frequency, location changes, and interactions to form a global understanding of the inspection area. This includes identifying potential patterns such as the correlation between floating object accumulation areas and wind direction and the correlation between illegal personnel activities and holidays.

[0050] Multimodal report generation and air-ground data fusion Generate a structured inspection report based on YOLO detection results and VLM semantic analysis: 1. Report generation logic: VLM converts YOLO's target data into natural language descriptions (e.g., "A 1.2-meter-diameter plastic floating object was found 30 meters from the spillway") and combines it with time series analysis results to generate a report containing a list of potential hazards, risk levels (high / medium / low), and recommended actions. 2. Air-ground data association: VLM uses feature matching to correlate distant drone targets with close-up details on the ground. For example, it can fuse the attributes of a "suspicious floating object" captured by a drone with the attributes of a "2-meter diameter plastic bucket" captured by a ground camera, supplementing key information such as target size and material. Enhanced visualization: Automatically generate a 3D hidden danger distribution map (based on point cloud, accuracy ≤ 0.1 meter), a target motion trajectory time sequence diagram, and an air-ground image linkage comparison diagram. Click on the reported target to jump to the corresponding video clip.

[0051] Finally, a few points should be explained: First, in the description of this application, it should be noted that, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense, and may refer to mechanical or electrical connections, internal communication between two components, or direct connection. "Up," "down," "left," and "right" are only used to indicate relative positional relationships. When the absolute positions of the objects being described change, the relative positional relationships may also change. Secondly: The drawings of the embodiments disclosed in the present invention only involve structures related to the embodiments disclosed in the present invention. Other structures may refer to conventional designs. The same embodiment and different embodiments of the present invention may be combined with each other without conflict. Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A reservoir safety intelligent inspection method based on the fusion of YOLO and VLM, characterized by: The following steps are included: S1. Multi-source data collection and preprocessing: Use drones and ground equipment to collect image / video data, and perform noise reduction, enhancement, and spatiotemporal alignment on the data; S2. Improve YOLO target detection: Optimize network structure and training strategies to achieve high-precision detection of dam cracks and floating objects in drone scenarios; S3, VLM post-link analysis: Extract the visual feature vector of the target in step S2, and at the same time extract the global visual feature vector of the entire image, build a text dictionary containing the core semantic concepts of reservoir inspection, and use VLM to achieve target / scene association judgment, zero-shot target recognition and video time sequence understanding; S4. Report generation: Based on the detection results and semantic analysis, a structured report containing hidden danger information, risk assessment and disposal suggestions is generated.

2. The method for intelligent inspection of reservoir safety based on the fusion of YOLO and VLM according to claim 1 is characterized by: In step S1, it specifically includes: Data collection: A multi-rotor drone equipped with an industrial-grade camera is used to collect images / videos of the reservoir dam, water surface, spillway, and surrounding dangerous areas. Ground-based high-definition fixed cameras are used to supplement data collection of the dam bottom and gate area. The timestamp and GPS coordinates of each frame of data are recorded during collection. Noise reduction: Gaussian filtering is used to remove high-frequency noise from the original image / video. The video data is additionally processed using inter-frame differencing to eliminate dynamic noise. Enhancement processing: For complex lighting scenes, the Retinex algorithm is used to adjust image brightness and contrast to highlight key target features such as dam cracks and floating objects; Spatiotemporal alignment processing: Based on timestamps and GPS coordinates, the RANSAC algorithm is used to achieve spatial alignment of drone and ground equipment data, and linear interpolation is used to complete video frame time synchronization to form a standardized multi-source fusion dataset.

3. The intelligent inspection method for reservoir safety based on the fusion of YOLO and VLM according to claim 1 is characterized by: In step S2, it specifically includes: Network structure optimization: Based on the YOLOv8 architecture, the backbone convolution is replaced with depthwise separable convolution to reduce the number of parameters. An attention module is added to the neck part to enhance the feature extraction capability of small-sized dam crack targets. Training strategy optimization: A dedicated reservoir inspection dataset containing multiple categories of targets, including dam cracks and floating objects, was constructed. The model was trained using a cosine annealing learning rate strategy combined with a Focal Loss function until the validation set accuracy stabilized. Target detection execution: The preprocessed dataset is input into the improved YOLO model, reasonable confidence and IOU thresholds are set, and duplicate detection frames are removed through non-maximum suppression to achieve target detection of dam cracks, floating objects, and illegal personnel, and output the target category, bounding box coordinates, and confidence.

4. The method for intelligent inspection of reservoir safety based on the fusion of YOLO and VLM according to claim 1 is characterized by: In step S3, the ROI region corresponding to the target detection bounding box is input into the pre-trained ResNet model, and then the visual feature vector is extracted to extract the global visual feature vector of the entire image; In step S3, a text dictionary containing the core semantic concepts of reservoir inspection is constructed, and the text concepts are input into the pre-trained BERT model to generate text embedding vectors, which are unified in dimension through linear transformation and matched with the visual feature vectors.

5. The method for intelligent inspection of reservoir safety based on the fusion of YOLO and VLM according to claim 1 is characterized in that: In step S3, it specifically includes: Target / scene association judgment: Calculate the cosine similarity between the target ROI visual feature vector and the text embedding vector of the scene key area. When the similarity reaches the set threshold, the target and scene are determined to be associated; Zero-shot object recognition: For undetected unknown targets, the ROI visual feature vector is extracted, and the distance between it and the text embedding vector of the candidate categories in the text dictionary is calculated. The category with the smallest distance is selected as the recognition result, without the need for additional training; Video timing understanding: Analyze the target motion trajectory and state changes of continuous multi-frame video detection results, and output target timing correlation information.

6. The method for intelligent inspection of reservoir safety based on the fusion of YOLO and VLM according to claim 1 is characterized by: In step S4, it specifically includes: Data integration: Associating and integrating target detection results, VLM analysis results with original spatiotemporal data to form structured data containing target information, scene-related information, and spatiotemporal information; Risk assessment: Using the fuzzy comprehensive evaluation method, the risk level of hidden dangers is calculated from three dimensions: severity, impact scope, and development trend; Generation of treatment suggestions: Based on the risk level, the preset treatment suggestion knowledge base is called up, and targeted treatment suggestions are generated based on the temporal and spatial information of hidden dangers; Structured report output: Generates standardized reports containing basic inspection information, hidden danger lists, risk assessments, and disposal recommendations, and stores them synchronously in the database to support subsequent queries.

7. The method for intelligent inspection of reservoir safety based on the fusion of YOLO and VLM according to claim 5 is characterized by: In the VLM post-link analysis of step S3, the cosine similarity calculation between the visual features and the text embedding specifically includes: Extract visual feature vectors from the target ROI area and generate text embedding vectors for text concepts in key areas of the scene; Unify the dimensions of text embedding vector and visual feature vector through linear transformation matrix; Calculate the vector cosine similarity. When the similarity reaches the set threshold, determine that the target is associated with the key area of ​​the scene, and realize the scene association judgment of floating objects / spilling outlets and personnel / dangerous areas to meet the actual inspection accuracy requirements.

8. The method for intelligent inspection of reservoir safety based on the fusion of YOLO and VLM according to claim 5 is characterized by: The zero-sample target recognition in step S3 specifically includes: For unknown targets that are not detected, the complete ROI area is extracted by improving the YOLO candidate box generation module; Extract the visual feature vector of the unknown target ROI area; Select candidate categories from the text dictionary, generate text embedding vectors for each category and unify the dimensions; Calculate the distance between the unknown target feature vector and the candidate category vector, and select the category with the smallest distance as the recognition result; This process does not require additional training and is suitable for new hidden danger scenarios where data sets are difficult to construct, meeting actual identification needs.

9. The reservoir safety intelligent inspection method based on the fusion of YOLO and VLM according to any one of claims 1 to 8, wherein the reservoir safety intelligent inspection system is as follows, characterized in that: It includes the following functional layers, and each layer works together in sequence: Data collection layer: It consists of a six-rotor drone with camera, GPS, and timestamp recording functions and a ground-based high-definition fixed camera to collect images / videos of key areas of the reservoir and transmit them to the pre-processing module; Preprocessing sublayer: Integrates Gaussian filtering, Retinex enhancement, and spatiotemporal alignment modules, performs the preprocessing operations of step S1, and outputs a standardized multi-source fusion dataset; YOLO detection layer: Contains a model training module and a real-time detection module, receives preprocessed data and performs the detection operation in step S2, and outputs the target detection result; VLM analysis layer: contains visual feature extraction, text embedding generation, association judgment, zero-shot recognition, and time series understanding modules. It receives the detection results and performs the analysis operation in step S3 to output association, recognition, and time series information. Report generation layer: It consists of data integration, risk assessment, disposal suggestion and report output modules. It receives analysis results and raw data, executes step S4, and outputs standardized reports. Storage layer: Adopts an architecture that combines relational database and object storage to store structured data, raw data, model files, and reports, supporting data storage and expansion. The collaborative logic of each layer: data acquisition layer → preprocessing sublayer → YOLO detection layer → VLM analysis layer → report generation layer. Finally, all data is stored in the storage layer, forming a closed loop of data flow.

Citation Information

Patent Citations

  • Intelligent construction site safety monitoring method and system fused with multi-modal large model

    CN119399702A

  • Electric power field operation safety detection method and system based on hybrid expert model

    CN119478626A

  • Bridge disease detection method and system based on improved multi-modal visual language model

    CN119649177A

  • Construction site safety risk intelligent assessment method and system

    CN120181586A

  • Blind guiding scene identification method based on multi-modal visual large model

    CN120472387A

Cited By

  • Water conservancy multi-source image disease identification method and system based on artificial intelligence

    CN121053557A

  • Intelligent identification and risk early warning system for all-time highway slope abnormity

    CN121259767A

  • A real-time intelligent identification and risk warning system for highway slope anomalies

    CN121259767B

  • Autonomous navigation system and method for electric power inspection unmanned aerial vehicle

    CN121275000A

  • Liquid level measurement system and method based on deep learning image recognition

    CN121363995A