Intelligent multi-mode collaborative reservoir unmanned aerial vehicle inspection system
Through the intelligent multi-modal collaborative reservoir drone inspection system, data enhancement, multi-scale feature fusion and loss function optimization technology are used to solve the problems of low accuracy and poor robustness of rockfall detection in complex environments, high-precision identification and early warning of rockfall targets, and the system's intelligence level and security guarantee capabilities are improved.
Patent Information
- Application Number
- CN202510705211.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-08
AI Technical Summary
The existing rockfall detection methods have low accuracy and poor robustness for high-speed motion, weak texture, fuzzy boundaries and high- and low-displacement rockfall targets in complex environments, making it difficult to achieve multi-stage and multi-feature joint recognition, resulting in unsatisfactory detection accuracy and robustness.
The intelligent multimodal collaborative reservoir drone patrol system is adopted to expand the training data set through the data enhancement module, combine the multi-scale feature fusion module and the loss function optimization module to build a feature extraction model, conduct multi-stage rockfall detection and judgment, and achieve high-precision and strong robustness identification and early warning of rockfall targets.
In the complex and changeable reservoir mountain slope environment, high-precision identification and early warning of rock falling targets is achieved, the model's generalization ability of different scales, angles, speeds and texture characteristics is improved, and the system's intelligence level and security guarantee capabilities are enhanced.
Smart Images

Figure CN120279450A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reservoir safety protection and early warning, and in particular to an intelligent multi-modal collaborative reservoir UAV inspection system. Background Art
[0002] At present, in the natural environment of reservoir mountain slopes, due to complex geological structures, drastic climate changes, and frequent human activities, geological disasters such as rockfalls are extremely likely to occur. This not only poses a threat to the safety of reservoir dams and water body stability, but also seriously affects the operation safety and inspection efficiency of inspection personnel. To improve the level of inspection automation, UAV inspection has gradually become an important means for monitoring reservoir mountain slopes. However, existing rockfall detection methods still mainly rely on manual inspections or traditional computer vision algorithms, and it is difficult to meet the comprehensive requirements for real-time performance, accuracy, and adaptability in complex environments. In the prior art, rockfall targets are usually detected by means of image feature extraction at a single scale in combination with simple inter-frame difference, background modeling, or optical flow estimation. Such methods have certain effects in relatively static and single-environment scenarios, but perform poorly in actual scenarios facing reservoir slopes with high-steep terrains, strong occlusions, strong illumination changes, and non-linear moving targets. Specifically, existing methods have low recognition accuracy for high-speed moving and complex-trajectory rockfall targets, and are prone to false detections and missed detections when processing rockfall images with weak textures, blurred boundaries, or partial occlusions. In particular, there is a lack of an effective modeling and fusion mechanism for the change of motion characteristics during the process of rockfall sliding from a high position to a low position, and it is difficult to achieve joint recognition of multiple stages and multiple features, ultimately resulting in unsatisfactory detection accuracy and robustness. Summary of the Invention
[0003] To solve the problems of low detection accuracy and poor robustness for high-speed moving, weak-texture, blurred-boundary, and high-low displacement rockfall targets in complex environments, the present application provides an intelligent multi-modal collaborative reservoir UAV inspection system.
[0004] An intelligent multi-modal collaborative reservoir UAV inspection system, the intelligent multi-modal collaborative reservoir UAV inspection system includes:
[0005] A data enhancement module, configured to expand a preliminary training data set including rockfall images through image enhancement methods to generate an enhanced training data set for rockfall recognition;
[0006] A multi-scale feature fusion module, configured to build a feature extraction model based on the enhanced training data set, introduce a multi-scale feature fusion mechanism in the feature extraction model, and generate a corresponding multi-scale fusion model;
[0007] A loss function optimization module, which is used to call a preset multi-task loss function to evaluate and optimize the calculation results obtained by the multi-scale fusion model, and generate a corresponding optimized fusion model;
[0008] A detection module, which is used to obtain real-time geological scene video images, perform dynamic target detection on each frame image data of the real-time geological scene video images, and generate corresponding candidate target regions;
[0009] A discrimination module, which is used to perform multi-stage rockfall detection and discrimination on the dynamic targets in the candidate target regions based on the optimized fusion model, so as to perform corresponding early warning operations.
[0010] By adopting the above technical solutions, through integrating multiple technical means such as data augmentation, deep feature fusion, loss optimization and multi-stage intelligent discrimination, it is possible to identify and give early warnings for rockfall targets with high precision and strong robustness in the complex and changeable reservoir mountain slope environment. By performing multi-dimensional enhancement processing on the original training data, the generalization ability of the model for rockfall targets with different scales, angles, speeds and texture features is improved; combined with the multi-scale feature extraction structure, the model can simultaneously perceive the global structure and local details of the target; with the joint optimization of the multi-task loss function, the performance of the model in terms of classification accuracy, boundary fitting ability and feature stability is improved; combined with the dynamic target detection and discrimination module, the system can efficiently identify candidate targets in the real-time video stream and complete the rockfall discrimination based on motion, texture and displacement trend, so as to realize the intelligent early warning of potential rockfall risks, and significantly enhance the intelligent level and safety guarantee ability of the UAV inspection of the reservoir slope.
[0011] Preferably, the discrimination module includes a motion feature recognition unit, and the motion feature recognition unit includes:
[0012] A first judgment sub-unit, which is used to judge the motion features of the dynamic target in multiple consecutive frame image data, and the motion features include pixel difference, optical flow estimation and target motion trajectory;
[0013] A first extraction sub-unit, which is used to extract corresponding dynamic change features according to the motion features, and the dynamic change features include displacement speed, acceleration and trajectory direction.
[0014] By adopting the above technical solution, by setting up a motion feature recognition unit and introducing a judgment subunit and an extraction subunit for analyzing the motion state of dynamic targets in consecutive image frames, the system can accurately obtain the motion features of the target in the sequential images, including pixel differences, optical flow estimation, and trajectory information, and further extract dynamic change features such as displacement speed, acceleration, and trajectory direction reflecting the motion state of the target, thereby significantly improving the system's recognition ability for falling rock targets with fast movement or complex trajectories, and enhancing its adaptability and detection accuracy for high-dynamic scenes.
[0015] Preferably, the discrimination module includes a texture feature recognition unit, and the texture feature recognition unit includes:
[0016] A second extraction subunit for performing deep convolution on the frame image data to extract corresponding deep texture features, where the deep texture features include surface texture patterns, edge structures, and gray-scale distributions;
[0017] A capture subunit for capturing scale features in the frame image data from multiple scales, where the scale features include weak textures and / or blurred boundaries;
[0018] A determination subunit for determining the texture difference features between the dynamic target and the surrounding background and non-falling-rock objects according to the deep texture features and the scale features.
[0019] By adopting the above technical solution, by setting up a texture feature recognition unit and integrating multiple functional subunits for deep texture extraction, scale feature capture, and texture difference judgment therein, the system can perform multi-level convolution processing on the frame image data, extract deep texture features including texture patterns, edge structures, and gray-scale distributions, simultaneously obtain key scale features such as weak textures and blurred boundaries by combining multi-scale information, and further identify the texture differences between the target and the background or non-falling-rock objects, thereby improving the system's recognition accuracy and anti-interference ability for blurred boundary and non-typical texture falling rock targets in complex scenes.
[0020] Preferably, the discrimination module includes a drop feature recognition unit, and the drop feature recognition unit includes:
[0021] A second judgment subunit for judging whether the dynamic target is a falling rock target according to the dynamic change features and the texture difference features;
[0022] A calculation subunit for continuously tracking the position changes of the falling rock target in multiple frame image data if the dynamic target is a falling rock target, and calculating the height displacement difference of the falling rock target;
[0023] A second judgment subunit, configured to judge whether the height displacement difference exceeds a preset displacement threshold. If it exceeds the preset displacement threshold, continue to execute the re-detection operation in the early warning operation, where the re-detection operation is used to perform multi-stage falling rock detection and discrimination on other dynamic targets. If it does not exceed the preset displacement threshold, execute the falling rock early warning operation in the early warning operation.
[0024] By adopting the above technical solution, by setting the drop feature recognition unit, after the system completes the joint analysis of the dynamic change feature and the texture difference feature, it can accurately judge whether the current dynamic target is a falling rock target, and track the position change of the target in consecutive image frames on the premise of judging it as a falling rock, and then calculate its height displacement difference; by introducing the displacement threshold judgment mechanism, the subsequent re-detection operation or falling rock early warning operation can be triggered when the target movement trend is obvious, so as to realize the efficient confirmation and hierarchical response of the target with an obvious falling trend, and significantly improve the intelligence and reliability of the system in the accuracy of falling rock recognition and the early warning response strategy.
[0025] Preferably, the multi-scale feature fusion module includes:
[0026] A third extraction subunit, configured to extract the target map in the enhanced training dataset at multiple different downsampling scales based on the feature pyramid structure, where the target map is used to capture the global structure information of the first-size target and the local texture details of the second-size target, and the first size is greater than the second size;
[0027] A first generation subunit, configured to input the target map into the feature extraction model to sequentially perform channel division operation, convolution operation and recombination fusion operation, and generate a corresponding fusion result;
[0028] A second generation subunit, configured to perform global average pooling operation on the fusion result for dimensionality reduction processing, and then generate a corresponding multi-scale fusion model.
[0029] By adopting the above technical solution, by setting the multi-scale feature fusion module and configuring a multi-scale extraction subunit, a channel fusion generation subunit and a dimensionality reduction generation subunit based on the feature pyramid structure therein, the system can extract the global structure and local detail features of the target image at multiple downsampling scales, realize the overall structure perception of large-size targets and the texture detail recognition of small-size targets, further improve the richness and hierarchy of feature expression through channel division, convolution processing and recombination fusion operations, and complete efficient dimensionality reduction in combination with global average pooling, and finally generate a multi-scale fusion model with high expression ability and excellent calculation efficiency, so as to improve the adaptability and accuracy of falling rock detection in multi-scale target scenarios.
[0030] Preferably, the loss function optimization module includes:
[0031] A classification loss calculation unit for calculating a corresponding classification loss value ;
[0032] A regression loss calculation unit for calculating a corresponding regression loss value ;
[0033] A feature consistency loss calculation unit for calculating a corresponding feature consistency loss value ;
[0034] Determine and / or adjust weight coefficients, where the weight coefficients include a classification weight value corresponding to the classification loss value , a regression weight value corresponding to the regression loss value , and a feature consistency loss weight value corresponding to the feature consistency loss value ;
[0035] Calculate an evaluation optimization value, and the calculation formula of the evaluation optimization value is: .
[0036] By adopting the above technical solution, by setting a loss function optimization module and constructing a multi-task loss calculation path including classification loss, regression loss, and feature consistency loss, the system can respectively evaluate the performance of the model in terms of target recognition accuracy, bounding box fitting accuracy, and multi-scale feature expression stability, and perform weighted integration on each loss term in combination with preset or dynamically adjusted weight coefficients to calculate a comprehensive evaluation optimization value for guiding model training, thereby realizing the joint optimization of the multi-dimensional capabilities of the model, improving the overall detection accuracy, positioning accuracy, and robustness under complex input conditions.
[0037] Preferably, the calculation formula of the classification loss calculation unit is:
[0038] ;
[0039] Wherein, is the true class label of the i-th training image sample in the enhanced training dataset, is the predicted class probability output by the multi-scale fusion model when processing the i-th training image sample.
[0040] By adopting the above technical solution,
[0041] Preferably, the calculation formula of the classification loss calculation unit is:
[0042] ;
[0043] Wherein, is the true bounding box position parameter of the i-th training image sample in the enhanced training dataset, is the bounding box position parameter predicted by the multi-scale fusion model for the i-th training image sample.
[0044] By adopting the above technical solution, by introducing a classification loss calculation unit and setting corresponding calculation formulas, the system can calculate the recognition error of the model in the rockfall target classification task based on the true class label of each image sample in the enhanced training dataset and the predicted probability output by the multi-scale fusion model, so as to effectively evaluate the accuracy of the model in class discrimination, and use this as the optimization basis in the training process to promote the continuous improvement of the model's discrimination ability in rockfall and non-rockfall target recognition, and enhance the overall classification reliability and detection accuracy of the system.
[0045] Preferably, the calculation formula of the feature consistency loss calculation unit is:
[0046] ;
[0047] where, is the feature map extracted from the original image sample in a pair of image samples in the enhanced training dataset at the l-th layer in the multi-scale fusion model, is the feature map extracted from the corresponding enhanced image sample of the original image sample at the same layer.
[0048] By adopting the above technical solution, by setting a feature consistency loss calculation unit and performing a comparison calculation based on the feature maps extracted from the original image and the enhanced image of the same image sample in the enhanced training dataset at the same level in the multi-scale fusion model, the system can effectively measure the consistency of the model's feature expression before and after image perturbation, ensure that the model still maintains a stable feature response when processing different enhanced input forms, thereby improving the robustness of the model to input changes and the convergence stability of the training stage, and enhancing its generalization ability and detection reliability in complex environments.
[0049] Preferably, the data augmentation module includes:
[0050] A spatial augmentation unit for performing image spatial transformation operations on the preliminary training dataset, and the image spatial transformation operations include random cropping and scaling operations, affine transformation operations, perspective transformation operations, image rotation and flipping operations, and occlusion simulation operations;
[0051] A temporal augmentation unit for performing inter-frame sampling perturbation operations, motion blur simulation operations, and temporal difference map generation operations on the preliminary training dataset;
[0052] A texture enhancement unit for performing random noise operations, brightness and contrast perturbation operations, and weathered texture simulation operations on the preliminary training dataset.
[0053] By adopting the above technical solution, by setting up a data augmentation module and dividing it into three functional units: spatial augmentation, temporal augmentation, and texture augmentation, the system can perform multi-dimensional and multi-strategy image enhancement processing on the preliminary training dataset during the training phase. Among them, the spatial augmentation operation simulates different shooting angles, target scales, and occlusion situations, the temporal augmentation operation simulates different sampling frequencies, motion blur, and dynamic change processes, and the texture augmentation operation simulates noise interference in image acquisition and the weathering effect of the rockfall surface, thereby significantly improving the diversity and complexity of the training data, enabling the model to have stronger robustness and generalization ability, and enhancing its adaptability and detection stability to rockfall targets in real complex natural environments.
[0054] In summary, the present application includes at least one of the following beneficial technical effects:
[0055] By constructing an integrated intelligent multi-modal collaborative rockfall detection system, combining various technical means such as data augmentation before training, deep multi-scale feature extraction, loss function optimization, and multi-stage feature discrimination, the problems in the prior art such as low accuracy of rockfall detection in complex natural environments and poor recognition ability for high-speed moving and weak-texture rockfall targets are systematically solved. During the training phase, the system first introduces an image enhancement strategy oriented to the diversity of natural scenes, and expands the coverage of the original dataset through perturbation means in multiple dimensions of space, time, and texture, so that the model has stronger input adaptation ability. Then, based on the enhanced training samples, a lightweight feature extraction model is constructed and a multi-scale feature fusion mechanism is introduced, which can obtain multi-level feature expressions containing global structures and local details at different downsampling scales, and then effectively identify rockfall targets with different sizes and different motion characteristics. During the training process, the system further designs a multi-task loss function, incorporating classification accuracy, bounding box fitting ability, and feature consistency into the evaluation and optimization objectives at the same time, so that the model not only has accurate judgment ability, but also can maintain feature stability at different scales. During the deployment phase, the system accesses the video image sequence obtained by the drone in real time through the detection module, and performs motion trajectory analysis, texture structure analysis, and displacement trend calculation on the dynamic targets in each frame of the image, realizing a three-stage joint recognition process from motion perception to appearance discrimination to physical displacement judgment, effectively making up for the problems of false detection and missed detection caused by single scale, thin features, and rough discrimination mechanism in the traditional method in dynamic scenes, and then realizing the function of accurately, stably, and real-time warning of high-speed rockfall targets in the reservoir mountain slope environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1It is a flowchart of a process of an intelligent multi-modal collaborative reservoir UAV inspection system in an embodiment of the present application. Specific embodiments
[0057] The present application will be further described in detail below with reference to the accompanying drawings.
[0058] In one embodiment, as Figure 1 shown, the present application discloses an intelligent multi-modal collaborative reservoir UAV inspection system. An intelligent multi-modal collaborative reservoir UAV inspection system includes:
[0059] A data enhancement module for expanding a preliminary training data set containing rockfall images through image enhancement methods to generate an enhanced training data set for rockfall recognition;
[0060] A multi-scale feature fusion module for constructing a feature extraction model based on the enhanced training data set, introducing a multi-scale feature fusion mechanism into the feature extraction model to generate a corresponding multi-scale fusion model;
[0061] A loss function optimization module for calling a preset multi-task loss function to evaluate and optimize the calculation results obtained from the multi-scale fusion model to generate a corresponding optimized fusion model;
[0062] A detection module for acquiring real-time geological scene video images, performing dynamic target detection on each frame image data of the real-time geological scene video images to generate corresponding candidate target regions;
[0063] A discrimination module for performing multi-stage rockfall detection and discrimination on dynamic targets in the candidate target regions based on the optimized fusion model to execute corresponding warning operations.
[0064] In this embodiment, the preliminary training dataset refers to the original training sample set containing typical rockfall images and their annotation information, which is collected manually or sorted from public sources and has not undergone enhancement processing. The enhanced training dataset is a richer and more complex dataset generated on the basis of the preliminary training dataset through a series of image enhancement means, and is used to improve the adaptability of the model to rockfalls in different environments and with different forms. The image enhancement method refers to the technical means of applying specific perturbations or transformation operations to image samples before training, covering the simulation processing of multiple aspects such as the spatial position, scale, illumination, and texture of the image, so as to expand the distribution range of training samples. The feature extraction model refers to the neural network structure used to automatically learn and extract discriminative feature vectors from input images, usually including multiple convolutional layers, activation functions, and pooling layers, and is used to identify local and global information in the image. The multi-scale feature fusion mechanism refers to the cross-level feature fusion strategy introduced in the neural network structure, and its core lies in simultaneously retaining the structural information and detail information of the target image extracted at different scales to enhance the model's comprehensive recognition ability for large-size and small-size targets. The multi-scale fusion model refers to the model structure with multi-level feature expression ability formed after the feature fusion mechanism, and it can simultaneously perceive the large-scale contours and tiny area texture changes in the image. The multi-task loss function refers to the composite loss function designed to simultaneously optimize multiple subtasks (such as classification, localization, consistency, etc.) during model training, usually including multiple sub-loss terms and their weighted combinations. The optimized fusion model refers to the finally trained model for target recognition obtained after loss function evaluation and parameter optimization on the basis of the multi-scale fusion model, and its performance has reached a stable state in multiple evaluation dimensions. The real-time geological scene video image refers to the real-time video frame data collected by the drone during the inspection process, and the image contains the dynamic change information of the reservoir mountain slope environment. The frame image data refers to the continuous image frames extracted from the video stream, and each frame image represents the surface image information collected by the drone at a certain time point. Dynamic target detection refers to the method of identifying the target area with relative motion behavior in the image sequence, usually realized by means of temporal image difference, motion estimation, or target tracking technology. The candidate target area refers to the image area initially screened in the image by the detection algorithm that may contain the target object and is used for further judgment by the subsequent recognition module. The multi-stage rockfall detection and discrimination refers to the comprehensive judgment of the dynamic target from multiple dimensions and multiple stages in sequence, including the analysis and recognition of aspects such as motion behavior, texture features, and displacement trends. The warning operation refers to the response mechanism triggered by the system after identifying the target with rockfall risk, including processing processes such as result output, alarm prompt, or data upload.For example, a set of image frames captured by a drone during reservoir inspection is fed into the system. After the training set is expanded by the data augmentation module, the optimized fusion model generated by training can identify a fast-falling fuzzy boundary target as a suspected rockfall, track its displacement trend in consecutive frames, and finally trigger a warning operation.
[0065] Specifically, in response to three core defects of traditional rockfall detection methods in complex natural environments, this application proposes a targeted structural optimization scheme. First, to solve the problem of low detection accuracy of fast-moving rockfall targets in complex backgrounds, the system adopts an improved lightweight texture discrimination network, ShuffleNetv2, and combines it with a three-stage rockfall detection mechanism to hierarchically model and discriminate the motion features, texture features, and elevation difference features of the target, effectively improving the recognition accuracy of high-speed moving targets. Second, to address the problem of difficult identification of rockfall targets with weak texture or blurred boundaries, the system introduces a multi-scale feature fusion module that uses a feature pyramid structure to extract texture images at different scales and realizes the efficient integration of multi-channel cross-layer feature information through group convolution partitioning and recombination fusion strategies, enabling the model to accurately perceive rockfall targets with unclear boundaries or insignificant textures. Finally, to make up for the lack of fusion analysis of high and low displacement motion features in existing methods, the system constructs a continuous feature extraction and judgment mechanism from local motion trends to overall elevation changes in the model through a fusion structure and multi-stage feature evaluation paths, achieving a global understanding and accurate discrimination of rockfall behavior in the spatio-temporal dimension, and significantly improving the detection effect and application reliability of the system in the dynamic monitoring task of mountain rockfalls.
[0066] Furthermore, the discrimination module includes a motion feature recognition unit, and the motion feature recognition unit includes:
[0067] A first judgment subunit for judging the motion features of a dynamic target in consecutive frame image data, where the motion features include pixel differences, optical flow estimation, and target motion trajectories;
[0068] A first extraction subunit for extracting corresponding dynamic change features according to the motion features, where the dynamic change features include displacement speed, acceleration, and trajectory direction.
[0069] In this embodiment, the motion feature recognition unit in the discrimination module receives the candidate target regions output by the detection module and their corresponding consecutive multi-frame image data. First, the first judgment subunit performs pixel-by-pixel comparison on the candidate target image regions in adjacent frames, calculates the pixel differences of the target regions between adjacent frames, and identifies the motion target regions with obvious change trends by clustering and statistics on the change amplitude and direction of pixel gray values. On this basis, combined with the optical flow estimation method, the optical flow vectors of the candidate target regions in each frame of image are densely calculated to obtain the local motion direction and speed distribution of the target in the image plane. Subsequently, by the change of the center point or bounding box position of the target in consecutive image frames, the spatial coordinate points corresponding to each frame are recorded, and the motion trajectory line of the target is constructed in sequence.
[0070] After the trajectory is established, the first extraction subunit calculates the Euclidean distance between every two adjacent points according to the temporal change of the trajectory points in the consecutive frame images, and obtains the average speed of the target in combination with the time interval between frames; further, the acceleration information is calculated according to the rate of change of speed; at the same time, the change of the included angle between the trajectory points is extracted, and the change trend of the trajectory direction is obtained in combination with the time dimension. The above process depends on the temporal continuity of the input image frames and the spatial saliency of the candidate regions, improves the preliminary positioning accuracy of the motion regions through pixel differences and optical flow information, and then extracts the dynamic change features including displacement speed, acceleration and trajectory direction in the target trajectory analysis stage, realizing the accurate modeling of the motion behavior of high-speed moving and non-linear path rockfall targets, and providing basic feature support for the subsequent multi-stage rockfall discrimination.
[0071] Further, the discrimination module includes a texture feature recognition unit, and the texture feature recognition unit includes:
[0072] A second extraction subunit, configured to perform deep convolution on the frame image data to extract the corresponding deep texture features, where the deep texture features include surface texture patterns, edge structures and gray value distributions;
[0073] A capture subunit, configured to capture the scale features in the frame image data from multiple scales, where the scale features include weak textures and / or blurred boundaries;
[0074] A determination subunit, configured to determine the texture difference features between the dynamic target and the surrounding background and non-rockfall objects according to the deep texture features and the scale features.
[0075] In one embodiment, the texture feature recognition unit in the discrimination module is used to analyze the texture information of the candidate target area output by the detection module in the image frame. First, the second extraction subunit receives the frame image data, inputs the candidate target area into the constructed feature extraction model for deep convolution processing, and sequentially extracts the spatial texture features of the image in multiple convolutional layers. The primary edges and gray-scale change information are obtained through shallow convolution, and the surface texture pattern, detailed structure, and local gray-scale distribution of the target are further extracted during the deep convolution process, enabling the model to distinguish the texture pattern of the typical rockfall surface from the texture of the background area and enhancing the rejection ability for non-structural backgrounds. Subsequently, based on the completion of deep convolution, the capture subunit extracts scale information from the intermediate feature maps at multiple downsampling scales to capture detailed areas such as weak textures and blurred boundaries in the image. Under the action of the feature pyramid structure, the feature maps output at different levels respectively retain the texture details at high resolution and the target contours at low resolution. By aggregating these texture expressions at different scales, the system can identify rockfall targets that are not obvious in the high-frequency region but have typical boundary features. Finally, the determination subunit performs fusion calculation on the above-extracted deep texture features and scale features, constructs the target texture expression vector, and compares the difference degree with the texture features of the surrounding image area. By calculating the distance difference between the target area and the background area in the feature space, the significance of the candidate target in texture expression is determined, thereby completing the discrimination of the texture difference between it and the background or non-rockfall objects. This process can effectively identify rockfall targets with unclear textures or blurred boundaries in complex terrain and occlusion environments, improving the recognition accuracy and environmental adaptability of the overall discrimination module.
[0076] Furthermore, the discrimination module includes a drop feature recognition unit, and the drop feature recognition unit includes:
[0077] A second judgment subunit, configured to judge whether the dynamic target is a rockfall target according to the dynamic change feature and the texture difference feature;
[0078] A calculation subunit, configured to, if the dynamic target is a rockfall target, continuously track the position change of the rockfall target in multiple frame image data and calculate the height displacement difference of the rockfall target;
[0079] A second judgment subunit, configured to judge whether the height displacement difference exceeds a preset displacement threshold. If it exceeds the preset displacement threshold, continue to execute the re-detection operation in the early warning operation, and the re-detection operation is used for multi-stage rockfall detection and discrimination of other dynamic targets. If it does not exceed the preset displacement threshold, execute the rockfall early warning operation in the early warning operation.
[0080] In this embodiment, the elevation difference feature recognition unit in the discrimination module is used to further analyze the movement trend of the dynamic target in the vertical direction to assist in completing the final discrimination of the rockfall target. First, the system respectively extracts the dynamic change features and texture difference features obtained by the movement feature recognition unit and the texture feature recognition unit, and hands them over to the second judgment subunit for preliminary comprehensive analysis. The dynamic change features include the displacement speed, acceleration, and trajectory direction of the target in consecutive frame images, and the texture difference features reflect the texture space expression differences between the candidate target and the background or non-rockfall objects. This judgment subunit fuses and determines the movement amplitude, path direction, and surface texture saliency of the target by setting a threshold model or a classification model. If the typical rockfall behavior characteristics are met, the dynamic target is marked as a suspected rockfall target. Thereafter, the calculation subunit continuously tracks the image frame sequence marked as a suspected rockfall target. Relying on the target position annotation of the detection module, the vertical position coordinates of the target in the image coordinate system are recorded in each frame, and according to the frame rate of the image sequence and the calibration parameters of the UAV's downward shooting perspective, the image coordinate difference is converted into a relative height difference, and then the height displacement difference of the target within the entire observation period is calculated. During this height difference calculation process, the system can use a time series interpolation or noise filtering mechanism to improve the stability and accuracy of the displacement trend calculation. Subsequently, the second judgment subunit performs a threshold judgment on the calculated height displacement difference. If the difference exceeds the preset displacement threshold, it is considered that the target exhibits obvious high and low displacement characteristics, which conforms to the typical spatio-temporal behavior characteristics of a rockfall falling rapidly from a high place. At this time, the system triggers the re-detection operation in the early warning operation to perform the three-stage rockfall detection and discrimination on other dynamic targets in the pending confirmation state again; if the height displacement difference of the target does not reach the threshold standard, it is regarded as a confirmed rockfall event, and the system directly enters the rockfall early warning operation process to issue a risk prompt or record an alarm message at the relevant position. Through the above multi-level and high-time-series-precision height displacement analysis mechanism, the system can accurately capture the full-process characteristics of the rockfall target from the initial movement to the large-scale vertical fall, significantly improving the accuracy and scene adaptability of rockfall discrimination.
[0081] Furthermore, the multi-scale feature fusion module includes:
[0082] A third extraction subunit, which is used to extract the target map in the enhanced training dataset at multiple different downsampling scales based on the feature pyramid structure. The target map is used to capture the global structure information of the target of the first size and the local texture details of the target of the second size, where the first size is larger than the second size;
[0083] A first generation subunit, which is used to input the target map into the feature extraction model to sequentially perform channel division operations, convolution operations, and recombination and fusion operations to generate the corresponding fusion result;
[0084] A second generation subunit, configured to perform global average pooling operation on the fusion result for dimensionality reduction processing, and then generate a corresponding multi-scale fusion model.
[0085] In this embodiment, the multi-scale feature fusion module is used to construct a feature extraction structure with hierarchical expression ability to enhance the model's recognition ability for different-sized rockfall targets. First, the third extraction subunit performs multi-scale processing on the image samples in the enhanced training dataset based on the feature pyramid structure, and extracts the feature maps of the target regions at multiple different downsampling scales. The feature maps at low resolution scales are mainly used to capture the overall shape and edge contours of larger-sized rockfall targets, while the feature maps at high resolution scales retain more detailed texture information and are suitable for identifying target regions with small sizes but significant textures. After the feature maps at each scale are extracted, the first generation subunit inputs the above multi-scale target maps into the feature extraction model. First, the features of each channel are divided through grouped convolution to extract information segments with consistent semantics within the channels, and then the divided feature maps are processed through independent convolution kernels respectively to extract the structural features corresponding to each channel. Finally, through feature recombination and fusion operations, the cross-channel and cross-scale feature expressions are integrated to form a fusion feature map with complete semantics and rich spatial details. After the fusion is completed, the second generation subunit performs global average pooling operation on the fusion result. By averaging and pooling the entire feature map in the spatial dimension, the spatial redundant information is compressed to obtain a low-dimensional and highly expressive feature vector, thereby generating the final multi-scale fusion model. This fusion model retains both the global structure information of large-sized targets and the local texture features of small-sized targets, effectively improving the adaptability, stability, and computational efficiency of the model in multi-scale rockfall detection tasks.
[0086] Furthermore, the loss function optimization module includes:
[0087] A classification loss calculation unit, configured to calculate the corresponding classification loss value ;
[0088] A regression loss calculation unit, configured to calculate the corresponding regression loss value ;
[0089] A feature consistency loss calculation unit, configured to calculate the corresponding feature consistency loss value ;
[0090] Determine and / or adjust weight coefficients, where the weight coefficients include a classification weight value corresponding to the classification loss value 、a regression weight value corresponding to the regression loss value , and a feature consistency loss weight value corresponding to the feature consistency loss value ;
[0091] The evaluation optimization value is calculated, and the calculation formula of the evaluation optimization value is: .
[0092] In this embodiment, the loss function optimization module is used to evaluate the performance and update the parameters of the feature extraction model during the training process. This module first includes a classification loss calculation unit, which receives the true class label of each training image sample in the augmented training dataset and the predicted probability value output by the model on this sample, and calculates using the cross-entropy loss function to obtain the classification loss value , which is used to measure the accuracy of the model in the classification judgment of rockfall and non-rockfall targets. Next, the regression loss calculation unit is used to evaluate the fitting accuracy of the model in target position prediction. Specifically, by comparing the true bounding box position parameters of each training sample with the difference between the predicted bounding boxes of the model, and calculating using the SmoothL1 loss function to obtain the regression loss value , which reflects the error size of the model in bounding box localization;
[0093] Furthermore, the feature consistency loss calculation unit is used to evaluate the consistency of the feature extraction results of the model under different input perturbation conditions. This unit selects each pair of original images and their augmented image samples in the augmented training dataset, extracts the feature maps at the same level in the multi-scale feature extraction model for both, calculates the Euclidean distance of their feature expressions, and obtains the feature consistency loss value , which is used to constrain the model to maintain stability in multi-scale feature expression and avoid being overly sensitive to image perturbations or feature collapse. To comprehensively measure the contribution degrees of the three types of losses, the system sets a weight coefficient adjustment mechanism, and assigns weights to the classification loss, to the regression loss, and to the feature consistency loss, and dynamically adjusts according to the actual performance in the training stage. Finally, the system performs weighted summation of the above three loss terms and the corresponding weight coefficients to obtain the final total loss function value , which is used as the optimization target during the model training process, participates in backpropagation and gradient update, so as to guide the feature extraction model to achieve balanced improvement in classification accuracy, localization accuracy, and feature stability.
[0094] Furthermore, the calculation formula of the classification loss calculation unit is:
[0095] ;
[0096] where is the true class label of the i-th training image sample in the augmented training dataset, and is the predicted class probability output by the multi-scale fusion model when processing the i-th training image sample.
[0097] In this embodiment, by setting up a loss function optimization module and constructing a multi-task loss calculation path including classification loss, regression loss, and feature consistency loss, the system can separately evaluate the performance of the model in terms of target recognition accuracy, bounding box fitting accuracy, and multi-scale feature expression stability, and combine preset or dynamically adjusted weight coefficients to weight and integrate each loss term, calculate a comprehensive evaluation optimization value for guiding model training, so as to realize the joint optimization of the multi-dimensional capabilities of the model, improve the overall detection accuracy, positioning accuracy, and robustness under complex input conditions.
[0098] Furthermore, the calculation formula of the classification loss calculation unit is:
[0099] ;
[0100] where is the true bounding box position parameter of the i-th training image sample in the enhanced training dataset, is the bounding box position parameter predicted by the multi-scale fusion model for the i-th training image sample.
[0101] In this embodiment, by introducing a classification loss calculation unit and setting the corresponding calculation formula, the system can calculate the recognition error of the model in the rockfall target classification task based on the true class label of each image sample in the enhanced training dataset and the predicted probability output by the multi-scale fusion model, so as to effectively evaluate the accuracy of the model in class discrimination, and use this as the optimization basis in the training process to promote the continuous improvement of the discrimination ability of the model in rockfall and non-rockfall target recognition, and enhance the classification reliability and detection accuracy of the overall system.
[0102] Furthermore, the calculation formula of the feature consistency loss calculation unit is:
[0103] ;
[0104] where is the feature map extracted from the l-th layer of the multi-scale fusion model for the original image sample in a pair of image samples in the enhanced training dataset, is the feature map extracted from the same layer for the enhanced image sample corresponding to the original image sample.
[0105] In this embodiment, by setting a feature consistency loss calculation unit and calculating the contrast between the feature maps extracted from the original image and the enhanced image of the same image sample in the enhanced training dataset at the same level in the multi-scale fusion model, the system can effectively measure the consistency of the model's feature representation before and after image perturbation, ensure that the model still maintains a stable feature response when processing inputs of different enhancement forms, thereby improving the robustness of the model to input changes and the convergence stability during the training phase, and enhancing its generalization ability and detection reliability in complex environments.
[0106] Furthermore, the data enhancement module includes:
[0107] A spatial enhancement unit for performing image spatial transformation operations on the preliminary training dataset. The image spatial transformation operations include random cropping and scaling operations, affine transformation operations, perspective transformation operations, image rotation and flipping operations, and occlusion simulation operations;
[0108] A temporal enhancement unit for performing inter-frame sampling perturbation operations, motion blur simulation operations, and temporal difference map generation operations on the preliminary training dataset;
[0109] A texture enhancement unit for performing random noise operations, brightness and contrast perturbation operations, and weathered texture simulation operations on the preliminary training dataset.
[0110] In this embodiment, the data enhancement module is used to perform multi-dimensional processing on the preliminary training dataset before model training to expand the sample distribution range and enhance the model's adaptability to complex rockfall scenarios. First, the spatial enhancement unit receives the original training image samples and perturbs the structure and composition of the images by performing image spatial transformation operations including random cropping and scaling, affine transformation, perspective transformation, image rotation and flipping, and occlusion simulation. Among them, random cropping and scaling are used to simulate the visual presentation of rockfall targets at different distances or different perspectives. By setting the cropping boundaries that preserve the integrity of the target and normalizing the image size, multiple versions of images that meet the input specifications of the detection model are generated; affine transformation and perspective transformation are used to simulate the imaging angle changes under drone downshooting, upshooting, or tilted flight trajectories, enabling the model to have the recognition ability under non-frontal viewing conditions; the rotation and flipping operations mainly use random rotation within the range of ±15° and horizontal mirroring to enhance the model's recognition robustness under terrain tilt and pose symmetry conditions; occlusion simulation randomly superimposes occlusion blocks on the target area of the image to simulate interference objects such as branches, stones, and dust that appear in the natural environment, effectively improving the model's detection ability for partially occluded targets.
[0111] Next, the time enhancement unit performs inter-frame perturbation operations on the image sequence samples with temporal information, including inter-frame sampling perturbation, motion blur simulation, and differential map generation. Inter-frame sampling perturbation simulates the variation of feature distribution caused by frame rate fluctuations or target speed changes in actual acquisition by changing the image sampling frame rate, enabling the model to adapt to the motion performance at different acquisition frequencies; motion blur simulation blurs the target motion trajectory area based on Gaussian kernel or directional convolution operations to simulate the dynamic blur phenomenon of a high-speed falling rock during shooting; by performing pixel-level difference calculations on adjacent frame images, a temporal difference image is generated to enhance the model's ability to model local motion features, which is particularly suitable for accurately capturing the edges of falling rock motions in complex backgrounds.
[0112] Finally, the texture enhancement unit performs visual style perturbations at multiple levels on the training image samples, including adding Gaussian noise, adjusting brightness, contrast, and saturation, and performing weathered texture simulation operations. Among them, Gaussian noise and compression artifacts simulate the common quality losses during image acquisition and transmission, enhancing the model's stability against image sharpness changes; brightness and contrast perturbations generate different versions of images under different lighting conditions through random parameter adjustment methods, enabling the model to adapt to adverse factors such as variable weather conditions and water body reflections; weathered texture simulation is based on texture mapping to superimpose natural aging effects such as sand erosion and surface particle spalling on the target area, so that the model can learn the texture feature changes of the surface of falling rocks in real scenes during the training stage and improve its recognition effect for targets with weak textures. Through the above multi-dimensional enhancement processes of space, time, and texture, the system finally generates an enhanced training dataset with diverse structures, rich content, and extensive visual perturbations, providing solid data support for the high-robustness training of subsequent models.
[0113] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An intelligent multi-modal collaborative reservoir UAV inspection system, characterized in that The intelligent multi-modal collaborative reservoir UAV inspection system includes: A data enhancement module, which is used to expand the preliminary training dataset containing rockfall images through image enhancement methods to generate an enhanced training dataset for rockfall recognition; A multi-scale feature fusion module, which is used to construct a feature extraction model based on the enhanced training dataset, introduce a multi-scale feature fusion mechanism into the feature extraction model, and generate a corresponding multi-scale fusion model; A loss function optimization module, which is used to call a preset multi-task loss function to evaluate and optimize the calculation results obtained by the multi-scale fusion model, and generate a corresponding optimized fusion model; A detection module, which is used to obtain real-time geological scene video images, perform dynamic target detection on each frame image data of the real-time geological scene video images, and generate corresponding candidate target areas; A discrimination module, which is used to perform multi-stage rockfall detection and discrimination on the dynamic targets in the candidate target areas based on the optimized fusion model to execute corresponding warning operations.
2. The intelligent multi-modal collaborative reservoir UAV inspection system according to claim 1, characterized in that, The discrimination module includes a motion feature recognition unit, and the motion feature recognition unit includes: A first judgment subunit, which is used to judge the motion features of the dynamic target in a continuous plurality of the frame image data, and the motion features include pixel difference, optical flow estimation, and target motion trajectory; A first extraction subunit, which is used to extract corresponding dynamic change features according to the motion features, and the dynamic change features include displacement speed, acceleration, and trajectory direction.
3. The intelligent multi-modal collaborative reservoir UAV inspection system according to claim 2, wherein, The discrimination module includes a texture feature recognition unit, and the texture feature recognition unit includes: A second extraction subunit, which is used to perform deep convolution on the frame image data to extract corresponding deep texture features, and the deep texture features include surface texture patterns, edge structures, and gray-scale distributions; A capture subunit, which is used to capture scale features in the frame image data from multiple scales, and the scale features include weak textures and / or blurred boundaries; A determination subunit, which is used to determine the texture difference features between the dynamic target and the surrounding background and non-rockfall objects according to the deep texture features and the scale features.
4. The intelligent multi-modal collaborative reservoir UAV inspection system according to claim 3, wherein, The discrimination module includes a drop feature recognition unit, and the drop feature recognition unit includes: A second judgment subunit, which is used to judge whether the dynamic target is a rockfall target according to the dynamic change features and the texture difference features; A calculation subunit, which is used to continuously track the position changes of the rockfall target in a continuous plurality of the frame image data if the dynamic target is a rockfall target, and calculate the height displacement difference of the rockfall target; A second judgment subunit, which is used to judge whether the height displacement difference exceeds a preset displacement threshold. If it exceeds the preset displacement threshold, the re-detection operation in the warning operation is continued, and the re-detection operation is used to perform multi-stage rockfall detection and discrimination on other dynamic targets. If it does not exceed the preset displacement threshold, the rockfall warning operation in the warning operation is executed.
5. The intelligent multi-modal collaborative reservoir UAV inspection system according to claim 1, characterized in that, The multi-scale feature fusion module includes: A third extraction subunit, configured to extract target graphs in the enhanced training dataset at multiple different downsampling scales based on a feature pyramid structure, where the target graphs are used to capture the global structural information of the first-size targets and the local texture details of the second-size targets, and the first size is greater than the second size; A first generation subunit, configured to input the target graphs into the feature extraction model to sequentially perform a channel division operation, a convolution operation, and a recombination and fusion operation, and generate corresponding fusion results; A second generation subunit, configured to perform global average pooling on the fusion results for dimensionality reduction processing, and then generate corresponding multi-scale fusion models.
6. The intelligent multi-modal collaborative reservoir UAV inspection system according to claim 1, characterized in that, The loss function optimization module includes: A classification loss calculation unit for calculating a corresponding classification loss value ; Regression loss calculation unit, configured to calculate a corresponding regression loss value ; A feature consistency loss calculation unit for calculating a corresponding feature consistency loss value ; Determine and / or adjust weight coefficients, where the weight coefficients include a classification weight value corresponding to the classification loss value , a regression weight value corresponding to the regression loss value , and a feature consistency loss weight value corresponding to the feature consistency loss value ; Calculate the evaluation optimization value, and the calculation formula of the evaluation optimization value is: .
7. An intelligent multi-modal collaborative reservoir UAV inspection system according to claim 6, characterized in that, The calculation formula of the classification loss calculation unit is: ; wherein, is the true class label of the i-th training image sample in the enhanced training dataset, is the predicted class probability output by the multi-scale fusion model when processing the i-th training image sample.
8. An intelligent multi-modal collaborative reservoir UAV inspection system according to claim 6, characterized in that, The calculation formula of the classification loss calculation unit is: ; wherein, is the true bounding box position parameter of the i-th training image sample in the enhanced training dataset, is the bounding box position parameter predicted by the multi-scale fusion model for the i-th training image sample.
9. An intelligent multi-modal collaborative reservoir UAV inspection system according to claim 6, characterized in that, The calculation formula of the feature consistency loss calculation unit is: ; Among them, is the feature map extracted from the original image sample in a pair of image samples in the enhanced training dataset at the l-th layer in the multi-scale fusion model, is the feature map extracted from the corresponding enhanced image sample of the original image sample at the same layer.
10. An intelligent multi-modal collaborative reservoir UAV inspection system according to claim 6, characterized in that, The data augmentation module includes: A spatial augmentation unit, configured to perform an image spatial transformation operation on the preliminary training dataset, and the image spatial transformation operation includes random cropping and scaling operations, affine transformation operations, perspective transformation operations, image rotation and flipping operations, and occlusion simulation operations; A temporal augmentation unit, configured to perform an inter-frame sampling perturbation operation, a motion blur simulation operation, and a temporal difference map generation operation on the preliminary training dataset; A texture augmentation unit, configured to perform a random noise operation, a brightness and contrast perturbation operation, and a weathered texture simulation operation on the preliminary training dataset.
Citation Information
Patent Citations
Mine rockfall detection method and device based on deep learning
CN113744291A
Real-time rockfall monitoring method, system and device based on machine vision and storage medium
CN118675106A
Road rockfall detection method and device based on deep learning and electronic equipment
CN119251689A
Cited By
Universe unattended intelligent monitoring method based on multi-source data fusion
CN120932182A