Fault identification methods, devices, equipment and storage media for drone inspection
By simultaneously collecting infrared and visible light image data from drones and combining them with environmental parameters, a multimodal dataset is constructed and a fault identification model is trained. This solves the problems of fault identification accuracy and adaptability of drone inspection systems in complex environments, achieving efficient intelligent diagnosis and long-term reliability.
Patent Information
- Application Number
- CN202411812097.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-10
AI Technical Summary
Existing UAV inspection systems suffer from low fault identification accuracy, poor generalization ability, and lack of adaptive mechanisms in complex environments, making it difficult to consistently guarantee long-term accuracy and reliability.
By using drones equipped with infrared and visible light cameras for synchronous inspections, multimodal image data is collected and combined with environmental parameters to construct a multimodal dataset. Fault type labeling and model training are then performed to achieve real-time fault identification, and the model is updated regularly to adapt to environmental changes.
It improves the accuracy of fault identification and the long-term reliability of the system, enables efficient inspection and intelligent diagnosis of target areas, and supports multi-user collaborative work and remote decision support.
Smart Images

Figure CN119942373B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of drone inspection technology, and in particular to a method, apparatus, equipment and storage medium for fault identification in drone inspection. Background Technology
[0002] Currently, most UAV inspection systems employ traditional image processing and machine learning methods to identify faults in collected infrared and visible light images. While existing UAV inspection systems can meet basic inspection needs to some extent, they still have some significant shortcomings. First, traditional image processing methods struggle to cope with complex and changing environmental conditions and fault types, resulting in low accuracy and robustness in fault identification. Second, existing fault identification models typically require large amounts of training data and long training times, and their generalization ability is poor when faced with new environments and fault types. Finally, existing UAV inspection systems lack effective adaptive mechanisms, failing to dynamically update image matching rules and algorithm models based on environmental changes and technological advancements, making it difficult to guarantee the long-term accuracy and reliability of the system. Summary of the Invention
[0003] The main objective of this invention is to provide a fault identification method, device, equipment, and storage medium for drone inspection, which can solve the problem that the accuracy of fault identification in existing drone inspection technologies needs to be improved.
[0004] To achieve the above objectives, the first aspect of the present invention provides a fault identification method for unmanned aerial vehicle (UAV) inspection, the method comprising:
[0005] Along the preset flight path, drones equipped with at least infrared and visible light cameras simultaneously patrol the target area, collecting infrared and visible light image data of the target area under different environmental conditions.
[0006] Based on the collected image data, a multimodal dataset is constructed. The multimodal dataset includes the image data and environmental parameters of the corresponding environmental conditions. The environmental parameters include at least timestamps, geographical locations, and meteorological parameters.
[0007] The multimodal dataset is used to label the fault type of the image data to obtain the labeled multimodal dataset.
[0008] Based on the labeled multimodal dataset, a fault identification model is trained. The fault identification model is used to automatically predict the fault type in the new infrared and visible light images when receiving new infrared and visible light image inputs, and to give corresponding early warning signals.
[0009] The fault identification model is integrated into the UAV inspection system of the UAV to realize real-time monitoring and intelligent diagnosis of fault types in the target area;
[0010] The labeled multimodal dataset and fault identification model are updated periodically to obtain an updated fault identification model.
[0011] To achieve the above objectives, a second aspect of the present invention provides a fault identification device for unmanned aerial vehicle (UAV) inspection, the device comprising:
[0012] Data collection module: Used to simultaneously inspect the target area along a preset flight path using a drone equipped with at least an infrared camera and a visible light camera, and collect infrared and visible light image data of the target area under different environmental conditions.
[0013] Collection construction module: used to construct a multimodal dataset based on the collected image data. The multimodal dataset includes the image data and environmental parameters of the corresponding environmental conditions. The environmental parameters include at least timestamps, geographical locations, and meteorological parameters.
[0014] Labeling module: used to label the fault type of the image data using the multimodal dataset, and obtain the labeled multimodal dataset;
[0015] Model training module: used to train a fault identification model based on the labeled multimodal dataset. The fault identification model is used to automatically predict the fault type in the new infrared and visible light images when receiving new infrared and visible light image inputs, and to give corresponding warning signals.
[0016] Fault monitoring module: used to integrate the fault identification model into the UAV inspection system of the UAV, so as to realize real-time monitoring and intelligent diagnosis of fault types in the target area;
[0017] Model update module: used to periodically update the labeled multimodal dataset and fault identification model to obtain an updated fault identification model.
[0018] To achieve the above objectives, a third aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method shown in the first aspect.
[0019] To achieve the above objectives, a fourth aspect of the present invention provides a computer device including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method shown in the first aspect.
[0020] The embodiments of the present invention have the following beneficial effects:
[0021] This invention provides a fault identification method for unmanned aerial vehicle (UAV) inspection. The method includes: simultaneously inspecting a target area along a preset flight path using a UAV equipped with at least an infrared camera and a visible light camera, collecting infrared and visible light image data of the target area under different environmental conditions; constructing a multimodal dataset based on the collected image data, the multimodal dataset including image data and environmental parameters corresponding to the image data, the environmental parameters including at least timestamps, geographical locations, and meteorological parameters; labeling the image data with fault type tags using the multimodal dataset to obtain a labeled multimodal dataset; training a fault identification model based on the labeled multimodal dataset, the fault identification model being used to automatically predict the fault type in new infrared and visible light images when receiving new infrared and visible light image inputs, and providing corresponding warning signals; integrating the fault identification model into the UAV inspection system to achieve real-time monitoring and intelligent diagnosis of fault types in the target area; and periodically updating the labeled multimodal dataset and the fault identification model to obtain an updated fault identification model. By introducing multimodal data acquisition, adaptive model training, and dynamic update techniques, we can achieve efficient inspection and intelligent diagnosis of the target area, thereby improving the accuracy of fault identification and the long-term reliability of the system. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] in:
[0024] Figure 1 This is a flowchart of a fault identification method for unmanned aerial vehicle (UAV) inspection according to an embodiment of the present invention;
[0025] Figure 2 This is a structural block diagram of a fault identification device for unmanned aerial vehicle (UAV) inspection according to an embodiment of the present invention;
[0026] Figure 3 This is a structural block diagram of a computer device in an embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] Please see Figure 1 , Figure 1 This is a flowchart of a fault identification method for drone inspection according to an embodiment of the present invention. This method can be applied to either a terminal or a server. The terminal can be a desktop terminal or a mobile terminal; a mobile terminal can be at least one of a mobile phone, tablet computer, or laptop computer. The server can be a standalone server or a server cluster composed of multiple servers. This embodiment uses a terminal application as an example. Figure 1 The method includes the following steps:
[0029] 101. Along the preset flight path, a drone equipped with at least an infrared camera and a visible light camera simultaneously inspects the target area and collects infrared and visible light image data of the target area under different environmental conditions.
[0030] It should be noted that, along a pre-set flight path, drones equipped with at least infrared and visible light cameras simultaneously inspect the target area, collecting infrared and visible light images of the target area under different environmental conditions. The drones follow the predetermined flight path to ensure coverage of all key locations within the target area. Infrared and visible light cameras simultaneously acquire image data; the infrared camera detects temperature anomalies, while the visible light camera captures visible light images of the target area. Simultaneously, the drones record the timestamp, geographic location information, and meteorological parameters for each image. This information is crucial for subsequent data analysis and fault identification. The target area can be a substation, power transmission line, or other power-related location, enabling intelligent inspection of power facilities using drones. The flight path is related to the inspection requirements and is used by the drones to collect data along that path.
[0031] 102. Based on the collected image data, construct a multimodal dataset, the multimodal dataset including the image data and environmental parameters of the environmental conditions corresponding to the image data, the environmental parameters including at least timestamps, geographical locations and meteorological parameters;
[0032] Furthermore, a multimodal dataset can be constructed based on the collected infrared and visible light images. This dataset includes not only the original image information but also timestamps, geographical locations, and meteorological parameters recorded by the drone. This constructed multimodal dataset forms the basis for subsequent analysis; each record in the dataset contains infrared images, visible light images, timestamps, geographical location information, and meteorological parameters. This data will be used to train and validate the fault identification model and provide a basis for subsequent fault detection.
[0033] In one feasible implementation, to avoid the impact of different environments on the acquired images, this application also considers environmental parameters, that is, constructing a set of multimodal data including image data and environmental parameters as the basis for subsequent fault identification model training, in order to reduce the influence of the environment and improve the accuracy of fault identification. Step 102 may include steps A01 to A02:
[0034] A01. Using real-time environmental perception technology, the drone's flight environment parameters are monitored in real time using sensors mounted on the drone. These environmental parameters include timestamps, meteorological parameters, and geographical location.
[0035] By introducing real-time environmental awareness technology, the sensors on board the drone monitor the flight environment in real time and make dynamic adjustments based on the preset path. When a sudden obstacle is detected ahead, the drone can automatically replan its path, bypass the obstacle, and continue to complete the mission.
[0036] Furthermore, for the various types of data collected, multi-sensor data fusion technology is used to comprehensively process the data collected by different sensors. Algorithms such as Kalman filtering or particle filtering are used to fuse the multi-sensor data, improving the accuracy and reliability of the data. The fused data can provide more comprehensive and accurate environmental perception information. An intelligent path planning algorithm is designed by combining a preset path and real-time environmental perception data. This algorithm can dynamically adjust the flight path according to the current environmental conditions and the UAV's status. The optimal path is generated using A* and Dijkstra's algorithm path planning techniques.
[0037] A02. Multimodal fusion of image data and environmental parameters is performed to construct a multimodal dataset.
[0038] Furthermore, image data and environmental parameter data are fused using a multimodal approach to construct a comprehensive dataset. Deep learning models are used to extract features from the image data, while time series analysis models are used to extract features from the environmental parameter data. Through multimodal data fusion, the system can more comprehensively understand the relationship between environmental conditions and image quality in inspection tasks. Convolutional neural networks are used to extract features from infrared and visible light images, capturing texture, edge, and color features. Simultaneously, recurrent neural networks or long short-term memory networks are used to perform time series analysis on the environmental parameter data, extracting trends and periodic characteristics of environmental parameters. Multiple fully connected layers and nonlinear activation functions are used to further fuse and represent the comprehensive feature vector.
[0039] In one feasible implementation, the method shown in this application further includes steps A03 to A04 for path planning:
[0040] A03. Record the spatiotemporal distribution of the environmental parameters and establish a three-dimensional spatiotemporal model of the environmental parameters;
[0041] In addition to obtaining the environmental parameters, this application also records their spatiotemporal distribution, i.e., the changes in environmental parameters at different time periods and geographical locations. By using high-precision timestamps and geographical location information, a three-dimensional spatiotemporal model of the environmental parameters is established, providing richer background information for subsequent data analysis and fault identification.
[0042] Furthermore, in this application, the sensor samples data at a high frequency (e.g., multiple times per second) to ensure high temporal resolution of the recorded data. High-frequency sampling captures instantaneous changes in environmental parameters, providing detailed data support for subsequent analysis. A spatiotemporal data model is used to organize and manage the collected environmental parameters. This model includes not only timestamps and geographic location information but also multi-dimensional attributes of the environmental parameters. High-precision timestamps and geographic location information accurately locate each data point in three-dimensional space. The inspection area is divided into multiple spatiotemporal grids, each containing environmental parameter data within a specific time period and geographic range. This spatiotemporal grid division organizes a large amount of data into an ordered structure, facilitating subsequent analysis and processing. Spatiotemporal correlation analysis is used to study the relationships between different environmental parameters. Pearson correlation coefficient and mutual information methods are used to quantify the correlation between different environmental parameters, revealing the intrinsic mechanisms of environmental parameter changes.
[0043] A04. When an obstacle is detected in front of the UAV, the path is replanned based on the three-dimensional spatiotemporal model and the flight path to obtain an obstacle avoidance path to bypass the obstacle.
[0044] Furthermore, by recording the spatiotemporal distribution of environmental parameters—that is, the changes in environmental parameters at different time periods and geographical locations—this application can achieve high-precision environmental perception and dynamic spatiotemporal modeling, providing comprehensive and reliable environmental data. Utilizing multimodal sensor data acquisition and spatiotemporal data fusion technology, combined with intelligent data analysis and prediction, a high-precision spatiotemporal distribution map of environmental parameters is generated, providing rich background information and scientific basis for UAV inspection tasks. This not only improves the accuracy and timeliness of fault identification, but also, when an obstacle is detected in front of the UAV, the path is replanned based on the three-dimensional spatiotemporal model and the flight path to obtain an obstacle avoidance path to bypass the obstacle. Furthermore, through intelligent path planning and dynamic obstacle avoidance strategies, it ensures that the UAV can safely and efficiently complete inspection tasks in complex environments, significantly improving inspection efficiency and the overall performance of the system.
[0045] 103. Use the multimodal dataset to label the fault type of the image data to obtain the labeled multimodal dataset;
[0046] It should be noted that after obtaining the set of multimodal data at various time points, each multimodal data can be labeled. Specifically, fault type labels are assigned based on the feature representation of the image data. Fault type labels include, but are not limited to, abnormal temperature, abnormal texture, abnormal color, etc., such as abnormal temperature rise, color change, and texture breakage.
[0047] For example, step 103 includes: performing registration processing on the infrared image and the visible light image in the image data to obtain registered image data; using the registered image data to perform fault type labeling processing to obtain fault type labels for the image data, wherein the labeled multimodal dataset includes image data, fault type labels, timestamps, geographical locations, and meteorological parameters.
[0048] The registration process includes analyzing and determining the correspondence between the infrared image and the visible light image. This results in registered image data that includes not only visible light information but also temperature information.
[0049] It should be noted that, using the aforementioned multimodal dataset, the correspondence between infrared and visible light images is analyzed and determined, particularly focusing on image feature representations under specific types of faults or anomalies, to form a set of image matching rules capable of effectively identifying fault types. Through manual annotation and feature comparison analysis, characteristic features in fault images are identified, such as abnormal temperature increases, color changes, and texture breaks. Based on these features, image matching rules are generated, including but not limited to rules for temperature anomalies, color changes, texture breaks, and multimodal feature combinations. These rules will be used to guide the development and training of the fault identification model.
[0050] Furthermore, the data can be filtered to improve its accuracy. Advanced filtering algorithms can be used to remove noise from image data and environmental parameters in multiple dimensions, improving the signal-to-noise ratio and ensuring the accuracy of subsequent analysis. That is, steps B01 to B03 are included before step 103 to perform data filtering processing:
[0051] B01. Filter the image data using a preset wavelet filtering algorithm to obtain filtered image data;
[0052] B02. Filter the environmental parameters using a preset extended Kalman filter model to obtain the filtered environmental parameters;
[0053] Understandably, in processing complex data streams, an advanced signal processing technique is first employed. This technique, based on multidimensional filtering theory, aims to extract useful information from the mixed raw data and reduce unnecessary interference. This process is not merely simple noise reduction; rather, it involves a series of carefully designed mathematical operations to enhance the proportion of the target signal relative to background noise, thereby improving the overall quality of the data and laying a solid foundation for subsequent data analysis and processing.
[0054] For image data, a preset wavelet filtering algorithm is used to filter the image data, resulting in filtered image data. Further, inverse wavelet transform is used to recombine the denoised subbands to restore the original image structure. By optimizing the reconstruction algorithm, the image clarity and detail are further improved. By employing a reverse engineering method, namely inverse wavelet transform, the components of the initially processed data stream are reassembled to restore their original structural features. This method not only restores the basic form of the data but also, by introducing specific optimization strategies such as adaptive threshold setting, further enhances the clarity and detail of the data, making the final result closer to the original state while possessing higher information density and resolution.
[0055] For environmental data, an extended Kalman filter (EPF) model is constructed to address the nonlinear characteristics of environmental parameters. This model can handle state estimation problems in nonlinear systems, improving filtering accuracy. Considering the prevalent nonlinear relationships in environmental parameters, the EPF model is introduced to solve the state estimation problem. This model is developed based on the classical Kalman filter and is specifically designed to handle complex systems that cannot be directly expressed by linear equations. By cleverly adjusting the internal parameter settings of the model and implementing a real-time response mechanism to changes in the external environment, the EPF can effectively capture the changing trends of the system state while maintaining high computational efficiency, thereby improving filtering accuracy.
[0056] B03. Obtain multimodal features from the filtered image data and environmental parameters, and dynamically optimize the weights of each modal feature using an adaptive weight adjustment algorithm to obtain optimized image data and environmental parameters. The multimodal dataset includes the optimized image data and environmental parameters.
[0057] Furthermore, the weights of each modal feature are dynamically adjusted based on their importance and reliability. An adaptive weight adjustment algorithm automatically adjusts the weights according to the variance and correlation of the features, optimizing the filtering effect. Specifically, multi-dimensional feature vectors are obtained from image data and environmental parameters, and the weights of each modal feature are dynamically optimized using the adaptive weight adjustment algorithm. This algorithm evaluates the relationships between different features by calculating the covariance matrix of the feature vectors based on the variance and correlation of the features. By introducing adaptive thresholds and dynamic adjustment factors, the algorithm can automatically adjust the weight of each feature according to the variance and correlation strength. This process not only highlights important features but also suppresses noise and redundant information, thereby improving the filtering effect. A feedback loop is introduced, allowing the algorithm to dynamically adjust the weights based on real-time system performance evaluation results, ensuring the system is always in optimal working condition. By adjusting the weights of each modal feature, the adaptive weight adjustment algorithm effectively optimizes the filtering effect of multimodal data and improves the overall system performance.
[0058] Furthermore, adaptive filtering algorithms can be introduced to dynamically adjust the filter parameters based on changes in real-time data. Extended Kalman filtering and particle filtering can be used to handle state estimation problems in nonlinear systems, improving the robustness and adaptability of the filter.
[0059] Among them, the adaptive adjustment factor w k and v k These are process noise and observation noise, respectively, assuming they follow a mean of zero and a noise covariance matrix of Q. k and R k Gaussian distribution, state vector X k State transition matrix F k Observation matrix H k λ, δ, γ, and μ are hyperparameters. It is the variance of the particle's state. It is the observed value z k and state x k The covariance of the state vector X. k State transition matrix F k Observation matrix H k P k Current state covariance, P k-1 It is the covariance of the previous state;
[0060]
[0061] P k∣k =(IK k H k )P k∣k-1
[0062] Using the above formula, the adaptive filtering algorithm can dynamically adjust the filter parameters according to changes in real-time data, handle the state estimation problem of nonlinear systems, and significantly improve the robustness and adaptability of the filter.
[0063] 104. Based on the labeled multimodal dataset, train a fault identification model. The fault identification model is used to automatically predict the fault type in the new infrared and visible light images when receiving new infrared and visible light image inputs, and give corresponding early warning signals.
[0064] Specifically, step 104 includes steps C01 to C04:
[0065] C01. The labeled multimodal dataset is used as the training dataset. Each training sample in the training dataset contains image data, environmental parameters, timestamps, geographical locations, and fault type labels.
[0066] First, training data is prepared by using labeled faulty and normal images as training data to generate a training dataset. Each training sample contains image data, environmental parameters, timestamps, geographic location information, and fault labels. The generated training dataset not only contains rich image information but also detailed environmental parameters and spatiotemporal information, providing a high-quality data foundation for subsequent model training.
[0067] C02. Use the training dataset to train the preset deep learning model to generate a fault identification model;
[0068] Then, model training can be performed using a deep learning model on the training dataset to generate a fault identification model. The model parameters are optimized through multiple rounds of iteration to improve the model's accuracy and generalization ability. The fused feature vector is input into the deep learning model for training; the fused feature vector can be obtained through subsequent steps D01 to D05, which will not be elaborated here. The model parameters are optimized through multiple rounds of iteration, using the backpropagation algorithm and optimizer to update the model weights. During training, the model's performance metrics, such as loss function value, accuracy, recall, and F1 score, are monitored in real time. Based on the monitoring results, hyperparameters such as learning rate and batch size are dynamically adjusted to improve the model's training effect. The generated fault identification model can automatically assess whether a predefined fault mode exists in a new infrared and visible light image input and provide corresponding warning signals.
[0069] C03. Validate the fault identification model using a preset validation dataset and evaluate the performance metrics of the fault identification model;
[0070] C04. Optimize the model parameters of the fault identification model based on the performance indicators to obtain an optimized fault identification model.
[0071] Finally, model validation is performed by using an independent validation dataset to validate the trained model, evaluate the model's performance metrics, and adjust the model parameters based on the validation results to further optimize the model's performance.
[0072] Prepare a separate validation dataset containing faulty and normal images not used in training, along with corresponding environmental parameters, timestamps, geographic location information, and fault labels. Use a confusion matrix to analyze the model's classification performance in detail, identifying its weaknesses and limitations. The confusion matrix reveals the model's performance across different classes. The structure of the confusion matrix assumes a binary classification problem with positive and negative classes. The confusion matrix structure is as follows:
[0073]
[0074] In summary, by utilizing the confusion matrix to analyze the model's classification performance in detail and identify its weaknesses and shortcomings, this method can significantly improve the performance and reliability of the fault identification model. Specifically, the confusion matrix provides a wealth of performance metrics, such as accuracy, precision, recall, F1 score, specificity, false positive rate, false negative rate, and AUC value. These metrics can comprehensively evaluate the model's performance across different categories. By analyzing these metrics, the false positive and false negative rates when predicting positive and negative samples can be identified, allowing for adjustments to model parameters and structure, and optimization of the training process. For example, increasing the number of positive or negative samples, adjusting the threshold, or retraining the model can effectively improve the model's precision and recall, and reduce false positives and false negatives. Ultimately, the optimized model can more accurately identify fault modes when receiving new infrared and visible light image inputs, providing reliable early warning signals, significantly improving the fault identification accuracy and robustness of the UAV inspection system, and enhancing the efficiency and safety of inspection tasks.
[0075] In one feasible implementation, step C02 includes steps D01 to D08:
[0076] D01. Use a multi-scale convolutional neural network to extract image features of different scales from infrared and visible light images; use a multi-band recurrent neural network to extract environmental features of different frequency bands from environmental parameters.
[0077] It should be noted that multi-scale convolutional neural networks are used to extract features at different scales from infrared and visible light images. By using convolutional kernels of different sizes, local and global features in the images are captured, ensuring feature diversity and richness. Multi-band recurrent neural networks are used to extract features of different frequency bands from environmental parameters, capturing trends of change over time and enhancing the temporal feature representation of environmental parameters.
[0078] D02. Through an adaptive feature stitching algorithm, the weights of the image features and environmental features are dynamically adjusted to generate a high-dimensional comprehensive feature vector, and the comprehensive feature vector is enhanced by an autoencoder.
[0079] Furthermore, an adaptive feature concatenation algorithm dynamically adjusts the weights of each modality feature to generate a high-dimensional comprehensive feature vector, and an autoencoder is used to enhance the robustness and expressiveness of the features. The adaptive feature concatenation algorithm uses the variance and covariance matrices of the features to evaluate the importance of different modal features and adjusts the weights based on the evaluation results. The generated high-dimensional comprehensive feature vector not only includes image features and environmental parameter features, but also ensures the robustness and expressiveness of the features through adaptive adjustment.
[0080] D03. Introduce a cross-modal interaction module. Through attention and gating mechanisms, the features of different modalities of the comprehensive feature vector can exchange and complement each other to obtain the interactive comprehensive feature vector.
[0081] For the enhanced integrated feature vector, a cross-modal interaction module is introduced. Through attention and gating mechanisms, information exchange and complementarity between features of different modalities are achieved. This allows features of different modalities to influence and complement each other. The attention mechanism dynamically adjusts the weights of each modal feature by calculating the correlation between features of different modalities, highlighting important features. The gating mechanism controls the flow of information through gating units, ensuring the temporal consistency of features.
[0082] D04. Design a multimodal feature fusion network, and further fuse and represent the interactive comprehensive feature vector through multiple fully connected layers and nonlinear activation functions to obtain the fused comprehensive feature vector;
[0083] A multimodal feature fusion network was designed and implemented, employing multiple fully connected layers and nonlinear activation functions to further fuse and represent the interacting features. This multimodal feature fusion network can capture the complex relationships between features from different modalities, generating richer feature representations.
[0084] D05. Use the dynamic feature selection algorithm to reduce the dimensionality of the fused comprehensive feature vector to obtain the dimensionality-reduced comprehensive feature vector.
[0085] By using a dynamic feature selection algorithm, the most relevant subset of features is dynamically selected, reducing the feature dimension, improving the computational efficiency and generalization ability of the model, and obtaining a dimensionality-reduced comprehensive feature vector.
[0086] D06. Using the reduced-dimensional comprehensive feature vector and deep learning model, the fault type is predicted to obtain the prediction result;
[0087] This can be achieved by using independent deep learning models to train data from different modalities, generating their own prediction results, and then fusing the prediction results from each modality through a multimodal prediction result fusion algorithm and an adaptive weight adjustment mechanism to ensure the accuracy and robustness of the final prediction result. Specifically, a weighted average method can be used to dynamically adjust the weights by calculating the confidence level of each modality's prediction result, a voting mechanism can be used to select the final prediction result through majority voting, and a stacking method can be used to build a multi-layer model, taking the prediction results from different modalities as input to generate the final prediction result.
[0088] An adaptive weight adjustment mechanism can also be introduced to dynamically adjust the weights based on the confidence and consistency of the prediction results. This mechanism dynamically adjusts the weights according to the confidence and consistency of the prediction results for each modality. Specifically, by calculating the confidence and consistency of the prediction results for each modality, its reliability and relevance are assessed, and the weights are dynamically adjusted to ensure that the final prediction results have the highest accuracy and robustness.
[0089] In summary, this application enables multimodal feature extraction, cross-modal interaction, multimodal feature fusion, and prediction result fusion of UAV inspection data, significantly improving the overall system performance. Specifically, it utilizes multi-scale convolutional neural networks and multi-band recurrent neural networks to extract features at different scales from infrared and visible light images, and features at different frequency bands from environmental parameters, ensuring feature diversity and richness. Through adaptive feature concatenation algorithms and autoencoders, feature weights are dynamically adjusted to enhance feature robustness and expressive power, generating high-dimensional comprehensive feature vectors. The introduction of a cross-modal interaction module and a multimodal feature fusion network allows for information exchange and complementarity between features from different modalities, generating richer feature representations. A dynamic feature selection algorithm reduces feature dimensionality, improving the model's computational efficiency and generalization ability. Independent deep learning models are used to train data from different modalities, generating their respective prediction results. A multimodal prediction result fusion algorithm and an adaptive weight adjustment mechanism are then used to fuse the prediction results from each modality, ensuring the accuracy and robustness of the final prediction result. After the final training is completed, the fused prediction results can be integrated into the decision support system, providing a visual interface to display image data, environmental parameters, fault identification results and early warning signals. The accuracy and robustness of the model can be optimized through a real-time feedback mechanism, which significantly improves the efficiency and reliability of UAV inspection tasks.
[0090] D07. Determine whether the deep learning model has converged based on the prediction results and the fault type label;
[0091] D08. If convergence is achieved, the training of the deep learning model is terminated, and the fault identification model is obtained.
[0092] Understandably, the prediction result can be a weighted fusion prediction result. The convergence of the deep learning model is determined by the fault type and fault type label of the prediction result. For example, the convergence is determined by calculating the loss value between the predicted value and the true value. If convergence is achieved, the model training is completed and a fault identification model is obtained. If convergence is not achieved, the model parameters are adjusted and the deep learning model is trained again until convergence or the preset number of iterations is reached.
[0093] 105. Integrate the fault identification model into the UAV inspection system of the UAV to realize real-time monitoring and intelligent diagnosis of fault types in the target area;
[0094] Furthermore, the fault identification model (i.e., algorithm model) is integrated into the UAV inspection system to achieve real-time monitoring and intelligent diagnosis of target areas. It also supports remote access and data sharing to facilitate multi-user collaborative work and decision support. The integrated system can receive image data and environmental parameters transmitted by the UAV in real time, automatically assess the presence of predefined fault modes in the images through the fault identification model, and generate corresponding early warning signals. These early warning signals are transmitted in real time to the ground control center and user terminal equipment via a wireless communication module, notifying operators to take appropriate preventative measures. The system provides a visual interface displaying image data, environmental parameters, fault identification results, and early warning signals, supporting simultaneous access and collaborative work by multiple users, improving work efficiency and decision-making quality. In addition, the system supports data sharing, allowing users to export or share collected data and generated reports with other users and teams, ensuring data security and privacy protection.
[0095] 106. Periodically update the labeled multimodal dataset and fault identification model to obtain an updated fault identification model.
[0096] Finally, the image matching rules and algorithm models are regularly updated to ensure they continuously adapt to environmental changes and technological advancements, maintaining the system's accuracy and reliability. Regularly updating the matching rules yields new annotated multimodal datasets, which are then used to develop new algorithm models. New inspection data, including infrared images, visible light images, timestamps, geographic location information, and environmental parameters, is collected regularly and annotated and preprocessed. The fault identification model is retrained and optimized using the new data to generate new image matching rules. A real-time feedback mechanism monitors the system's operational status and performance indicators, enabling timely identification and resolution of problems. Operators can feed back actual processing results to the system for further optimization of the model and rules. The overall system performance is regularly evaluated, and the model and rules are continuously optimized based on environmental changes and technological advancements to ensure the system's accuracy and reliability.
[0097] This application enables efficient and accurate inspection of target areas, significantly improving the accuracy and reliability of fault identification. The system can monitor the target area in real time, automatically detect and warn of potential faults or anomalies, support multi-user collaborative work and remote decision support, greatly improving inspection efficiency and response speed. Regular updates to image matching rules and algorithm models ensure the system continuously adapts to environmental changes and technological advancements, maintaining long-term accuracy and reliability, thus providing strong technical support for UAV inspection missions.
[0098] In one feasible implementation, step 106 specifically includes the following steps E01 to E04:
[0099] E01. Utilize a multi-band recurrent neural network to extract features from environmental parameters across different frequency bands. Through multi-step time series analysis, capture the changing trends of environmental parameters at different time scales.
[0100] Among these methods, multi-band recurrent neural networks can not only handle high-frequency changes but also capture low-frequency changes. Through multi-level time windows, features of environmental parameters can be extracted from different time scales, from micro to macro, generating high-dimensional time-series feature vectors. These feature vectors not only contain dynamic changes in environmental parameters but also reflect periodicity and trends at different time scales, providing rich background information for subsequent feature fusion.
[0101] E02. An adaptive feature alignment algorithm is used to align the extracted image features and environmental parameter features. This algorithm calculates the similarity matrix between features and adjusts the dimension and order of the feature vectors to ensure that features of different modalities have similar structures and distributions before stitching.
[0102] The adaptive feature alignment algorithm first calculates the similarity matrix between image features and environmental parameter features. It then evaluates the correlation between different features using cosine similarity, Euclidean distance, or other similarity metrics. Based on the similarity matrix, it dynamically adjusts the dimension and order of the feature vectors to ensure that features from different modalities have similar structures and distributions before stitching. This adaptive alignment not only eliminates inconsistencies between features from different modalities but also enhances feature complementarity and complementarity, providing a more consistent and coordinated input for subsequent feature fusion.
[0103] E03. Dimensionality reduction techniques such as principal component analysis or linear discriminant analysis are used to reduce the dimensionality of the comprehensive feature vectors, thereby improving the computational efficiency and generalization ability of the model. The dimensionality-reduced feature vectors not only retain the main information of the original features but also reduce noise and redundant information.
[0104] To reduce the dimensionality of feature vectors and improve the computational efficiency and generalization ability of the model, this method utilizes dimensionality reduction techniques such as Principal Component Analysis (PCA) or Linear Discriminant Analysis (LDA) to reduce the dimensionality of the comprehensive feature vectors. PCA identifies the principal components of the feature vectors, retains the direction of maximum variance, and removes redundant information and noise to generate dimensionality-reduced feature vectors. LDA, on the other hand, maximizes inter-class distance and minimizes intra-class distance, preserving the discriminative information of the feature vectors and generating more discriminative feature vectors. The dimensionality-reduced feature vectors not only retain the main information of the original features but also reduce the dimensionality of the features, thus improving the computational efficiency and generalization ability of the model. Dimensionality reduction not only improves the training speed of the model but also reduces the risk of overfitting, enhancing the robustness and generalization ability of the model.
[0105] E04, a multimodal feature fusion network, further fuses and represents the concatenated comprehensive feature vector through multiple fully connected layers and nonlinear activation functions. This network can capture the interaction information between features of different modalities, generating richer feature representations.
[0106] To generate richer feature representations, this application designs a multimodal feature fusion network. This network further fuses and represents the concatenated comprehensive feature vector through multiple fully connected layers and nonlinear activation functions. First, the multimodal feature fusion network maps the dimensionality-reduced feature vector to a high-dimensional space through multiple fully connected layers, capturing the interaction information between different modal features. Then, a nonlinear transformation is introduced through a nonlinear activation function to enhance the expressive power and discriminative power of the features. The multimodal feature fusion network can capture not only linear relationships between different modal features but also nonlinear relationships, generating richer feature representations. This multi-layered nonlinear transformation not only improves the robustness and generalization ability of the features but also enhances the predictive performance of the model, providing a more accurate and reliable basis for the final fault identification and early warning.
[0107] In summary, through multi-band feature extraction, adaptive feature alignment, dimensionality reduction, and multimodal feature fusion, this method not only improves the environmental perception and data processing capabilities of the UAV inspection system but also significantly enhances the overall system performance. These innovative methods are not only theoretically profound but also demonstrate strong practical value, providing robust technical support for UAV inspection tasks and significantly improving inspection efficiency and fault identification accuracy.
[0108] This invention provides a fault identification method for unmanned aerial vehicle (UAV) inspection. The method includes: simultaneously inspecting a target area along a preset flight path using a UAV equipped with at least an infrared camera and a visible light camera, collecting infrared and visible light image data of the target area under different environmental conditions; constructing a multimodal dataset based on the collected image data, the multimodal dataset including image data and environmental parameters corresponding to the image data, the environmental parameters including at least timestamps, geographical locations, and meteorological parameters; labeling the image data with fault type tags using the multimodal dataset to obtain a labeled multimodal dataset; training a fault identification model based on the labeled multimodal dataset, the fault identification model being used to automatically predict the fault type in new infrared and visible light images when receiving new infrared and visible light image inputs, and providing corresponding warning signals; integrating the fault identification model into the UAV inspection system to achieve real-time monitoring and intelligent diagnosis of fault types in the target area; and periodically updating the labeled multimodal dataset and the fault identification model to obtain an updated fault identification model. By introducing multimodal data acquisition, adaptive model training, and dynamic update techniques, we can achieve efficient inspection and intelligent diagnosis of the target area, thereby improving the accuracy of fault identification and the long-term reliability of the system.
[0109] Please see Figure 2 , Figure 2 This is a structural block diagram of a fault identification device for drone inspection according to an embodiment of the present invention, as shown below. Figure 2 The apparatus shown includes:
[0110] Data collection module 201: Used to simultaneously inspect the target area along a preset flight path using a drone equipped with at least an infrared camera and a visible light camera, and collect infrared and visible light image data of the target area under different environmental conditions.
[0111] Collection construction module 202: used to construct a multimodal dataset based on the collected image data, the multimodal dataset including the image data and environmental parameters of the environmental conditions corresponding to the image data, the environmental parameters including at least timestamps, geographical locations and meteorological parameters;
[0112] Labeling module 203: Used to label the fault type labels of the image data using the multimodal dataset, to obtain the labeled multimodal dataset;
[0113] Model training module 204: used to train a fault identification model based on the labeled multimodal dataset. The fault identification model is used to automatically predict the fault type in the new infrared and visible light images when receiving new infrared and visible light image inputs, and to give corresponding warning signals.
[0114] Fault monitoring module 205: used to integrate the fault identification model into the UAV inspection system of the UAV, so as to realize real-time monitoring and intelligent diagnosis of fault types in the target area;
[0115] Model update module 206: used to periodically update the labeled multimodal dataset and fault identification model to obtain an updated fault identification model.
[0116] It should be noted that, Figure 2 The functions of each module in the device shown are as follows: Figure 1 The steps in the method shown are similar, and will not be repeated here to avoid repetition. For details, please refer to [reference needed]. Figure 1 The content of each step in the method shown.
[0117] This invention provides a fault identification device for drone inspection, the device comprising: (see details) Figure 2 , Figure 2 This is a structural block diagram of a fault identification device for drone inspection according to an embodiment of the present invention, as shown below. Figure 2 The device includes: a data collection module, used to simultaneously inspect a target area along a preset flight path using a drone equipped with at least an infrared camera and a visible light camera, collecting infrared and visible light image data of the target area under different environmental conditions; a dataset construction module, used to construct a multimodal dataset based on the collected image data, the multimodal dataset including image data and environmental parameters corresponding to the image data, the environmental parameters including at least timestamps, geographical locations, and meteorological parameters; a labeling module, used to label the image data with fault type labels using the multimodal dataset, obtaining a labeled multimodal dataset; a model training module, used to train a fault identification model based on the labeled multimodal dataset, the fault identification model automatically predicts the fault type in new infrared and visible light images when receiving new infrared and visible light image inputs, and provides corresponding warning signals; a fault monitoring module, used to integrate the fault identification model into the drone inspection system, realizing real-time monitoring and intelligent diagnosis of fault types in the target area; and a model update module, used to periodically update the labeled multimodal dataset and the fault identification model, obtaining an updated fault identification model. By incorporating multimodal data acquisition, adaptive model training, and dynamic update technologies through the aforementioned device, efficient inspection and intelligent diagnosis of the target area can be achieved, thereby improving the accuracy of fault identification and the long-term reliability of the system.
[0118] Figure 3 An internal structural diagram of a computer device in one embodiment is shown. This computer device can specifically be a terminal or a server. Figure 3As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program, which, when executed by the processor, causes the processor to perform the aforementioned methods. The internal memory may also store a computer program, which, when executed by the processor, causes the processor to perform the aforementioned methods. Those skilled in the art will understand that… Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0119] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform actions such as... Figure 1 The steps of the method shown.
[0120] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the following actions: Figure 1 The steps of the method shown.
[0121] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0122] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0123] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A fault identification method for unmanned aerial vehicle (UAV) inspection, characterized in that, The method includes: Along the preset flight path, drones equipped with at least infrared and visible light cameras simultaneously patrol the target area, collecting infrared and visible light image data of the target area under different environmental conditions. Based on the collected image data, a multimodal dataset is constructed. The multimodal dataset includes the image data and environmental parameters of the corresponding environmental conditions. The environmental parameters include at least timestamps, geographical locations, and meteorological parameters. The multimodal dataset is used to label the fault type of the image data to obtain the labeled multimodal dataset. Based on the labeled multimodal dataset, a fault identification model is trained. The fault identification model is used to automatically predict the fault type in the new infrared and visible light images when receiving new infrared and visible light image inputs, and to give corresponding early warning signals. The fault identification model is integrated into the UAV inspection system of the UAV to realize real-time monitoring and intelligent diagnosis of fault types in the target area; The labeled multimodal dataset and fault identification model are updated periodically to obtain an updated fault identification model; The step of training a fault identification model based on the labeled multimodal dataset includes: The labeled multimodal dataset is used as the training dataset. Each training sample in the training dataset contains image data, environmental parameters, timestamps, geographical locations, and fault type labels. The preset deep learning model is trained using the training dataset to generate a fault identification model; The fault identification model is validated using a pre-set validation dataset to evaluate its performance metrics. The model parameters of the fault identification model are optimized based on the performance indicators to obtain an optimized fault identification model. The step of training a preset deep learning model using the training dataset to generate a fault identification model includes: Multi-scale convolutional neural networks are used to extract image features of different scales from infrared and visible light images; multi-band recurrent neural networks are used to extract environmental features of different frequency bands from environmental parameters. The weights of the image features and environmental features are dynamically adjusted by an adaptive feature concatenation algorithm to generate a high-dimensional comprehensive feature vector, and the comprehensive feature vector is enhanced by an autoencoder. A cross-modal interaction module is introduced, which uses attention and gating mechanisms to enable information exchange and complementarity between features of different modalities in the comprehensive feature vector, resulting in an interactive comprehensive feature vector. A multimodal feature fusion network is designed, which further fuses and represents the interactive comprehensive feature vector through multiple fully connected layers and nonlinear activation functions to obtain the fused comprehensive feature vector. The dimensionality reduction of the fused comprehensive feature vector is obtained by using a dynamic feature selection algorithm. The fault type is predicted using the reduced-dimensional comprehensive feature vector and the deep learning model, and the prediction result is obtained. Based on the prediction results and the fault type labels, determine whether the deep learning model has converged; If convergence is achieved, the training of the deep learning model is terminated, and the fault identification model is obtained.
2. The method according to claim 1, characterized in that, The step of labeling image data with fault type labels using the multimodal dataset to obtain a labeled multimodal dataset includes: The infrared image and the visible light image in the image data are registered to obtain the registered image data. The registered image data is used to perform fault type labeling to obtain the fault type label of the image data. The labeled multimodal dataset includes image data, fault type label, timestamp, geographical location and meteorological parameters.
3. The method according to claim 1, characterized in that, The step of constructing a multimodal dataset based on the collected image data includes: Through real-time environmental perception technology, the environmental parameters of the drone's flight environment are monitored in real time using sensors mounted on the drone. These environmental parameters include timestamps, meteorological parameters, and geographical location. Multimodal fusion of image data and environmental parameters is performed to construct a multimodal dataset.
4. The method according to claim 1 or 3, characterized in that, The method further includes: Record the spatiotemporal distribution of the environmental parameters and establish a three-dimensional spatiotemporal model of the environmental parameters; When an obstacle is detected in front of the drone, the path is replanned based on the three-dimensional spatiotemporal model and the flight path to obtain an obstacle avoidance path to bypass the obstacle.
5. The method according to claim 1, characterized in that, The step of labeling image data with fault type labels using the multimodal dataset to obtain the labeled multimodal dataset also includes: The image data is filtered using a preset wavelet filtering algorithm to obtain filtered image data; The environmental parameters are filtered using a preset extended Kalman filter model to obtain the filtered environmental parameters. Multimodal features are obtained from filtered image data and environmental parameters. An adaptive weight adjustment algorithm is used to dynamically optimize the weights of each modal feature to obtain optimized image data and environmental parameters. The multimodal dataset includes the optimized image data and environmental parameters.
6. A fault identification device for unmanned aerial vehicle (UAV) inspection, characterized in that, The device includes: Data collection module: Used to simultaneously inspect the target area along a preset flight path using a drone equipped with at least an infrared camera and a visible light camera, and collect infrared and visible light image data of the target area under different environmental conditions. Collection construction module: used to construct a multimodal dataset based on the collected image data. The multimodal dataset includes the image data and environmental parameters of the corresponding environmental conditions. The environmental parameters include at least timestamps, geographical locations, and meteorological parameters. Labeling module: used to label the fault type of the image data using the multimodal dataset, and obtain the labeled multimodal dataset; Model training module: used to train a fault identification model based on the labeled multimodal dataset. The fault identification model is used to automatically predict the fault type in the new infrared and visible light images when receiving new infrared and visible light image inputs, and to give corresponding early warning signals. Fault monitoring module: used to integrate the fault identification model into the UAV inspection system of the UAV, so as to realize real-time monitoring and intelligent diagnosis of fault types in the target area; Model update module: used to periodically update the labeled multimodal dataset and fault identification model to obtain an updated fault identification model; Specifically, the model training module is used to: use the labeled multimodal dataset as a training dataset, wherein each training sample in the training dataset contains image data, environmental parameters, timestamps, geographical locations, and fault type labels; train a preset deep learning model using the training dataset to generate a fault identification model; validate the fault identification model using a preset validation dataset to evaluate the performance metrics of the fault identification model; and optimize the model parameters of the fault identification model based on the performance metrics to obtain an optimized fault identification model. The step of training a pre-defined deep learning model using the training dataset to generate a fault identification model includes: extracting image features of different scales from infrared and visible light images using a multi-scale convolutional neural network; extracting environmental features of different frequency bands from environmental parameters using a multi-band recurrent neural network; dynamically adjusting the weights of the image features and environmental features using an adaptive feature concatenation algorithm to generate a high-dimensional comprehensive feature vector, and enhancing the comprehensive feature vector using an autoencoder; and introducing a cross-modal interaction module to enable information exchange and complementarity between features of different modalities in the comprehensive feature vector through attention and gating mechanisms. The interaction-adapted comprehensive feature vector is obtained; a multimodal feature fusion network is designed, and the interaction-adapted comprehensive feature vector is further fused and represented through multiple fully connected layers and nonlinear activation functions to obtain a fused comprehensive feature vector; the dimensionality of the fused comprehensive feature vector is reduced using a dynamic feature selection algorithm to obtain a dimensionality-reduced comprehensive feature vector; the dimensionality-reduced comprehensive feature vector and a deep learning model are used to predict the fault type to obtain the prediction result; based on the prediction result and the fault type label, it is determined whether the deep learning model has converged; if it has converged, the training of the deep learning model is terminated to obtain the fault identification model.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the processor performs the steps of the method as described in any one of claims 1 to 5.
8. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
IoT sensor network artificial intelligence warning, control and monitoring systems and methods
US20200364583A1
Humanoid inspection operation method and system for semantic intelligent substation robot
WO2022021739A1