Video Stream Analysis Method and System Based on Drone Inspection
Through the drone, the video stream data is collected, spatial and temporal characteristics are extracted for comprehensive abnormality detection, which solves the problems of low efficiency of traditional inspection methods and fixed camera blind spots, and achieves efficient and accurate abnormality detection and dynamic path optimization.
Patent Information
- Application Number
- CN202510660701.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-22
AI Technical Summary
Traditional inspection methods are inefficient and have limited scope, and fixed camera monitoring has blind spots. The existing video stream analysis methods cannot comprehensively and accurately capture abnormal information, resulting in missed and missed inspections.
The drone is equipped with a camera device to collect video stream data, extract spatial correlation characteristics and temporal dynamic characteristics, combine it with an abnormality detection model for comprehensive analysis, generate dynamic path optimization instructions, dynamically adjust patrol paths, and update the model through incremental learning.
It has achieved flexible and comprehensive inspections, improved the accuracy of abnormal detection, reduced missed inspections and missed inspections, improved inspection efficiency and targetedness, and adapted to changes in the target area.
Smart Images

Figure CN120182873B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles, and more particularly, to a method and system for video stream analysis based on unmanned aerial vehicle patrol inspection. Background Art
[0002] In many fields today, such as power facility monitoring, urban infrastructure inspection, natural environment monitoring, etc., it is crucial to conduct efficient and accurate inspection work on target areas. Traditional inspection methods face many difficulties in practical applications.
[0003] The early manual inspection mode is the most basic means. Staff need to go to the target area on site and rely on visual observation and simple detection tools to identify potential problems. However, this method has significant limitations. For large-scale target areas, manual inspection not only consumes a large amount of time and labor costs, but also has extremely low work efficiency. In some complex geographical environments, such as mountainous areas, canyons, rivers and other regions, it is difficult for humans to reach, which makes it difficult to carry out inspection work in some areas, thus unable to achieve full coverage of the target area. In addition, manual inspection is also affected by human subjective factors. There are differences in the detection standards and judgment abilities of different personnel, which may lead to inaccurate and inconsistent detection results.
[0004] With the development of technology, fixed camera monitoring systems have gradually been applied. These cameras are installed at key positions in the target area and can obtain monitoring images in real time. However, the monitoring range of fixed cameras is fixed and can only cover a limited area. To achieve wider monitoring, a large number of cameras need to be installed, which not only increases the equipment cost and maintenance difficulty, but also there may be monitoring blind spots between multiple cameras, and a continuous and comprehensive monitoring network cannot be formed. Moreover, fixed cameras cannot dynamically adjust the monitoring perspective and range according to the actual situation, and for some suddenly emerging abnormal situations, they cannot be tracked and observed in detail in a timely manner.
[0005] In terms of video stream analysis, most of the existing technologies focus on the analysis of single features. Some methods only focus on the spatial features of video frames, such as the shape, color of objects, etc., while ignoring the dynamic change information of the video in the time dimension. Other methods only analyze the time dynamic features, such as the movement trajectory of objects, etc., without fully considering the spatial correlation relationship. This single-feature analysis method cannot comprehensively and accurately capture the abnormal information in the video stream, resulting in a low detection accuracy rate for abnormal situations and prone to missed detections and false detections. Summary of the Invention
[0006] In view of the problems mentioned above, in combination with the first aspect of the present invention, embodiments of the present invention provide a method for video stream analysis based on unmanned aerial vehicle patrol inspection, the method comprising:
[0007] Collect the original video stream data of the target area through a camera device carried by a drone, where the original video stream data includes a plurality of consecutive video frame sequences;
[0008] Extract video features from the original video stream data to obtain a spatial association feature set and a temporal dynamic feature set of the video frame sequences;
[0009] Based on a preset anomaly detection model, perform anomaly detection on the spatial association feature set and the temporal dynamic feature set to generate an anomaly distribution feature set of the target area;
[0010] Generate a dynamic path optimization instruction according to the anomaly distribution feature set and the real-time flight parameters of the drone, where the dynamic path optimization instruction is used to adjust the inspection path of the drone to cover the area corresponding to the anomaly distribution feature set;
[0011] Call the incremental learning module to update the parameters of the anomaly detection model based on the anomaly distribution feature set.
[0012] In another aspect, an embodiment of the present invention further provides a video stream analysis system based on drone inspection, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0013] Based on the above aspects, the embodiment of the present invention collects the original video stream data of the target area through a camera device carried by a drone, can flexibly and comprehensively inspect the target area, overcomes the problems of low efficiency and limited range of traditional inspection methods, extracts a spatial association feature set and a temporal dynamic feature set from the original video stream data, and performs comprehensive anomaly detection based on a preset anomaly detection model. Compared with single-feature analysis, by considering the spatial association and temporal dynamic changes between video frames, it can more accurately identify anomalies in the target area, greatly improving the accuracy of anomaly detection, effectively reducing the occurrence of missed detections and false detections. Generating a dynamic path optimization instruction according to the anomaly distribution feature set and the real-time flight parameters of the drone can enable the drone to dynamically adjust the inspection path, focus on covering the anomaly area, improve the inspection efficiency and pertinence. At the same time, calling the incremental learning module to update the parameters of the anomaly detection model based on the anomaly distribution feature set can enable the anomaly detection model to adapt to changes in the situation of the target area, maintain a high detection accuracy, and improve the effect and quality of the inspection of the target area. Description of the Drawings
[0014] Figure 1It is a schematic flowchart of the execution process of the video stream analysis method based on drone inspection provided by an embodiment of the present invention.
[0015] Figure 2 It is a schematic diagram of exemplary hardware and software components of the video stream analysis system based on drone inspection provided by an embodiment of the present invention. Detailed implementation manners
[0016] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 It is a schematic flowchart of the video stream analysis method based on drone inspection provided by an embodiment of the present invention. The video stream analysis method based on drone inspection will be introduced in detail below.
[0017] Step S110: Collect the original video stream data of the target area through the camera device carried by the drone. The original video stream data includes a plurality of consecutive video frame sequences.
[0018] In this embodiment, the drone is deployed to the target area to perform an inspection task. The camera device carried by the drone operates at a set frame rate of F frames per second. During the flight of the drone along the preset path, the camera device continuously captures images of the target area. Within each second, the camera device obtains F video frames. As time goes by, these video frames are arranged in chronological order to form the original video stream data. For example, in an inspection scenario of a large factory, the camera device will record video frames of the buildings, equipment, pipelines, etc. inside the factory at different times.
[0019] Step S120: Extract video features from the original video stream data to obtain a spatial association feature set and a temporal dynamic feature set of the video frame sequence.
[0020] After obtaining the original video stream data, in order to deeply mine the useful information therein, it is necessary to extract video features to separate the spatial association features and temporal dynamic features in the video frame sequence. The spatial association features can reflect the spatial relationships and structural information between different regions within a video frame, while the temporal dynamic features reflect the changes and motion information between adjacent video frames.
[0021] Step S121: Perform preprocessing operations on the original video stream data to generate a standardized video frame sequence. The preprocessing operations include geometric distortion correction, illumination equalization processing, and noise filtering processing.
[0022] Since the camera device is affected by various factors during the shooting process, the original video stream data may have problems such as geometric distortion, uneven illumination, and noise interference. Therefore, preprocessing is required. In terms of geometric distortion correction, the lens characteristics of the camera device or the flight attitude of the drone may cause the captured images to be distorted. For example, the buildings in the target area are originally regular rectangles, but may appear trapezoidal or other irregular shapes in the image. To correct this distortion, a method based on feature point matching can be used. First, representative feature points, such as corner points and edge points, are extracted from the original video frame. These feature points have a relatively stable position relationship in normal images and distorted images. Suppose there is a set of feature points A in the normal image, and the corresponding set of feature points in the distorted image is B. By finding the correspondence between set A and set B, and using a mathematical transformation model (such as affine transformation or perspective transformation), the transformation parameters are calculated. Then, according to these transformation parameters, the entire video frame is transformed to restore the distorted image to a normal geometric shape, obtaining the corrected video frame.
[0023] Illumination equalization processing is to solve the problem of excessive differences in illumination intensity in different regions of the video frame. In the target area, there may be a situation where some areas are shaded while other areas are illuminated by strong light, which will affect the accuracy of subsequent feature extraction. The method of histogram equalization can be used to achieve illumination equalization. For each video frame, first, the histogram of its pixel values is statistically calculated, which reflects the distribution of pixel values at different gray levels. Then, the histogram is normalized, and the cumulative distribution function of each gray level is calculated. Finally, according to the cumulative distribution function, each pixel value in the video frame is mapped to a new gray level, making the pixel value distribution of the entire video frame more uniform, thereby achieving the effect of illumination equalization.
[0024] Noise filtering processing is used to remove the noise generated by external interference or the device itself in the video frame. Common types of noise include Gaussian noise, salt-and-pepper noise, etc. Taking Gaussian noise as an example, the method of Gaussian filtering can be used for processing. Gaussian filtering is a linear smoothing filter, which performs weighted averaging on each pixel in the video frame and its neighboring pixels according to the weights of the Gaussian function. Suppose a neighborhood of a set size (such as 3×3, 5×5, etc.) is selected with a certain pixel as the center, the weights of each pixel in the neighborhood are calculated according to the Gaussian function, and then these pixel values are multiplied by the corresponding weights and summed to obtain the filtered value of the center pixel. By performing such operations on each pixel in the video frame, the influence of Gaussian noise can be effectively reduced, making the video frame clearer. After geometric distortion correction, illumination equalization processing, and noise filtering processing, the original video stream data is converted into a standardized video frame sequence, providing a high-quality data basis for subsequent feature extraction.
[0025] Step S122: Invoke the spatial feature encoder to perform frame-by-frame analysis on the standardized video frame sequence, extract the local texture features and global structure features of each video frame in the video frame sequence, and merge the local texture features and the global structure features into the spatial association feature set.
[0026] Input the standardized video frame sequence into the spatial feature encoder for frame-by-frame analysis to extract the local texture features and global structure features of each video frame. The spatial feature encoder mainly consists of a shallow convolutional module, a spatial attention module, and a deep convolutional module.
[0027] Step S1221: Input the standardized video frame sequence into the shallow convolutional module of the spatial feature encoder, and use multi-scale convolutional kernels to extract local texture features for each video frame, obtaining a multi-scale texture feature map set.
[0028] In the shallow convolutional module, use multiple convolutional kernels with different scales to perform convolutional operations on each video frame in the standardized video frame sequence. Convolutional kernels with different scales can capture texture information of different sizes in the video frame. For example, small-scale convolutional kernels can detect fine textures in the video frame, such as the small patterns on the surface of an object; while large-scale convolutional kernels can capture larger-scale texture features, such as the overall contour texture of an object. Suppose n different scales of convolutional kernels are used, denoted as K1, K2,..., Kn respectively. For each video frame V, perform convolutional operations with these convolutional kernels in sequence. For the convolutional kernel Ki (i = 1, 2,..., n), the convolutional operation will generate a corresponding texture feature map Fi. After processing by all convolutional kernels, a multi-scale texture feature map set {F1, F2,..., Fn} is obtained. Each texture feature map reflects the texture information of the video frame at the set scale.
[0029] Step S1222: Input the multi-scale texture feature map set into the spatial attention module of the spatial feature encoder to generate the weight distribution of different spatial positions in the multi-scale texture feature map set.
[0030] The role of the spatial attention module is to assign different weights to different spatial positions in the set of multi-scale texture feature maps to highlight important texture information. First, the set of multi-scale texture feature maps is concatenated to obtain a concatenated feature map F_combined. Then, a series of convolutional and non-linear transformation operations are performed on F_combined. For example, first, the number of channels of F_combined is compressed through a convolutional layer to obtain an intermediate feature map F_middle. Then, a non-linear activation function (such as the sigmoid function) is applied to F_middle to map its values to between 0 and 1. The mapped feature map represents the weight distribution W of different spatial positions in the set of multi-scale texture feature maps. The closer the weight value is to 1, the more important the texture information at that position; the closer the weight value is to 0, the relatively less important the texture information at that position.
[0031] Step S1223: Perform weighted fusion processing on the set of multi-scale texture feature maps according to the weight distribution to obtain enhanced local texture features.
[0032] According to the generated weight distribution W, perform weighted fusion on the set of multi-scale texture feature maps {F1, F2,..., Fn}. For each texture feature map Fi (i = 1, 2,..., n), multiply it by the element at the corresponding position in the weight distribution W to obtain the weighted texture feature map Fi_weighted. Then, add all the weighted texture feature maps element by element to obtain the enhanced local texture feature F_local_enhanced. Through this weighted fusion method, the texture information at important positions in the multi-scale texture feature maps can be highlighted, and the interference information at unimportant positions can be suppressed, so as to obtain more representative local texture features.
[0033] Step S1224: Input the standardized video frame sequence into the deep convolutional module of the spatial feature encoder, and extract the global structural features of each video frame through dilated convolutional kernels. The global structural features include edge distribution features and region segmentation features.
[0034] The deep convolutional module processes the normalized video frame sequence using dilated convolutional kernels to extract global structural features. The dilated convolutional kernel introduces the concept of dilation rate on the basis of the ordinary convolutional kernel, which can expand the receptive field of the convolutional kernel without increasing the number of parameters, thereby capturing a larger range of structural information in the video frame. The normalized video frame V is input into the deep convolutional module, and convolutional operations are performed using dilated convolutional kernels. During the convolution process, the dilated convolutional kernel extracts features from different regions in the video frame and gradually generates a feature map reflecting the global structure of the video frame. Through further processing and analysis, edge distribution features and region segmentation features can be extracted from this feature map. The edge distribution features can represent the position and intensity information of the object edges in the video frame, while the region segmentation features can divide the video frame into different regions, each region representing different objects or scene elements.
[0035] Step S1225: Concatenate the enhanced local texture features and the global structural features to obtain the spatial correlation feature set.
[0036] Perform a concatenation operation on the enhanced local texture feature F_local_enhanced and the global structural feature F_global. Concatenation is to connect the two features in the channel dimension. Assuming the number of channels of F_local_enhanced is C1 and the number of channels of F_global is C2, the number of channels of the obtained spatial correlation feature set F_space after concatenation is C1 + C2. Through the concatenation operation, the local texture information and the global structural information of the video frame are integrated together to form a more comprehensive and representative spatial correlation feature set.
[0037] Step S123: Invoke the temporal feature encoder to perform cross-frame analysis on the normalized video frame sequence, extract the motion trajectory features and inter-frame change features between adjacent video frames, and merge the motion trajectory features and the inter-frame change features into the temporal dynamic feature set.
[0038] The temporal feature encoder is used to perform cross-frame analysis on the normalized video frame sequence to capture the motion and change information between adjacent video frames. The specific steps are as follows:
[0039] Step S1231: Divide the normalized video frame sequence into multiple video segments, and each video segment contains a preset number of consecutive video frames.
[0040] To facilitate the analysis of the relationship between adjacent video frames, the standardized video frame sequence is divided into multiple video segments according to a preset quantity. Assume that each video segment contains m consecutive video frames. Select m video frames from the standardized video frame sequence in sequence to form a video segment, and repeat this process continuously until the entire standardized video frame sequence is divided. For example, if the standardized video frame sequence has N video frames in total, then N / m video segments can be obtained (assuming N is divisible by m). The video frames in each video segment are continuous in time, and there is a certain motion and change relationship between them. Subsequent analysis will be carried out based on these video segments.
[0041] Step S1232: Perform optical flow field calculation processing on the consecutive video frames in each video segment to generate a set of pixel displacement vectors between adjacent video frames.
[0042] Optical flow field calculation is used to describe the motion of pixels between adjacent video frames. For the consecutive video frames in each video segment, an optical flow algorithm (such as the Lucas-Kanade algorithm) is used to calculate the optical flow field between adjacent video frames. The optical flow field reflects the displacement information of each pixel between adjacent video frames, that is, the direction and distance of the pixel moving from one video frame to the next. For two adjacent video frames Vt and Vt+1, the displacement vector of each pixel is calculated through the optical flow algorithm, and the displacement vectors of all pixels form the set of pixel displacement vectors Dt,t+1 between adjacent video frames. For all pairs of adjacent video frames in each video segment, such optical flow field calculations are performed to obtain a series of sets of pixel displacement vectors, and these sets record the motion trajectory information of the pixels in the video segment.
[0043] Step S1233: Input the set of pixel displacement vectors into the motion estimation module of the temporal feature encoder, and extract the motion trajectory features of the video segment in the temporal dimension through a three-dimensional convolutional kernel.
[0044] Input the set of pixel displacement vectors into the motion estimation module, and this motion estimation module processes the set of pixel displacement vectors using a three-dimensional convolutional kernel. The three-dimensional convolutional kernel not only considers the information in the spatial dimension but also the information in the temporal dimension. Through three-dimensional convolutional operations, the motion patterns and trajectory features of the video segment in the temporal dimension can be captured. Assume that the set of pixel displacement vectors D is a three-dimensional data structure that contains the pixel displacement information between multiple adjacent video frames. The three-dimensional convolutional kernel performs sliding convolution on D to extract features from the pixel displacement information at different temporal and spatial positions. After three-dimensional convolutional operations, the motion trajectory features F_motion of the video segment in the temporal dimension are obtained, and this motion trajectory feature reflects the motion trend and pattern of the object in the video segment.
[0045] Step S1234: Perform differential processing on consecutive video frames in each video segment to generate a sequence of inter-frame difference maps.
[0046] Perform differential processing on consecutive video frames in each video segment to obtain the change information between adjacent video frames. For two adjacent video frames Vt and Vt+1, calculate the difference in their corresponding pixel values to obtain an inter-frame difference map Dt,t+1. By performing such differential processing on all pairs of adjacent video frames in the video segment, a sequence of inter-frame difference maps {D1,2, D2,3,..., Dm-1,m} is obtained. The sequence of inter-frame difference maps reflects the pixel value changes between adjacent video frames in the video segment. These changes may be caused by factors such as object motion and lighting changes, providing a basis for subsequent extraction of inter-frame change features.
[0047] Step S1235: Input the sequence of inter-frame difference maps into the change detection module of the temporal feature encoder, and extract the inter-frame change features of the video segment in the temporal dimension through temporal convolutional kernels.
[0048] Input the sequence of inter-frame difference maps into the change detection module, which uses temporal convolutional kernels to process the sequence of inter-frame difference maps. Temporal convolutional kernels are specifically designed to process data with time series characteristics and can capture the change patterns and features of the video segment in the temporal dimension. Assume that the sequence of inter-frame difference maps is a two-dimensional image sequence. The temporal convolutional kernel slides and convolves on this image sequence to extract features from the inter-frame difference information at different time points. After the temporal convolution operation, the inter-frame change features F_change of the video segment in the temporal dimension are obtained. This inter-frame change feature reflects the dynamic changes of objects in the video segment.
[0049] Step S1236: Fuse the motion trajectory features and the inter-frame change features to obtain the set of temporal dynamic features.
[0050] Fuse the motion trajectory features F_motion and the inter-frame change features F_change to form a more comprehensive set of temporal dynamic features. The fusion method can be concatenation, that is, connecting F_motion and F_change in the channel dimension. Assume that the number of channels of F_motion is C3 and the number of channels of F_change is C4. The number of channels of the fused set of temporal dynamic features F_time is C3 + C4. Through the fusion operation, the motion trajectory information and the inter-frame change information of the objects in the video segment are integrated together to form a more representative set of temporal dynamic features.
[0051] Step S124: Among them, the spatial feature encoder and the temporal feature encoder are constructed by a joint training method, and the joint training method includes: taking the standardized video frame sequence as the input, and synchronously updating the parameters of the spatial feature encoder and the temporal feature encoder by minimizing the reconstruction error of the spatial correlation feature set and the prediction error of the temporal dynamic feature set.
[0052] When constructing the spatial feature encoder and the temporal feature encoder, a joint training method is adopted to improve their performance and collaborative working ability. The specific training process is as follows:
[0053] Take the standardized video frame sequence as the input and input it into the spatial feature encoder and the temporal feature encoder at the same time. The spatial feature encoder outputs the spatial correlation feature set F_space, and the temporal feature encoder outputs the temporal dynamic feature set F_time.
[0054] To evaluate the performance of the spatial feature encoder, the reconstruction error of the spatial correlation feature set is defined. First, use a decoder network to reconstruct the spatial correlation feature set F_space to obtain the reconstructed video frame V_reconstructed_space. Then, calculate the difference between the reconstructed video frame V_reconstructed_space and the original standardized video frame V_original as the reconstruction error E_space of the spatial correlation feature set.
[0055] To evaluate the performance of the temporal feature encoder, the prediction error of the temporal dynamic feature set is defined. The prediction error can be obtained by predicting the features of future video frames and then calculating the difference between the predicted features and the actual features. For example, according to the current temporal dynamic feature set F_time, predict the temporal dynamic feature set F_time_predicted at the next moment, and then calculate the difference between F_time_predicted and the actual temporal dynamic feature set F_time_next at the next moment as the prediction error E_time of the temporal dynamic feature set.
[0056] By minimizing the weighted sum of the reconstruction error E_space and the prediction error E_time, that is, E_total = α * E_space + β * E_time (where α and β are weight coefficients), use an optimization algorithm (such as the stochastic gradient descent algorithm) to synchronously update the parameters of the spatial feature encoder and the temporal feature encoder. In each iteration, adjust the parameters of the encoder according to the gradient information of the error, so that the reconstruction error and the prediction error gradually decrease, thereby improving the performance of the spatial feature encoder and the temporal feature encoder, and enabling them to better extract the spatial correlation features and temporal dynamic features of video frames.
[0057] Step S130: Perform anomaly detection on the spatial association feature set and the temporal dynamic feature set based on a preset anomaly detection model to generate an anomaly distribution feature set for the target area.
[0058] After obtaining the spatial association feature set and the temporal dynamic feature set, use a preset anomaly detection model to perform anomaly detection on them to determine whether there are anomalies in the target area and the distribution of the anomalies.
[0059] Step S131: Perform spatio-temporal alignment processing on the spatial association feature set and the temporal dynamic feature set to generate a spatio-temporal fusion feature map.
[0060] Since the spatial association feature set and the temporal dynamic feature set respectively reflect the spatial and temporal features of video frames, in order to perform anomaly detection more comprehensively, it is necessary to perform spatio-temporal alignment processing on them. First, it is necessary to adjust the dimensions of the spatial association feature set and the temporal dynamic feature set so that they have the same resolution and size in the spatial and temporal dimensions. Dimension adjustment can be achieved through operations such as interpolation, downsampling, or upsampling. Then, the adjusted spatial association feature set and temporal dynamic feature set are concatenated in the channel dimension to obtain a spatio-temporal fusion feature map F_spatio-temporal. The spatio-temporal fusion feature map combines the spatial and temporal information of video frames and provides richer data for subsequent anomaly detection.
[0061] Step S132: Input the spatio-temporal fusion feature map into the feature pyramid module of the anomaly detection model, and generate a multi-scale anomaly feature map set through multi-level downsampling and upsampling operations.
[0062] The feature pyramid module is used to perform multi-scale analysis on the spatio-temporal fusion feature map to capture anomaly information at different scales. Input the spatio-temporal fusion feature map F_spatio-temporal into the feature pyramid module, and the feature pyramid module will perform multi-level downsampling and upsampling operations.
[0063] During the multi-level downsampling process, use convolutional kernels and pooling layers to gradually downsample the spatio-temporal fusion feature map to obtain feature maps of different scales. For example, after the first downsampling operation, a feature map F_down1 with a smaller scale is obtained, and its size is smaller than the spatio-temporal fusion feature map F_spatio-temporal; after the second downsampling operation, a feature map F_down2 with an even smaller scale is obtained, and so on, to obtain a series of downsampled feature maps of different scales. These downsampled feature maps can capture the global information and macroscopic features at different scales in the spatio-temporal fusion feature map.
[0064] During the multi-level upsampling process, a transposed convolution layer or an interpolation method is used to gradually upsample the feature map obtained by downsampling so that it is restored to a size close to that of the spatio-temporal fusion feature map. For example, the feature map F_downn with the smallest scale is upsampled to obtain a feature map F_up1 with a slightly larger scale; then F_up1 is upsampled to obtain a feature map F_up2 with an even larger scale, until the size of the upsampled feature map is close to that of the spatio-temporal fusion feature map. The upsampling operation can restore the detailed information of the feature map, so that feature maps of different scales all contain rich feature information.
[0065] Through multi-level downsampling and upsampling operations, a multi-scale abnormal feature map set is finally generated. This multi-scale abnormal feature map set contains abnormal feature maps at different scales, such as {F_down1, F_down2,..., F_downn, F_up1, F_up2,..., F_upm}, where n and m represent the number of levels of downsampling and upsampling respectively. These multi-scale abnormal feature maps can reflect the abnormal information in the spatio-temporal fusion feature map from different perspectives, providing more comprehensive data support for subsequent anomaly detection.
[0066] Step S133: Perform cross-scale fusion processing on the multi-scale abnormal feature map set to obtain a fused abnormal probability map.
[0067] Select representative abnormal feature maps of different scales from the multi-scale abnormal feature map set for cross-scale fusion processing to synthesize the abnormal information at different scales and obtain a more accurate abnormal probability map.
[0068] Step S1331: Extract the first-scale abnormal feature map, the second-scale abnormal feature map, and the third-scale abnormal feature map from the multi-scale abnormal feature map set, where the first-scale abnormal feature map has the highest resolution and the third-scale abnormal feature map has the lowest resolution.
[0069] In the multi-scale abnormal feature map set, select the first-scale abnormal feature map F1, the second-scale abnormal feature map F2, and the third-scale abnormal feature map F3 according to the resolution. The first-scale abnormal feature map F1 has the highest resolution and can provide the most detailed abnormal information, such as the subtle features of objects in the target area and local abnormal conditions; the third-scale abnormal feature map F3 has the lowest resolution and mainly reflects the macroscopic abnormal information and global features of the target area; the resolution of the second-scale abnormal feature map F2 is between the first scale and the third scale and can provide abnormal information at the intermediate scale.
[0070] Step S1332: Upsample the third-scale anomaly feature map to the resolution of the second-scale anomaly feature map, adjust the channel dimension through a convolutional layer, and then perform element-wise addition with the second-scale anomaly feature map to obtain the first fused feature map.
[0071] First, perform an upsampling operation on the third-scale anomaly feature map F3 to make its resolution the same as that of the second-scale anomaly feature map F2. The upsampling operation can use methods such as bilinear interpolation or transposed convolution to enlarge the size of F3 to be the same as F2, obtaining the upsampled third-scale anomaly feature map F3_up.
[0072] Then, in order to effectively fuse F3_up and F2, it is necessary to adjust the channel dimension of F3_up through a convolutional layer so that its number of channels is the same as that of F2. Assume the number of channels of F2 is C2, and the feature map F3_up_adj with the same number of channels C2 is obtained after F3_up is processed by the convolutional layer.
[0073] Finally, perform element-wise addition on F3_up_adj and F2, that is, add the elements at their corresponding positions to obtain the first fused feature map F_merged1. The element-wise addition operation can initially fuse the anomaly information at different scales, making the first fused feature map contain both the intermediate-scale anomaly information of the second scale and the adjusted macroscopic anomaly information of the third scale.
[0074] Step S1333: Upsample the first fused feature map to the resolution of the first-scale anomaly feature map and perform feature concatenation with the first-scale anomaly feature map to obtain the second fused feature map.
[0075] Perform an upsampling operation on the first fused feature map F_merged1 to make its resolution the same as that of the first-scale anomaly feature map F1. Similarly, use methods such as bilinear interpolation or transposed convolution to enlarge the size of F_merged1 to be the same as F1, obtaining the upsampled first fused feature map F_merged1_up.
[0076] Next, perform feature concatenation on F_merged1_up and F1. Feature concatenation is to connect the two feature maps in the channel dimension. Assume the number of channels of F1 is C1 and the number of channels of F_merged1_up is C2. The number of channels of the concatenated second fused feature map F_merged2 is C1 + C2. Through the feature concatenation operation, the anomaly information at different scales is more comprehensively fused, making the second fused feature map contain both the detailed anomaly information of the first scale and the fused intermediate-scale and macroscopic-scale anomaly information.
[0077] Step S1334: Input the second fused feature map into the output layer of the anomaly detection model, and generate the fused anomaly probability map through an activation function, where each pixel value in the anomaly probability map represents the probability that the corresponding position belongs to the abnormal area.
[0078] The second fused feature map F_merged2 is input to the output layer of the anomaly detection model. The output layer typically contains one or more convolutional layers and an activation function. The convolutional layer further extracts and transforms the second fused feature map, mapping it to an appropriate feature space.
[0079] The output of the convolutional layer is then processed using an activation function (such as the sigmoid function), mapping the output values to a range between 0 and 1. The result is the fused anomaly probability map P. Each pixel value in the anomaly probability map P represents the probability that the corresponding location in the target area is an anomaly. The closer the pixel value is to 1, the more likely the location is an anomaly; the closer the pixel value is to 0, the more likely the location is a normal area.
[0080] Step S134: determining the boundary coordinates and the abnormal type label of the abnormal area in the target area according to the fused abnormal probability map, and associating the boundary coordinates and the abnormal type label into the abnormal distribution feature set.
[0081] After obtaining the fused anomaly probability map P, it is necessary to determine the boundary coordinates of the abnormal area and the anomaly type label in the target area to form an abnormal distribution feature set.
[0082] First, perform threshold processing on the abnormal probability map P. Set a suitable threshold T, mark the pixels in the abnormal probability map with pixel values greater than the threshold T as possible abnormal points, and mark the pixels less than or equal to the threshold T as normal points. Figure 2 The value of the abnormal point is 1, and the value of the normal point is 0.
[0083] Next, use an image segmentation algorithm (such as connected component analysis) to process the binary image B and identify all connected abnormal regions. For each connected abnormal region, calculate its boundary coordinates. This can be done by traversing the pixels in the abnormal region and finding the coordinates of its leftmost, rightmost, topmost, and bottommost pixels. This allows you to determine the bounding box of the abnormal region. The coordinates of the four vertices of the bounding box are the boundary coordinates of the abnormal region.
[0084] Next, in order to determine the anomaly type labels of the anomaly regions, a pre-trained classification model needs to be used. The original video frames or feature maps corresponding to each anomaly region are input into the classification model, and the classification model classifies according to the input data and outputs the possible anomaly type labels of the anomaly region, such as equipment failure, personnel violation, etc.
[0085] Finally, the boundary coordinates of each anomaly region are associated with the corresponding anomaly type labels to form an anomaly distribution feature set A. The anomaly distribution feature set A records the location and type information of all anomaly regions in the target region, providing an important basis for subsequent path optimization and anomaly handling.
[0086] Step S135: Among them, the anomaly detection model is constructed by an adversarial training method, and the adversarial training method includes: inputting the spatio-temporal fusion feature map of normal samples into a generator network to generate a reconstructed feature map, and calling a discriminator network to distinguish the reconstructed feature map from the feature map of real normal samples, and updating the parameters of the anomaly detection model by minimizing the reconstruction loss of the generator network and maximizing the discrimination error of the discriminator network.
[0087] The anomaly detection model is constructed by an adversarial training method and mainly consists of a generator network and a discriminator network.
[0088] During the training process, first, a large number of spatio-temporal fusion feature maps of normal samples are prepared. These spatio-temporal fusion feature maps of normal samples are input into the generator network. The role of the generator network is to try to reconstruct the input spatio-temporal fusion feature map to generate a reconstructed feature map. The generator network usually consists of multiple convolutional layers, deconvolutional layers, and non-linear activation functions, and through a series of feature transformation and mapping operations, it converts the input spatio-temporal fusion feature map into a reconstructed feature map.
[0089] At the same time, the reconstructed feature map and the feature map of real normal samples are input into the discriminator network. The task of the discriminator network is to distinguish whether the input feature map is the reconstructed feature map generated by the generator network or the feature map of real normal samples. The discriminator network also consists of multiple convolutional layers and non-linear activation functions, and through feature extraction and classification of the input feature map, it outputs a probability value indicating the probability that the input feature map is the feature map of real normal samples.
[0090] To evaluate the performance of the generator network, a reconstruction loss is defined. The reconstruction loss can be obtained by calculating the difference between the reconstructed feature map and the feature map of real normal samples, such as using the mean square error loss function. The goal of the generator network is to minimize the reconstruction loss, that is, to reconstruct the input spatio-temporal fusion feature map as accurately as possible.
[0091] To evaluate the performance of the discriminator network, the discrimination error is defined. The discrimination error can be obtained by calculating the classification error of the discriminator network for the reconstructed feature map and the feature map of the true normal samples. The goal of the discriminator network is to maximize the discrimination error, that is, to distinguish the reconstructed feature map and the feature map of the true normal samples as accurately as possible.
[0092] Adversarial training is achieved by alternately updating the parameters of the generator network and the discriminator network. In each iteration, first fix the parameters of the discriminator network, and update the parameters of the generator network according to the reconstruction loss, so that the generator network generates a reconstructed feature map closer to the feature map of the true normal samples; then fix the parameters of the generator network, and update the parameters of the discriminator network according to the discrimination error, so that the discriminator network can better distinguish the reconstructed feature map and the feature map of the true normal samples. Through continuous iterative training, the performance of the generator network and the discriminator network is continuously improved, and finally a well-performing anomaly detection model is obtained.
[0093] Step S140: Generate a dynamic path optimization instruction according to the abnormal distribution feature set and the real-time flight parameters of the UAV. The dynamic path optimization instruction is used to adjust the inspection path of the UAV to cover the area corresponding to the abnormal distribution feature set.
[0094] After obtaining the abnormal distribution feature set, combined with the real-time flight parameters of the UAV, a dynamic path optimization instruction is generated to adjust the inspection path of the UAV to ensure that the UAV can conduct a more detailed inspection of the abnormal area.
[0095] Step S141: Obtain the real-time flight parameters of the UAV. The real-time flight parameters include the current flight altitude, remaining flight time, and camera pitch angle.
[0096] The real-time flight parameters of the UAV are obtained through the sensors and flight control system of the UAV. The current flight altitude can be measured by an altitude sensor, which reflects the vertical distance of the UAV from the ground; the remaining flight time can be calculated by the battery power monitoring system and the flight power consumption model, which represents the time the UAV can still fly at the current battery level; the camera pitch angle can be measured by the attitude sensor of the camera, which represents the tilt angle of the camera relative to the horizontal direction.
[0097] Step S142: According to the boundary coordinates in the abnormal distribution feature set, combined with the field of view angle parameter of the imaging device, the current flight altitude, and the camera pitch angle, calculate the actual coverage area and spatial distribution density of the abnormal area.
[0098] Among them, the camera pitch angle is used to correct the three-dimensional projection relationship of the field of view coverage, and the current flight altitude is used to determine the actual ground resolution corresponding to a unit pixel.
[0099] First, according to the field of view angle parameter of the imaging device, the current flight altitude, and the camera pitch angle, calculate the field of view coverage range of the imaging device on the ground. The field of view angle parameter determines the horizontal and vertical angle ranges that the imaging device can capture. The current flight altitude determines the size of the field of view coverage range, and the camera pitch angle affects the projection shape of the field of view coverage range on the ground.
[0100] Then, according to the boundary coordinates in the abnormal distribution feature set, determine the positions of each abnormal area within the field of view coverage range. Convert the boundary coordinates of the abnormal area into actual ground coordinates, and combine the actual size of the field of view coverage range and the actual ground resolution corresponding to each unit pixel to calculate the actual coverage area of each abnormal area. The actual ground resolution corresponding to each unit pixel can be calculated based on the current flight altitude and the parameters of the imaging device, which represents the actual distance corresponding to each pixel point in the image on the ground.
[0101] Finally, according to the actual coverage areas of all abnormal areas and their distribution in the target area, calculate the spatial distribution density of the abnormal areas. The spatial distribution density can be represented by the ratio of the total area of the abnormal areas to the total area of the target area, which reflects the concentration degree of the abnormal areas in the target area.
[0102] Step S143: Based on the actual coverage area and the spatial distribution density, combined with the heading overlap rate constraint corresponding to the current flight altitude and the adjustable range of the camera pitch angle, generate a preliminary path adjustment plan.
[0103] The preliminary path adjustment plan includes the flight altitude adjustment amount determined based on the mapping relationship between the current flight altitude and the target detection accuracy, the flight speed adjustment amount based on the comparison of the remaining battery life and the range margin, and the camera angle adjustment amount of the optimal observation angle based on the camera pitch angle.
[0104] According to the actual coverage area and spatial distribution density of the abnormal areas, as well as the heading overlap rate constraint corresponding to the current flight altitude and the adjustable range of the camera pitch angle, formulate a preliminary path adjustment plan.
[0105] Determination of the flight altitude adjustment amount: According to the mapping relationship between the current flight altitude and the target detection accuracy, analyze the detection accuracy of the abnormal areas at different flight altitudes. If the spatial distribution density of the abnormal areas is large, in order to improve the detection accuracy, it may be necessary to reduce the flight altitude; if the spatial distribution density of the abnormal areas is small, the flight altitude can be appropriately increased to expand the inspection range. By analyzing this mapping relationship, determine the adjustment amount of the flight altitude so that the drone can better detect the abnormal areas at the new flight altitude.
[0106] Determination of the flight speed adjustment amount: Consider the comparison between the remaining flight time and the margin of the flight range. If the remaining flight time is sufficient and the actual coverage area of the abnormal area is large, the flight speed can be appropriately reduced to have more time for detailed detection of the abnormal area; if the remaining flight time is limited and the actual coverage area of the abnormal area is small, the flight speed can be appropriately increased to cover more areas within the limited time. Based on this comparison result, determine the adjustment amount of the flight speed.
[0107] Determination of the camera angle adjustment amount: Calculate the optimal observation angle of the camera according to the adjustable range of the camera's pitch angle and the position of the abnormal area. In order to clearly capture the abnormal area, it is necessary to adjust the pitch angle of the camera. By analyzing the spatial distribution of the abnormal area and the field of view angle of the camera, determine the adjustment amount of the camera angle so that the camera can capture the abnormal area at the best angle.
[0108] Step S144: Perform a constraint matching process on the preliminary path adjustment plan and the real-time flight parameters to determine whether the remaining flight time meets the execution requirements of the preliminary path adjustment plan.
[0109] The remaining flight time is used to calculate the maximum adjustable flight range, the current flight altitude is used to constrain the dynamic feasibility of altitude adjustment, and the camera pitch angle is used to verify the mechanical limit conditions of the angle adjustment mechanism.
[0110] To ensure the feasibility of the preliminary path adjustment plan, it is necessary to perform a constraint matching process with the real-time flight parameters of the UAV.
[0111] Step S1441: According to the flight altitude adjustment amount in the preliminary path adjustment plan, combine the vertical distance difference between the current flight altitude and the target altitude and the unit altitude energy consumption coefficient of the UAV's climb / descent to calculate the first energy consumption increment required for flight altitude adjustment.
[0112] According to the flight altitude adjustment amount in the preliminary path adjustment plan, determine the vertical distance difference between the current flight altitude and the target altitude. Then, combine the unit altitude energy consumption coefficient of the UAV's climb / descent to calculate the first energy consumption increment required for flight altitude adjustment. The unit altitude energy consumption coefficient represents the energy consumed by the UAV for each unit increase or decrease in altitude. By multiplying the vertical distance difference by the unit altitude energy consumption coefficient, the first energy consumption increment required for flight altitude adjustment is obtained.
[0113] Step S1442: According to the flight speed adjustment amount, combine the air density parameter corresponding to the current flight altitude and the thrust-power conversion curve of the UAV's power system to calculate the second energy consumption increment required for flight speed adjustment.
[0114] According to the flight speed adjustment amount in the preliminary path adjustment plan, combined with the air density parameter corresponding to the current flight altitude and the thrust-power conversion curve of the UAV power system, calculate the second energy consumption increment required for flight speed adjustment. The air density parameter affects the air resistance during the UAV flight. At different flight speeds and air densities, the thrust and power required by the UAV power system are also different. By querying the thrust-power conversion curve, according to the flight speed adjustment amount and the current air density parameter, calculate the additional power required after the flight speed adjustment, and then obtain the second energy consumption increment required for flight speed adjustment.
[0115] Step S1443: According to the camera angle adjustment amount, combined with the angle difference between the current angle value and the target angle value of the camera pitch angle, and the power consumption per unit angle adjustment of the servo motor, calculate the third energy consumption increment required for camera angle adjustment.
[0116] According to the camera angle adjustment amount in the preliminary path adjustment plan, determine the angle difference between the current angle value and the target angle value of the camera pitch angle. Then, combined with the power consumption per unit angle adjustment of the servo motor, calculate the third energy consumption increment required for camera angle adjustment. The power consumption per unit angle adjustment of the servo motor represents the energy consumed by the servo motor for each unit angle adjustment. By multiplying the angle difference by the power consumption per unit angle adjustment of the servo motor, the third energy consumption increment required for camera angle adjustment is obtained.
[0117] Step S1444: Based on the first energy consumption increment, the second energy consumption increment, and the third energy consumption increment being superimposed on the base power consumption of the current flight state, calculate the total power consumption rate of the preliminary path adjustment plan.
[0118] Superimpose the first energy consumption increment required for flight altitude adjustment, the second energy consumption increment required for flight speed adjustment, and the third energy consumption increment required for camera angle adjustment on the base power consumption of the current flight state to obtain the total power consumption rate of the preliminary path adjustment plan. The base power consumption represents the power consumption of the UAV in the current flight state (without path adjustment). By adding each energy consumption increment to the base power consumption, the adjusted total power consumption rate is obtained, which reflects the energy consumption situation of the UAV when executing the preliminary path adjustment plan.
[0119] Step S1445: Obtain the current remaining battery power of the UAV, and calculate the maximum executable duration according to the total power consumption rate and the remaining flight time.
[0120] Among them, the remaining flight time is obtained by dividing the current remaining battery power by the real-time total power consumption in the current flight state.
[0121] Obtain the current remaining power of the drone, and calculate the maximum executable duration based on the total power consumption rate and the remaining flight duration of the preliminary path adjustment plan. The remaining flight duration is obtained by dividing the current remaining power by the real-time total power consumption in the current flight state. According to the total power consumption rate and the current remaining power, the maximum duration that the drone can continuously fly when executing the preliminary path adjustment plan can be calculated, that is, the maximum executable duration.
[0122] Step S1446: Compare the estimated execution time of the preliminary path adjustment plan with the maximum executable duration.
[0123] If the estimated execution time is less than or equal to the maximum executable duration, it is determined that the execution requirement of the preliminary path adjustment plan is met; if the estimated execution time is greater than the maximum executable duration, it is determined that the execution requirement of the preliminary path adjustment plan is not met.
[0124] According to the specific parameters of flight altitude adjustment, flight speed adjustment, and camera angle adjustment in the preliminary path adjustment plan, estimate the time required to execute this plan to obtain the estimated execution time of the preliminary path adjustment plan. Compare this estimated execution time with the maximum executable duration. If the estimated execution time is less than or equal to the maximum executable duration, it means that the current remaining power and flight endurance of the drone can support it to execute the preliminary path adjustment plan, and it is determined that the execution requirement of the preliminary path adjustment plan is met; if the estimated execution time is greater than the maximum executable duration, it means that the current power and flight endurance of the drone are insufficient to complete the preliminary path adjustment plan, and it is determined that the execution requirement of the preliminary path adjustment plan is not met.
[0125] Step S145: If it is satisfied, directly use the preliminary path adjustment plan as the dynamic path optimization instruction.
[0126] When it is determined that the execution requirement of the preliminary path adjustment plan is met, it indicates that the plan is feasible under the constraints of the drone's power, flight performance, etc. At this time, directly use the preliminary path adjustment plan as the dynamic path optimization instruction. This dynamic path optimization instruction includes specific information such as the flight altitude adjustment amount, flight speed adjustment amount, and camera angle adjustment amount. The flight control system of the drone will make corresponding adjustments to the flight parameters according to this information, so that the drone can conduct inspections on the abnormal area according to the optimized path. For example, if the flight altitude adjustment amount is to decrease a certain altitude, the flight control system will control the drone to descend to the target altitude; if the flight speed adjustment amount is to slow down the speed, the flight control system will reduce the flight speed of the drone; if the camera angle adjustment amount is to adjust to a certain specific angle, the flight control system will control the camera to adjust to this angle, so as to ensure that the drone can more effectively cover the area corresponding to the abnormal distribution feature set.
[0127] Step S146: If not satisfied, downgrade the preliminary path adjustment plan according to the remaining battery life to generate an optimized dynamic path optimization instruction.
[0128] Among them, the downgrading process includes preferentially retaining the connectivity area coverage rate index according to the spatial topological distribution characteristics of the abnormal area, reducing the scanning range of non-critical areas, and recalculating the feasible coverage density based on the current flight altitude and camera pitch angle.
[0129] When it is determined that the execution requirements of the preliminary path adjustment plan are not met, it is necessary to downgrade the preliminary path adjustment plan to ensure that the abnormal area can still be inspected as effectively as possible under the limitation of the remaining battery life of the UAV.
[0130] First, analyze the spatial topological distribution characteristics of the abnormal area. The spatial topological distribution characteristics describe information such as the relative position, connection relationship, and distribution form of the abnormal area in the target area. By analyzing these characteristics, the connected areas in the abnormal area can be identified, that is, the set of interconnected abnormal areas. The connectivity area coverage rate index reflects the coverage degree of these connected abnormal areas during the UAV inspection process and is an important inspection effect evaluation index.
[0131] During the downgrading process, the connectivity area coverage rate index is preferentially retained. This means ensuring that the UAV covers as much of the connected part of the abnormal area as possible within the limited battery life. To achieve this goal, it is necessary to reduce the scanning range of non-critical areas. Non-critical areas refer to those areas that have little impact on the overall abnormal detection result or are weakly associated with the connected abnormal areas. By reducing the scanning of these areas, the power and flight time of the UAV can be saved.
[0132] At the same time, recalculate the feasible coverage density based on the current flight altitude and camera pitch angle. The feasible coverage density represents the ratio of the area of the abnormal area that the UAV can effectively cover to the total area of the target area under the current flight conditions. After reducing the scanning range of non-critical areas, it is necessary to recalculate the feasible coverage density according to the new scanning range and current flight parameters to evaluate the inspection effect after downgrading.
[0133] After the above downgrading process, an optimized dynamic path optimization instruction is generated. This dynamic path optimization instruction optimizes the preliminary path adjustment plan on the premise of considering the remaining battery life of the UAV, enabling the UAV to cover the key parts of the abnormal area as comprehensively as possible with limited resources, and improving the efficiency and accuracy of abnormal detection.
[0134] Step S150: Call the incremental learning module to update the parameters of the abnormal detection model based on the abnormal distribution feature set.
[0135] To enable the anomaly detection model to continuously adapt to newly emerging anomaly situations and improve its detection performance, it is necessary to call the incremental learning module to update the parameters of the anomaly detection model based on the anomaly distribution feature set.
[0136] Step S151: Extract target anomaly samples from the anomaly distribution feature set, where the target anomaly samples include the spatio-temporal fusion feature map generated by performing spatio-temporal alignment processing on the spatial association feature set and the temporal dynamic feature set of the anomaly region.
[0137] In the anomaly distribution feature set, detailed information of all anomaly regions in the target area is included, such as boundary coordinates, anomaly type labels, etc. To update the anomaly detection model, it is necessary to extract target anomaly samples from it. The target anomaly samples are the spatio-temporal fusion feature maps generated by performing spatio-temporal alignment processing on the spatial association feature set and the temporal dynamic feature set of the anomaly region. First, according to the boundary coordinates in the anomaly distribution feature set, the feature information of the corresponding anomaly regions is extracted from the original spatial association feature set and temporal dynamic feature set. Then, spatio-temporal alignment processing is performed on the spatial association features and temporal dynamic features of these anomaly regions. The specific method is the same as the spatio-temporal alignment processing process described in step S131, that is, matching and integrating them in the spatial and temporal dimensions, and finally obtaining the spatio-temporal fusion feature map of the target anomaly samples.
[0138] Step S152: Input the target anomaly samples into the feature buffer of the incremental learning module for temporary storage until the number of samples in the feature buffer reaches a preset threshold.
[0139] The incremental learning module includes a feature buffer for temporarily storing newly obtained target anomaly samples. The extracted target anomaly samples are input into the feature buffer for temporary storage. The feature buffer will manage and store the input samples and record the relevant information of the samples. The preset threshold is a pre-set standard for the number of samples. When the number of samples in the feature buffer reaches this threshold, it means that enough new anomaly samples have been accumulated and a model update operation can be performed. Before the number of samples reaches the preset threshold, the feature buffer will continuously receive new target anomaly samples and store them.
[0140] Step S153: Randomly sample a batch of training samples from the feature buffer and obtain the corresponding batch in the historical training sample set.
[0141] When the number of samples in the feature buffer reaches a preset threshold, a batch of training samples is randomly sampled from the feature buffer. The purpose of random sampling is to ensure the randomness and diversity of the samples and avoid the adverse effects of sample bias on model updating. The batch of training samples is a part of the target abnormal samples selected from the feature buffer, and its quantity is determined according to the actual requirements and the requirements of model training.
[0142] Meanwhile, obtain the corresponding batch in the historical training sample set. The historical training sample set is the sample set used in the previous model training process and contains information on various normal and abnormal samples. The corresponding batch refers to a sample set with the same quantity and similar feature distribution as the batch of training samples sampled from the feature buffer. This part of the samples is selected from the historical training sample set and used for joint training together with the new batch of training samples.
[0143] Step S154: Input the batch of training samples and the corresponding batch in the historical training samples into the anomaly detection model for joint training processing, and calculate the current model loss.
[0144] For example, step S1541: Perform feature enhancement processing on the target abnormal samples in the batch of training samples to generate an enhanced abnormal sample set. The feature enhancement processing includes spatial rotation, brightness transformation, and temporal jitter.
[0145] To increase the diversity and robustness of the training samples, feature enhancement processing is performed on the target abnormal samples in the batch of training samples. The feature enhancement processing includes operations such as spatial rotation, brightness transformation, and temporal jitter. Spatial rotation means rotating the spatio-temporal fusion feature map of the target abnormal sample in the spatial dimension to simulate abnormal situations at different angles. By randomly selecting the rotation angle, each target abnormal sample is rotated to obtain the rotated feature map. Brightness transformation means adjusting the brightness value of the target abnormal sample to simulate abnormal situations under different lighting conditions. The brightness of each target abnormal sample can be changed by randomly adjusting the brightness coefficient to obtain the feature map after brightness transformation. Temporal jitter means making a small offset of the target abnormal sample in the time dimension to simulate the uncertainty of the abnormal occurrence time. By randomly selecting the time offset, the time series of each target abnormal sample is adjusted to obtain the feature map after temporal jitter. After these feature enhancement processes, an enhanced abnormal sample set is generated. The number of samples in this enhanced abnormal sample set is the same as the number of samples in the batch of training samples, but each sample has undergone different degrees of feature enhancement.
[0146] Step S1542: Merge the enhanced abnormal sample set and the corresponding batch in the historical training samples into mixed training data.
[0147] The generated enhanced abnormal sample set is merged with the corresponding batches obtained from the historical training sample set to form mixed training data. The merging operation is to splice the two sample sets in the sample dimension, that is, each sample in the enhanced abnormal sample set is arranged in sequence with the samples in the corresponding batches of the historical training samples to obtain a new sample set, which is the mixed training data. The mixed training data contains new target abnormal samples and historical training samples, and can provide richer information for the update of the abnormal detection model.
[0148] Step S1543: Invoke the abnormal detection model to perform forward propagation processing on the mixed training data to obtain a predicted abnormal probability distribution.
[0149] The mixed training data is input into the abnormal detection model for forward propagation processing. The abnormal detection model will perform feature extraction and classification prediction on the input mixed training data according to its internal network structure and parameters. After inputting the mixed training data, the model will be processed layer by layer, such as convolutional layer, pooling layer, fully connected layer, etc., and finally output a predicted abnormal probability distribution. The predicted abnormal probability distribution represents the probability estimation of the abnormal detection model for each sample in the mixed training data belonging to the abnormal category. Each sample corresponds to a probability value, which reflects the likelihood that the model believes the sample is an abnormal sample.
[0150] Step S1544: Calculate the classification loss according to the difference between the predicted abnormal probability distribution and the true abnormal label.
[0151] The true abnormal label is the actual abnormal category information of each sample in the mixed training data, which is known. The predicted abnormal probability distribution is compared with the true abnormal label to calculate the difference between them to obtain the classification loss. The classification loss is used to measure the error degree between the prediction result of the model and the actual situation. Common classification loss functions include cross-entropy loss function, etc. By inputting the predicted abnormal probability distribution and the true abnormal label into the classification loss function, a loss value is calculated, which is the classification loss. The smaller the classification loss, the closer the prediction result of the model is to the actual situation, and the better the performance of the model.
[0152] Step S1545: Calculate the current model loss according to the classification loss and the regularization term, where the regularization term is used to constrain the parameter change amplitude of the abnormal detection model.
[0153] To prevent the overfitting phenomenon from occurring in the anomaly detection model during the update process, it is necessary to introduce a regularization term. The regularization term is an additional penalty term used to constrain the variation amplitude of the model parameters. Common regularization methods include L1 regularization and L2 regularization, etc. When calculating the current model loss, the classification loss and the regularization term are weighted and summed to obtain the current model loss. By adjusting the weight of the regularization term, the variation degree of the model parameters can be controlled. The current model loss comprehensively considers the classification error of the model and the stability of the parameters, and can more comprehensively evaluate the performance of the model.
[0154] Step S155: Update the parameters of the anomaly detection model based on the current model loss, and synchronously update the sample weight distribution in the feature buffer.
[0155] According to the current model loss, use an optimization algorithm (such as the stochastic gradient descent algorithm) to update the parameters of the anomaly detection model. The optimization algorithm will adjust the model parameters according to the gradient information of the current model loss, so that the model loss gradually decreases. In each iteration, the optimization algorithm will calculate the gradient of the current model loss with respect to the model parameters, and then update the model parameters according to the direction and magnitude of the gradient. Through continuous iterative updates, the model parameters will gradually converge to an optimal value, improving the performance of the model.
[0156] At the same time, synchronously update the sample weight distribution in the feature buffer. The sample weight distribution represents the importance of each sample in the feature buffer during model training. During the model update process, different samples may have different impacts on the model. Therefore, it is necessary to adjust the weight distribution of the samples according to the model update situation and the characteristics of the samples. For example, for those samples that play an important role in model update, their weights can be increased; for those samples that have less impact on model update, their weights can be decreased. By updating the sample weight distribution, the sample information in the feature buffer can be better utilized, improving the model update efficiency and performance.
[0157] During the implementation process of the above embodiments, in the data acquisition stage, the unmanned aerial vehicle carries a camera device to collect video data of the target area. To avoid infringing on others' privacy, the use of the camera device will be strictly limited within the target area, and the scope, method, and purpose of data acquisition will be clearly publicized in advance. For example, in the industrial park inspection scenario, only public areas such as buildings and equipment within the park will be photographed, and the camera will not be aimed at the private spaces of the personnel in the park, such as the interior of the employee dormitory. At the same time, strict security storage measures will be taken for the collected video data, access permissions will be set, and data leakage will be prevented to protect the security of the relevant information involved.
[0158] In terms of label management, when annotating abnormal types for abnormal areas, objective and fair annotation rules need to be formulated. Annotators will undergo professional training to ensure that they can annotate according to clear standards, avoiding inaccurate or discriminatory labels caused by personal subjective factors or biases. For example, when judging the type of equipment failure, accurate annotation will be carried out based on the actual operating parameters and failure manifestations of the equipment, and unreasonable label settings will not be made due to factors such as the brand and manufacturer of the equipment. Moreover, multiple review processes will be set for the labels to ensure the accuracy and fairness of the labels.
[0159] In the rule setting process, whether it is the training rules of the anomaly detection model or the formulation of path optimization rules, fairness and reasonableness will be the principles. In the training of the anomaly detection model, the normal samples and abnormal samples used will be widely representative, covering all possible situations within the target area, and no samples will be excluded due to certain specific factors to ensure the generalization ability and fairness of the model. In setting the path optimization rules, various factors such as the actual situation of the abnormal area and the flight performance of the drone will be comprehensively considered, and no unreasonable preference or neglect will be given to certain areas to ensure that each abnormal area has a reasonable opportunity to be inspected.
[0160] In terms of recommendation decisions, the generated dynamic path optimization instructions will be based on objective data and reasonable algorithms. No excessive attention or neglect will be given to certain abnormal areas due to human preferences or unreasonable factors. For example, when deciding whether to conduct key inspections on a certain abnormal area, comprehensive judgments will be made based on factors such as the actual coverage area, spatial distribution density of the abnormal area, and its impact on the overall target area, rather than based on other irrelevant factors. And the results of the recommendation decisions will be evaluated and adjusted regularly to ensure the fairness and reasonableness of the decisions.
[0161] For the processing of privacy-sensitive data, a variety of privacy protection and anti-leakage technical means are adopted throughout the process. In the data collection stage, the collected video data will be anonymized to remove any personal identity information it may contain. In terms of data storage, encryption technology will be used to encrypt and store the data to prevent it from being illegally obtained during storage. During the data transmission process, a secure transmission protocol will be used to encrypt and transmit the data to prevent it from being stolen or tampered with during transmission. At the same time, a strict access control mechanism needs to be established so that only authorized personnel can access and process this data, and detailed records and audits of access operations will be carried out to enable traceability and investigation in case of problems.
[0162] Figure 2FIG. 0 shows a schematic diagram of exemplary hardware and software components of a video stream analysis system 100 based on drone inspection that can implement the ideas of the present application. For example, the processor 120 can be used on the video stream analysis system 100 based on drone inspection and is used to execute the functions in the present application.
[0163] The video stream analysis system 100 based on drone inspection can be a general-purpose server or a special-purpose server, both of which can be used to implement the video stream analysis method based on drone inspection of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0164] For example, the video stream analysis system 100 based on drone inspection can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the video stream analysis system 100 based on drone inspection can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The video stream analysis system 100 based on drone inspection also includes an I / O interface 150 between the computer and other input / output devices.
[0165] For ease of explanation, only one processor is described in the video stream analysis system 100 based on drone inspection. However, it should be noted that the video stream analysis system 100 in the present application can also include multiple processors. Therefore, the steps executed by one processor described in the present application can also be jointly executed or separately executed by multiple processors. For example, if the processor of the video stream analysis system 100 based on drone inspection executes step A and step B, it should be understood that step A and step B can also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.
[0166] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned video stream analysis method based on drone inspection is implemented.
[0167] It should be noted that, in order to simplify the expression of the present invention disclosure and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.
Claims
1. A video stream analysis method based on drone inspection, characterized in that The method includes: Collecting original video stream data of a target area through a camera device carried by a drone, where the original video stream data includes a plurality of consecutive video frame sequences; Performing video feature extraction on the original video stream data to obtain a spatial association feature set and a temporal dynamic feature set of the video frame sequences; Performing anomaly detection on the spatial association feature set and the temporal dynamic feature set based on a preset anomaly detection model to generate an anomaly distribution feature set of the target area; Generating a dynamic path optimization instruction according to the anomaly distribution feature set and the real-time flight parameters of the drone, where the dynamic path optimization instruction is used to adjust the inspection path of the drone to cover the area corresponding to the anomaly distribution feature set; Invoking an incremental learning module to update the parameters of the anomaly detection model based on the anomaly distribution feature set. The performing video feature extraction on the original video stream data to obtain a spatial association feature set and a temporal dynamic feature set of the video frame sequences includes: Performing preprocessing operations on the original video stream data to generate a standardized video frame sequence, where the preprocessing operations include geometric distortion correction, illumination equalization processing, and noise filtering processing; Invoking a spatial feature encoder to perform frame-by-frame analysis processing on the standardized video frame sequence, extracting local texture features and global structure features of each video frame in the video frame sequence, and combining the local texture features and the global structure features into the spatial association feature set; Invoking a temporal feature encoder to perform cross-frame analysis processing on the standardized video frame sequence, extracting motion trajectory features and inter-frame change features between adjacent video frames, and combining the motion trajectory features and the inter-frame change features into the temporal dynamic feature set; Wherein, the spatial feature encoder and the temporal feature encoder are constructed through a joint training method, and the joint training method includes: using the standardized video frame sequence as an input, and synchronously updating the parameters of the spatial feature encoder and the temporal feature encoder by minimizing the reconstruction error of the spatial association feature set and the prediction error of the temporal dynamic feature set.
2. The video stream analysis method based on drone inspection according to claim 1, wherein The invoking a spatial feature encoder to perform frame-by-frame analysis processing on the standardized video frame sequence, extracting local texture features and global structure features of each video frame in the video frame sequence, and combining the local texture features and the global structure features into the spatial association feature set includes: Inputting the standardized video frame sequence into a shallow convolutional module of the spatial feature encoder, and extracting local texture features of each video frame through multi-scale convolutional kernels to obtain a multi-scale texture feature map set; Inputting the multi-scale texture feature map set into a spatial attention module of the spatial feature encoder to generate a weight distribution of different spatial positions in the multi-scale texture feature map set; Performing weighted fusion processing on the multi-scale texture feature map set according to the weight distribution to obtain enhanced local texture features; Input the standardized video frame sequence into the deep convolutional module of the spatial feature encoder, and extract the global structure features of each video frame through dilated convolutional kernels. The global structure features include edge distribution features and region segmentation features; Concatenate the enhanced local texture features and the global structure features to obtain the spatial correlation feature set.
3. The video stream analysis method based on drone inspection according to claim 1, characterized in that Call the temporal feature encoder to perform cross-frame analysis on the standardized video frame sequence, extract the motion trajectory features and inter-frame change features between adjacent video frames, and merge the motion trajectory features and the inter-frame change features into the temporal dynamic feature set, including: Divide the standardized video frame sequence into multiple video segments, and each video segment contains a preset number of consecutive video frames; Perform optical flow field calculation on the consecutive video frames in each video segment to generate a set of pixel displacement vectors between adjacent video frames; Input the set of pixel displacement vectors into the motion estimation module of the temporal feature encoder, and extract the motion trajectory features of the video segment in the time dimension through 3D convolutional kernels; Perform differential processing on the consecutive video frames in each video segment to generate a sequence of inter-frame difference maps; Input the sequence of inter-frame difference maps into the change detection module of the temporal feature encoder, and extract the inter-frame change features of the video segment in the time dimension through temporal convolutional kernels; Fuse the motion trajectory features and the inter-frame change features to obtain the temporal dynamic feature set.
4. The video stream analysis method based on drone inspection according to claim 1, characterized in that Based on a preset anomaly detection model, perform anomaly detection on the spatial correlation feature set and the temporal dynamic feature set to generate an anomaly distribution feature set of the target region, including: Perform spatio-temporal alignment processing on the spatial correlation feature set and the temporal dynamic feature set to generate a spatio-temporal fusion feature map; Input the spatio-temporal fusion feature map into the feature pyramid module of the anomaly detection model, and generate a set of multi-scale anomaly feature maps through multi-level downsampling and upsampling operations; Perform cross-scale fusion processing on the set of multi-scale anomaly feature maps to obtain a fused anomaly probability map; Determine the boundary coordinates and anomaly type labels of the anomaly regions in the target region according to the fused anomaly probability map, and associate the boundary coordinates and the anomaly type labels as the anomaly distribution feature set; Among them, the anomaly detection model is constructed by an adversarial training method. The adversarial training method includes: inputting the spatio-temporal fusion feature map of normal samples into the generator network to generate a reconstructed feature map, and calling the discriminator network to distinguish the reconstructed feature map from the feature map of real normal samples, and updating the parameters of the anomaly detection model by minimizing the reconstruction loss of the generator network and maximizing the discrimination error of the discriminator network.
5. The method for video stream analysis based on drone inspection according to claim 4, wherein The performing cross-scale fusion processing on the set of multi-scale anomaly feature maps to obtain a fused anomaly probability map includes: Extract the first-scale anomaly feature map, the second-scale anomaly feature map, and the third-scale anomaly feature map from the set of multi-scale anomaly feature maps, where the first-scale anomaly feature map has the highest resolution and the third-scale anomaly feature map has the lowest resolution; Upsample the third-scale anomaly feature map to the resolution of the second-scale anomaly feature map, adjust the channel dimension through a convolutional layer, and perform element-wise addition with the second-scale anomaly feature map to obtain a first fused feature map; Upsample the first fused feature map to the resolution of the first-scale anomaly feature map and perform feature concatenation with the first-scale anomaly feature map to obtain a second fused feature map; Input the second fused feature map into the output layer of the anomaly detection model, and generate the fused anomaly probability map through an activation function. Each pixel value in the anomaly probability map represents the probability that the corresponding position belongs to the anomaly region.
6. The method for video stream analysis based on drone inspection according to claim 1, wherein The generating of the dynamic path optimization instruction according to the anomaly distribution feature set and the real-time flight parameters of the drone includes: Obtain the real-time flight parameters of the drone, where the real-time flight parameters include the current flight altitude, the remaining flight duration, and the camera pitch angle; According to the boundary coordinates in the anomaly distribution feature set, combine the field of view angle parameter of the imaging device, the current flight altitude, and the camera pitch angle to calculate the actual coverage area and spatial distribution density of the anomaly region. The camera pitch angle is used to correct the three-dimensional projection relationship of the field of view coverage range, and the current flight altitude is used to determine the actual ground resolution corresponding to a unit pixel; Based on the actual coverage area and the spatial distribution density, combine the heading overlap rate constraint corresponding to the current flight altitude and the adjustable range of the camera pitch angle to generate a preliminary path adjustment plan. The preliminary path adjustment plan includes a flight altitude adjustment amount determined based on the mapping relationship between the current flight altitude and the target detection accuracy, a flight speed adjustment amount based on the comparison between the remaining flight duration and the range margin, and a camera angle adjustment amount of the optimal observation angle based on the camera pitch angle; Perform constraint matching processing on the preliminary path adjustment plan and the real-time flight parameters to determine whether the remaining flight duration meets the execution requirements of the preliminary path adjustment plan. The remaining flight duration is used to calculate the maximum adjustable range, the current flight altitude is used to constrain the dynamic feasibility of altitude adjustment, and the camera pitch angle is used to verify the mechanical limit conditions of the angle adjustment mechanism; If it is satisfied, directly use the preliminary path adjustment plan as the dynamic path optimization instruction; If it is not satisfied, downgrade the preliminary path adjustment plan according to the remaining flight duration to generate an optimized dynamic path optimization instruction. The downgrading process includes, according to the spatial topological distribution characteristics of the anomaly region, preferentially retaining the connected region coverage rate index, reducing the scanning range of non-critical regions, and recalculating the feasible coverage density based on the current flight altitude and the camera pitch angle.
7. The video stream analysis method based on drone inspection according to claim 6, wherein Constraining and matching the preliminary path adjustment scheme with the real-time flight parameters to determine whether the remaining flight time meets the execution requirements of the preliminary path adjustment scheme includes: According to the flight altitude adjustment amount in the preliminary path adjustment scheme, combining the vertical distance difference between the current flight altitude and the target altitude and the unit altitude energy consumption coefficient of the UAV's climb / descent, calculate the first energy consumption increment required for flight altitude adjustment; According to the flight speed adjustment amount, combining the air density parameter corresponding to the current flight altitude and the thrust-power conversion curve of the UAV's power system, calculate the second energy consumption increment required for flight speed adjustment; According to the camera angle adjustment amount, combining the angle difference between the current angle value and the target angle value of the camera's pitch angle and the power consumption per unit angle adjustment of the servo motor, calculate the third energy consumption increment required for camera angle adjustment; Based on the superposition of the first energy consumption increment, the second energy consumption increment, and the third energy consumption increment on the basic power consumption of the current flight state, calculate the total power consumption rate of the preliminary path adjustment scheme; Obtain the current remaining battery power of the UAV, and calculate the maximum executable duration according to the total power consumption rate and the remaining flight time, where the remaining flight time is obtained by dividing the current remaining battery power by the real-time total power consumption in the current flight state; Compare the estimated execution time of the preliminary path adjustment scheme with the maximum executable duration: If the estimated execution time is less than or equal to the maximum executable duration, it is determined that the execution requirements of the preliminary path adjustment scheme are met; If the estimated execution time is greater than the maximum executable duration, it is determined that the execution requirements of the preliminary path adjustment scheme are not met.
8. The video stream analysis method based on drone inspection according to claim 1, characterized in that, The calling of the incremental learning module to update the parameters of the anomaly detection model based on the set of anomaly distribution features includes: Extract target anomaly samples from the set of anomaly distribution features, where the target anomaly samples include a spatio-temporal fusion feature map generated by spatio-temporal alignment processing of the spatial association feature set and the time dynamic feature set of the anomaly region; Temporarily store the target anomaly samples in the feature buffer of the incremental learning module until the number of samples in the feature buffer reaches a preset threshold; Randomly sample a batch of training samples from the feature buffer and obtain the corresponding batch in the historical training sample set; Input the batch of training samples and the corresponding batch in the historical training samples into the anomaly detection model for joint training processing, and calculate the current model loss; Update the parameters of the anomaly detection model based on the current model loss, and synchronously update the sample weight distribution in the feature buffer.
9. A video stream analysis system based on drone patrol, characterized in that, The video stream analysis system based on UAV inspection includes a processor and a memory. The memory is connected to the processor. The memory is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the memory to implement the video stream analysis method based on UAV inspection according to any one of claims 1-8 above.
Citation Information
Patent Citations
Three-dimensional reconstruction-based unmanned aerial vehicle image stitching method and system
CN108765298A
Transformer substation unmanned aerial vehicle inspection path planning method
CN118518103A