Power distribution network equipment hidden danger detection method based on large-view-field spliced video data and related device

Through the drone collecting multi-view video data and performing large-field splicing processing, combining video processing technology and hidden danger detection model, the problems of low manual data acquisition efficiency and low image quality in the existing power inspection technology are solved, and efficient and accurate hidden danger detection of distribution network equipment is achieved.

CN120013921AActive Publication Date: 2025-05-16ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD

Patent Information

Application Number
CN202510155595.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-16
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The existing drone power inspection technology requires manual data collection multiple times, resulting in large workload, low efficiency, and susceptible to environmental interference, resulting in low image quality.

Method used

Using a method based on large field of view stitching video data, multi-view video data is collected through drones, and spatial and temporal parameters are aligned using video processing technology, video frame image features are extracted for matching, stitching video data is fused, and pre-trained hidden danger detection model is used for automatic detection.

Benefits of technology

It realizes efficient and accurate detection of hidden dangers for distribution network equipment, reduces manual workload, improves detection speed and efficiency, reduces false detection and missed detection, and enhances the reliability and objectivity of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013921A_ABST
    Figure CN120013921A_ABST
Patent Text Reader

Abstract

According to the power distribution network equipment hidden danger detection method based on the large-view-field spliced video data and the related device provided by the invention, the video data under different view angles are combined, so that the detection view field range is remarkably expanded, and the problem of repeated shooting of multiple pieces of power distribution network equipment at the same position is solved through the large view field. And in combination with the synchronization timestamp, accurate alignment of space and time parameters is performed on the multi-channel video data, so that the continuity and consistency of the spliced video in time and space are ensured, and the reliability and effectiveness of a detection result are improved. Through accurate extraction and matching of multi-channel video frame image features, visual distortion and seam traces in the splicing process are reduced, high-quality fused video data are generated, and a clear and continuous visual basis is provided for subsequent hidden danger detection. And the pre-trained hidden danger detection model is used for automatically detecting the fused video data, so that the detection speed and the working efficiency are greatly improved, and the objectivity and the standardization of the detection process are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electric power inspection, and in particular relates to a method and a related device for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data. Background Art

[0002] The safe and stable operation of distribution network equipment is an important part of the normal operation of the power system. Therefore, inspection of distribution network equipment is essential. With the rapid progress of drone technology, power inspection methods based on drones have begun to be widely promoted.

[0003] Nowadays, the commonly used drone inspection technology needs to rely on manual data collection and identification of hidden dangers. Multiple distribution network devices at the same location need to collect data multiple times. The surge in data volume will also greatly increase the workload of manual identification. In addition, drone flight and camera shooting are easily disturbed by the surrounding environment, and often collect out-of-focus and jittery image data, which will also cause a single distribution network device to need to perform multiple refined shots to obtain usable high-quality image data. This not only requires high technical requirements for operators, but also leads to a further increase in inspection workload, thereby reducing the efficiency of power inspection.

[0004] Therefore, developing a distribution network inspection method that can obtain high-quality distribution network equipment image data and reduce the workload of inspection personnel to ensure the efficient operation of distribution network equipment hidden danger inspection has become a key issue that needs to be urgently solved in the power industry. Summary of the invention

[0005] In view of this, the present invention aims to provide a distribution network equipment hidden danger detection method and related devices based on large-field-of-view stitching video data. By comprehensively utilizing multi-view video data collected by drones and advanced video processing technology, efficient detection of distribution network equipment hidden dangers based on large-viewing-angle video data is achieved, thereby improving the accuracy of detection and the convenience of operation.

[0006] In order to achieve the above object, the technical solution provided by the present invention is as follows:

[0007] In a first aspect, the present invention provides a method for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data, comprising the following steps:

[0008] Acquire multi-channel distribution network equipment video data with synchronized timestamps from different viewing angles;

[0009] According to the perspective differences of the video data of multiple distribution network devices, the spatial and temporal parameters of the video data of multiple distribution network devices are aligned in combination with the synchronization timestamp;

[0010] Based on the aligned multi-channel distribution network equipment video data, the features of the multi-channel video frame images are extracted and matched to obtain the splicing relationship between the multi-channel video frame images;

[0011] Determine a stitching seam between two video frame images to be stitched based on the stitching relationship;

[0012] Based on the splicing relationship and splicing seams of the multi-channel video frame images, the multi-channel video frame images are fused and spliced ​​to obtain fused video data;

[0013] The pre-trained hidden danger detection model is used to detect the fused video data and obtain the hidden danger detection results of distribution network equipment.

[0014] Furthermore, features of multiple video frame images are extracted and matched, including:

[0015] A neural network is used to extract multi-level features of multiple video frame images, and a coarse feature map that meets the small size requirement and a fine feature map that meets the large size requirement are obtained for each video frame image;

[0016] After adding learnable position encoding to each coarse feature map, it is converted into a one-dimensional vector, and then converted into an easily matched feature representation through a Transformer-based feature matching module to obtain a feature vector.

[0017] Taking any one channel as the main channel, the similarity matrix between the video frame images of other channels and the video frame images of the main channel is calculated by using the inner product of pixel-by-pixel vectors, and the optimal match is calculated using the dual-Softmax method. Some outlier matching pairs are filtered out by the mutual nearest neighbor algorithm to obtain the rough matching point pairs that match the multi-channel video frame images two by two. The similarity matrix is ​​used to represent the similarity between the feature vectors of two video frame images;

[0018] The coarse matching point pairs are mapped to the corresponding fine feature maps, and the mapping area is input into the Transformer-based feature matching module. The matching probability between the central features of the main video frame image and all the features of the matching other video frame images is calculated, and the feature position with the highest matching probability is used as the matching point position with sub-pixel accuracy. The matching point position with sub-pixel accuracy represents the splicing relationship between the two matching video frame images.

[0019] Furthermore, the stitching seam is an optimal stitching seam, and the optimal stitching seam is determined by a dynamic programming method with the goal of minimizing the image difference. The determination process includes:

[0020] Determine the pixel-by-pixel energy sum function of the two spliced ​​video frame images, where the energy sum function is used to quantify the image difference degree of the overlapping area of ​​the two video frame images;

[0021] Initialize the stitching path, and take the first row of the overlapping area as the minimum energy and path of the first step; iterate based on the minimum energy and, traverse each row in turn, obtain the location of the minimum energy and path source, and record the minimum energy and vector reaching each row and column, as well as the location of the path source reaching each row and column;

[0022] The seams are backtracked by recording the source position matrix from the last row, and the seam with the smallest total energy is selected as the best seam.

[0023] Furthermore, the multi-channel video frame images are fused and spliced, including:

[0024] The video frames to be spliced ​​are all subjected to wavelet transformation and decomposed using the Mallat algorithm to obtain low-frequency components and high-frequency components;

[0025] A weighted average fusion method is used for the low-frequency components to obtain the fused low-frequency components;

[0026] The high-frequency components are fused by taking the maximum absolute value of the window coefficient to obtain the fused high-frequency components;

[0027] The fused low-frequency component and high-frequency component are subjected to inverse wavelet transform to obtain the fused image.

[0028] Furthermore, according to the perspective difference of the video data of the multiple distribution network devices, the spatial and temporal parameters of the video data of the multiple distribution network devices are aligned in combination with the synchronization timestamp, including:

[0029] Time alignment: select the main channel and obtain the video frame data of the main channel at the set time, index the time and video frame data of the frames before and after the set time of other channels, and obtain the video frame data of other channels at the set time through linear interpolation method;

[0030] Spatial alignment: According to the camera calibration parameters of different viewing angles, geometric correction is performed on each channel of video data to eliminate the different degrees of distortion that occurs when each camera is imaging;

[0031] Brightness alignment: after converting the frame images of each channel of video data from RGB format to HSV format, unify the brightness of the video frame images at the same moment in each channel of video data;

[0032] Coordinate system alignment: the image reference system obtained from the main channel is used as the reference reference system. According to the camera external parameters between other channels and the main channel, the projection transformation model of the video frame images of other channels is calculated, and the multi-channel video frame data is projected to the reference reference system.

[0033] Furthermore, the hidden danger detection model is trained using the YOLOv8 network, and the pre-trained hidden danger detection model is used to detect the fused video data, including:

[0034] The fused video data is divided into several sub-images according to the input image size of the YOLOv8 network, and all sub-images are stacked into a batch;

[0035] Input data into the hidden danger detection model in batches, output the detection results of hidden danger targets of distribution network equipment in all sub-images in the batch, merge all the detection results, and convert the position coordinates of the target in the sub-image into the position coordinates of the original image;

[0036] All transformed detection results are subjected to non-maximum suppression processing to obtain the final hidden danger detection results.

[0037] Furthermore, video data of multiple distribution network devices with synchronized timestamps from different viewing angles are obtained, including:

[0038] Arrange multiple viewing angles according to the rule of overlapping adjacent viewing angles by a set degree, and construct a multi-view shooting system, in which the video shooting parameters of each viewing angle are the same;

[0039] Place the multi-view shooting system in the distribution network equipment scene to be tested, and set the video shooting mode to continuously acquire data.

[0040] In a second aspect, the present invention provides a distribution network equipment hidden danger detection device based on large field of view spliced ​​video data, comprising:

[0041] A data acquisition module is used to acquire video data of multiple distribution network devices with synchronized time stamps from different viewing angles;

[0042] A preprocessing module, used for aligning the spatial and temporal parameters of the video data of multiple distribution network devices according to the perspective differences of the video data of multiple distribution network devices in combination with synchronization timestamps;

[0043] The video fusion and splicing module is used to extract and match the features of multiple video frame images based on the aligned video data of multiple distribution network equipment to obtain the splicing relationship between the multiple video frame images; it is also used to determine the splicing seam between two video frame images to be spliced ​​based on the splicing relationship; it is also used to fuse and splice the multiple video frame images based on the splicing relationship and splicing seam of the multiple video frame images to obtain fused video data;

[0044] The hidden danger detection module is used to detect the fused video data using a pre-trained hidden danger detection model to obtain the hidden danger detection results of the distribution network equipment.

[0045] Accordingly, the present invention provides a computer device, the device comprising a processor and a memory:

[0046] The memory is used to store the computer program and send the instructions of the computer program to the processor;

[0047] The processor executes a method for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data as described in the first aspect according to the instructions of the computer program.

[0048] Accordingly, the present invention provides a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, a distribution network equipment hidden danger detection method based on large field of view spliced ​​video data as in the first aspect is implemented.

[0049] In summary, the present invention provides a method and related device for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data. The method significantly expands the field of view of detection by merging video data under different viewing angles, and solves the problem of repeated shooting of multiple distribution network equipment at the same location through a large field of view. The spatial and temporal parameters of multiple video data are precisely aligned in combination with synchronous timestamps, ensuring the continuity and consistency of the spliced ​​video in time and space, avoiding false detection and missed detection due to time dislocation or viewing angle differences, and improving the reliability and effectiveness of the detection results. Through the precise extraction and matching of multi-channel video frame image features, the method can efficiently determine the splicing relationship and optimal splicing seam between video frames, reduce visual distortion and seam traces in the splicing process, generate high-quality fused video data, and provide a clear and continuous visual basis for subsequent hidden danger detection. The pre-trained hidden danger detection model is used to automatically detect the fused video data, without the need to manually check the video one by one, which greatly improves the detection speed and work efficiency, while also reducing errors caused by human factors and enhancing the objectivity and standardization of the detection process. The present invention solves the problem that a single distribution network device needs to be filmed multiple times by efficiently integrating multiple video resources for unified detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0051] Figure 1 A flow chart of a method for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data provided by an embodiment of the present invention;

[0052] Figure 2 A technical roadmap for a method for detecting hidden dangers in distribution network equipment based on large-field-of-view spliced ​​video data provided by an embodiment of the present invention;

[0053] Figure 3 A block diagram of a device for detecting hidden dangers in distribution network equipment based on large-field-of-view spliced ​​video data provided by an embodiment of the present invention;

[0054] Figure 4 A block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0056] See also Figure 1 This embodiment provides a method for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data, comprising the following steps:

[0057] S1: Acquire multi-channel distribution network equipment video data with synchronized timestamps from different perspectives.

[0058] It should be noted that in the specific implementation of this step, multiple cameras or video acquisition devices can be arranged at different locations and angles of the distribution network to ensure that they record video data at the same time. The key is that all devices are equipped with a time synchronization mechanism (such as GPS clock synchronization), so that even if they are shot from different angles, each frame of video has a common timestamp. This provides a time reference system for subsequent processing and ensures data consistency.

[0059] S2: According to the perspective differences of the video data of multiple distribution network devices, the spatial and temporal parameters of the video data of multiple distribution network devices are aligned in combination with the synchronization timestamp.

[0060] It should be noted that due to the different positions and directions of the cameras, the captured videos are spatially inconsistent. This step uses the timestamps in the video data and the known camera position information to adjust the spatiotemporal parameters of each video stream through calculation, so that the videos from different perspectives are aligned in time and space, preparing for subsequent stitching.

[0061] S3: Based on the aligned multi-channel distribution network equipment video data, extract the features of the multi-channel video frame images and match them to obtain the splicing relationship between the multi-channel video frame images.

[0062] It should be noted that matching feature points are found between video frames, such as SIFT, ORB and other feature descriptors, and the corresponding relationship between these feature points is used to determine how different video frames are connected to each other. This process involves image processing technology, the purpose of which is to find frame pairs that can be seamlessly connected, thereby determining the order and method of splicing video frame images.

[0063] S4: Determine a stitching seam between two video frame images to be stitched based on the stitching relationship.

[0064] It should be noted that after clarifying the splicing relationship of the video frames, the optimal overlapping area (i.e., the splicing seam) between each pair of adjacent frames is further accurately calculated to minimize the splicing traces while maintaining the continuity of the scene. A smooth transition is usually achieved through an image fusion algorithm.

[0065] S5: Based on the splicing relationship and splicing seams of the multiple video frame images, the multiple video frame images are fused and spliced ​​to obtain fused video data.

[0066] It should be noted that, using the information obtained in the previous steps, multiple video frames are fused into a single large-field-of-view video in the correct order and overlapping areas through image stitching technology. This process may include color and brightness adjustments to ensure the visual coherence and naturalness of the final video.

[0067] S6: Use the pre-trained hidden danger detection model to detect the fused video data and obtain the hidden danger detection results of the distribution network equipment.

[0068] It should be noted that this step is to apply machine learning or deep learning hidden danger detection models to analyze the spliced ​​video data. These models learn to identify common hidden dangers of distribution network equipment during the training phase, such as insulation damage and loose components. Through model reasoning, the system can automatically mark suspected hidden danger areas, provide maintenance personnel with accurate inspection clues, and speed up the speed of hidden danger detection and repair.

[0069] The present embodiment provides a method for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data. The method significantly expands the field of view of detection by merging video data from different perspectives, and solves the problem of repeated shooting of multiple distribution network equipment at the same location through a large field of view. The spatial and temporal parameters of multiple video data are precisely aligned in combination with synchronous timestamps, ensuring the continuity and consistency of the spliced ​​video in time and space, avoiding false detection and missed detection caused by time dislocation or perspective difference, and improving the reliability and effectiveness of the detection results. By accurately extracting and matching the image features of multiple video frames, the method can efficiently determine the splicing relationship and optimal splicing seam between video frames, reduce visual distortion and seam traces in the splicing process, generate high-quality fused video data, and provide a clear and continuous visual basis for subsequent hidden danger detection. The pre-trained hidden danger detection model is used to automatically detect the fused video data, without the need to manually check the video one by one, which greatly improves the detection speed and work efficiency, while also reducing errors caused by human factors, and enhancing the objectivity and standardization of the detection process. The present invention solves the problem that a single distribution network device needs to be shot multiple times in a refined manner by efficiently integrating multiple video resources for unified detection. It provides a new type of intelligent inspection technology for the detection of hidden dangers in distribution network equipment during power inspection, improves inspection efficiency, and promotes the intelligent development of power inspection.

[0070] See also Figure 2 , Figure 2 The following is a technical roadmap of a method for detecting hidden dangers in distribution network equipment based on large field of view spliced ​​video data. Figure 2 The present invention is further introduced.

[0071] In a preferred embodiment of the present invention, for step S1, obtaining video data of multiple power distribution network devices with synchronized timestamps at different viewing angles includes:

[0072] Multiple perspectives are arranged according to the rule of setting the degree of overlap of adjacent perspectives to build a multi-perspective shooting system, in which the video shooting parameters of each perspective are the same; the multi-perspective shooting system is placed in the distribution network equipment scene to be tested, and the video shooting mode is set to continuously acquire data.

[0073] In some specific implementations of this embodiment, a multi-view camera can be mounted on a drone using a gimbal. The multi-view camera is composed of multiple color cameras with the same parameters, and the field of view of each camera is set to overlap by 5°. The multi-view camera is used to collect images of distribution network equipment from different perspectives.

[0074] Among them, multi-view cameras have the following advantages in acquiring video data. After multi-view videos are stitched together, a large field of view video is obtained, which can include multiple distribution network equipment and large-size distribution network equipment in a single scene. In addition, video data has stronger information redundancy than image data, and there is no need to deliberately shoot high-quality and multi-angle image data, which effectively reduces the technical requirements of inspection personnel and improves the data acquisition efficiency of distribution network equipment during power inspections.

[0075] In a preferred embodiment of the present invention, for step S2, according to the perspective difference of the video data of the multiple distribution network devices, the spatial and temporal parameters of the video data of the multiple distribution network devices are aligned in combination with the synchronization timestamp, including:

[0076] S21: Time alignment, selecting the main channel and obtaining the video frame data of the main channel at the set time, indexing the time and video frame data of the frames before and after the set time of other channels, and obtaining the video frame data of other channels at the set time by linear interpolation method.

[0077] In some embodiments, the video data is time-aligned, including selecting a main camera and obtaining the main camera exist Data at this moment, indexing other cameras exist The frame time before and after the moment , And data, through linear interpolation method to obtain other cameras in Data at the moment:

[0078]

[0079] in, , Represents the main camera and other cameras respectively; , Respectively represent other cameras exist , Video frame data collected at all times; Indicates other cameras exist Time-aligned video frame data.

[0080] S22: Spatial alignment: According to the camera calibration parameters of different viewing angles, geometric correction is performed on each channel of video data to eliminate the different degrees of distortion that occurs when each camera is imaging.

[0081] In some specific implementations, spatially aligning the video data includes geometrically correcting each channel of video data according to calibration parameters of each camera to eliminate different degrees of distortion occurring in the imaging process of each camera.

[0082] S23: Brightness alignment: after converting the frame images of each channel of video data from RGB format to HSV format, unify the brightness of the video frame images of each channel of video data at the same time.

[0083] In some specific implementations, brightness alignment of video data includes converting the frame image from RGB format to HSV format, and then using the illumination gain compensation coefficient to adaptively adjust and unify the brightness of the video frame images in each channel of video data at the same time:

[0084]

[0085] in, , Respectively represent the brightness values ​​of the image before and after illumination gain compensation; Represents the light gain compensation coefficient.

[0086] The illumination gain compensation coefficient of each camera image is solved by the extreme point of the illumination gain error function between the main camera image and other camera images:

[0087]

[0088] in, represents the gain compensation error, , Respectively represent the gain coefficients of the main camera image and other camera images; , Respectively represent the brightness values ​​of the main camera and other camera images; , Represent the pixel coordinates of the main camera and other camera images respectively; , They represent the overlapping areas between the main camera and other camera images respectively.

[0089] In order to simplify the calculation and improve the robustness of the gain, the empirical formula of the illumination gain error function is:

[0090]

[0091] in, , Respectively represent the average brightness value of the overlapping area of ​​the main camera image and other camera images; , Represent the standard deviation of error and gain respectively. Usually Approximately 0.7, Approximately 0.1, Indicates the number of pixels that overlap.

[0092] To solve for the gain coefficient, a closed-form solution can be obtained by taking the derivative of the above function to be 0. , The following linear equations are obtained by derivation:

[0093]

[0094] Obtain is 1, for .

[0095] S24: Coordinate system alignment: the image reference system obtained from the main channel is used as the reference reference system, and according to the camera external parameters between other channels and the main channel, the projection transformation model of the video frame images of other channels is calculated, and the multi-channel video frame data is projected to the reference reference system.

[0096] In some specific embodiments, aligning the coordinate system of the video data includes taking the image reference system obtained by the main camera as the reference reference system, calculating the projection transformation model of the video frame images of other cameras based on camera external parameters such as the angle difference and displacement difference between the other cameras and the main camera, projecting multiple channels of video frame data into a unified coordinate system, and reducing the parallax between the video data of each channel.

[0097] The multi-channel video frame images are projected into a unified coordinate system using cylindrical projection, and the plane image is projected onto the surface of the cylinder:

[0098]

[0099] in, Represents the pixel coordinates of the image after cylindrical projection; , Represents the image width and height respectively.

[0100] In this embodiment, a generalized normalization operation is performed on multiple video data, and multiple video data are preprocessed in four aspects: time, space, brightness, and coordinate system. Time alignment is used to unify the video frame images of each video at the same time, space alignment is used to eliminate the difference in data acquisition of cameras with the same parameters caused by camera distortion, brightness alignment is used to unify the brightness of the video frame images of each video at the same time, and coordinate system alignment is used to eliminate the spatial coordinates of video data with different viewing angles.

[0101] Among them, the generalized normalization of multi-channel video data has the following advantages: it can reduce or eliminate the differences between video data caused by objective factors such as environment and equipment, reduce the difficulty of video data fusion, improve the quality of fused video data, and provide a high-quality data foundation for subsequent distribution network equipment hidden danger detection.

[0102] In a preferred embodiment of the present invention, extracting features of multiple video frame images and matching them includes:

[0103] S31: extracting multi-level features of multiple video frame images using a neural network to obtain a coarse feature map that meets a small size requirement and a fine feature map that meets a large size requirement for each video frame image.

[0104] In some specific embodiments, feature extraction is performed on the video frame image, and the multi-level features of the video frame image are extracted using the ResNet deep neural network loaded with ImageNet pre-trained weights, including the original image Coarse feature map and original image of small size Fine feature maps of large size.

[0105] S32: Add a learnable position code to each coarse feature map and convert it into a one-dimensional vector. Then, use a Transformer-based feature matching module to convert the one-dimensional vector into an easily matched feature representation to obtain a feature vector.

[0106] It should be noted that learnable position encoding is added pixel by pixel to the coarse feature map of multiple video frame images and converted into a one-dimensional vector. The information between multiple video frame images is fused through the Transformer-based feature matching module and converted into a feature representation that is easy to match.

[0107] S33: Taking any one channel as the main channel, the similarity matrix between the video frame images of other channels and the video frame images of the main channel is calculated by using the pixel-by-pixel vector inner product method, and the optimal match is calculated using the dual-Softmax method. Some outlier matching pairs are filtered out through the mutual nearest neighbor algorithm to obtain coarse matching point pairs that match multiple video frame images pairwise. The similarity matrix is ​​used to represent the similarity between the feature vectors of two video frame images.

[0108] It should be noted that the calculation formula of dual-Softmax is as follows:

[0109]

[0110] in, represents the matching probability matrix; , Represent the video frame image features of the main camera and other cameras respectively; , Represents the pixel coordinates of the video frame image features of the main camera and other cameras respectively.

[0111] S34: Map the coarse matching point pairs to the corresponding fine feature maps, and input the mapping area into the Transformer-based feature matching module to calculate the matching probability between the central features of the main video frame image and all the features of the matching other video frame images, and use the feature position with the highest matching probability as the matching point position with sub-pixel accuracy. The matching point position with sub-pixel accuracy represents the splicing relationship between the two matching video frame images.

[0112] It should be noted that the point pairs in the rough matching , Mapped to the corresponding fine feature map, and input the mapped area into the Transformer-based feature matching module to further refine the matching features and calculate Central features and The matching probability of all features in is obtained, and the corresponding probability distribution is obtained. The matching point positions with sub-pixel accuracy in .

[0113] In this embodiment, a ResNet neural network pre-trained on ImageNet is used to extract features from video frame images to obtain small-sized coarse feature maps and large-sized fine feature maps. Coarse and fine feature matching is achieved through similarity measurement between feature maps of multiple video frame images and LoFTR feature matching module, and the sub-pixel matching point position is obtained by calculating the probability distribution between features. A dynamic programming method is used to find the best seam between multiple video frame images, and an image fusion algorithm based on wavelet transform is used to achieve the splicing of video frame images. The spliced ​​video frame images are post-processed to obtain high-quality large-field spliced ​​video.

[0114] In a preferred embodiment of the present invention, the stitching seam is an optimal stitching seam, and the optimal stitching seam is determined by a dynamic programming method with the goal of minimizing image differences. The determination process includes:

[0115] S41: Determine a pixel-by-pixel energy sum function of the two spliced ​​video frame images, where the energy sum function is used to quantify the degree of image difference in the overlapping area of ​​the two video frame images.

[0116] In some specific implementations, the idea of ​​minimizing differences is used to construct a pixel-by-pixel energy sum function for multiple video frame images. , evaluate the structural and color differences between images at the seams, and the energy function is calculated as follows:

[0117]

[0118] in, Represents all the pixels through which the seam passes; Represents the weight used to balance the color difference term and the structure difference term; Indicates the color difference of pixels at the same position in the overlapping area; Indicates the structural difference of pixels at the same position in the overlapping area.

[0119]

[0120]

[0121] in, , Respectively , Sobel operator in direction; , Represent the video frame images of the main camera and other cameras respectively.

[0122] S42: Initialize the stitching seam path, and take the first row of the overlapping area as the minimum energy and path of the first step; iterate based on the minimum energy and, traverse each row in turn, obtain the location of the minimum energy and path source, and record the minimum energy and vector reaching each row and column, as well as the location of the path source reaching each row and column.

[0123] It should be noted that the best stitching seam is found by combining the dynamic programming method. Initialize the stitching seam path, and take the first row as the minimum energy and path of the first step; iterate according to the minimum energy and, traverse each row in turn, and get the location of the minimum energy and path source: upper left, top, and upper right. When traversing, it is necessary to record the minimum energy and vector to reach each row and column, as well as the location of the path source to reach each row and column, for use in backtracking.

[0124] S43: backtracking the seam, backtracking the seam from the last row by recording the source position matrix, and selecting the seam with the smallest total energy as the best seam.

[0125] It should be noted that the seam is backtracked by recording the source position matrix from the last row. In terms of direction, if it is in the upper left, subtract 1; if it is in the upper right, add 1; if it is in the top, leave it unchanged. In the direction, subtract 1. Compare all the stitching seams and select the stitching line with the smallest total energy as the best stitching line.

[0126] In a preferred embodiment of the present invention, fusing and splicing multiple video frame images includes:

[0127] S51: performing wavelet transform on the video frame images to be spliced, and decomposing them using Mallat algorithm to obtain low-frequency components and high-frequency components.

[0128] S52: A weighted average fusion method is used for the low-frequency components to obtain fused low-frequency components.

[0129] In some specific implementations, a weighted average fusion method is used for low-frequency components, and the fusion algorithm is as follows:

[0130]

[0131] in, , Represent the low-frequency coefficients of the video frame images of the main camera and other cameras respectively; , They represent the weight values ​​corresponding to the low-frequency coefficients respectively, and the sum is 1.

[0132] S53: The high-frequency components are fused by taking the maximum absolute value of the window coefficients to obtain fused high-frequency components.

[0133] In some specific implementations, a method based on taking the maximum absolute value of window coefficients is used to fuse high-frequency components. The fusion algorithm is as follows:

[0134]

[0135] in, , Represent the high-frequency coefficients of the video frame images of the main camera and other cameras respectively.

[0136] S54: performing inverse wavelet transform on the fused low-frequency component and high-frequency component to obtain a fused image.

[0137] S55: Post-processing of large field of view stitching videos.

[0138] The details and edges of the stitched image are enhanced by sharpening filtering and local contrast enhancement. Secondly, the noise level of the stitched image is reduced by image denoising technology. Finally, the stitched video frame images are saved as a video sequence to output large field of view video data.

[0139] In a preferred embodiment of the present invention, the hidden danger detection model is trained using the YOLOv8 network, and the training steps include:

[0140] S61: Construct a distribution network equipment target detection dataset based on large field of view stitching video data.

[0141] First, data cleaning is performed to remove video frame images with poor imaging effects caused by, for example, camera overexposure, camera underexposure, camera defocus, motion blur, and internal parameter calibration errors.

[0142] Secondly, the hidden dangers of distribution network equipment in the video data are annotated and screened frame by frame, and the data is divided into training set and verification set to form a standard data set that the neural network can learn. The data set contains the images, categories and location information of the hidden dangers of distribution network equipment.

[0143] S62: Use the hidden danger detection training data set to train the YOLOV8 target detection neural network.

[0144] Firstly, minimizing the difference between the predicted value and the true value is used as the neural network optimization criterion, and focalloss is used as the classification loss function of the distribution network equipment hidden danger target detection network to reduce the impact of the imbalance of hidden danger data volume and identification complexity of different distribution network equipment on network stability; intersection-over-union loss is used as the distribution network equipment hidden danger target box positioning loss function:

[0145]

[0146]

[0147] in, represents the classification loss function of hidden dangers of distribution network equipment; Indicates the closeness between the predicted category and the actual category of the hidden danger of the distribution network equipment. The closer it is to 1, the more accurate the prediction result is; Indicates an adjustable factor, greater than 0; Indicates the loss of the target frame positioning of hidden dangers of distribution network equipment; , Indicates the predicted area and actual area of ​​the target box for hidden dangers of distribution network equipment.

[0148] Secondly, the gradient descent method is used to update the network parameters, and the small batch stochastic gradient descent method is used as the parameter optimization method:

[0149]

[0150] in, , , Represents the learnable parameters of the YOLOV8 network; lr represents the learning rate, which represents the step size of each optimization; represents the exponential moving average of the gradient; represents the exponential moving average of the squared gradient; Indicates a very small non-zero number to avoid network optimization failure caused by the divisor being 0 during the optimization process. The default value is .

[0151] S63: Evaluate the trained network, and use mAP and IoU to evaluate the effect of YOLOV8 target detection neural network on the classification and location of hidden dangers in distribution network equipment. Among them, the calculation of mAP adopts the method of calculating the average precision AP of each category and taking the average. For the average precision AP of each category, the PR curve of precision (Precision) and recall rate (Recall) is constructed, and the area between the PR curve and the coordinate axis is calculated to obtain the average precision AP of each category; IoU adopts the calculation method of the intersection and union ratio of each predicted target and the true target, and after calculating the common area and the total area respectively, the common area is divided by the total area to obtain the final result.

[0152] Calculate mAP:

[0153]

[0154] Calculate IoU:

[0155]

[0156] Where, P(r) represents the PR curve function; Indicates the maximum recall rate; Indicates the prediction target area; Indicates the true target area.

[0157] In a further embodiment of the present invention, the fused video data is detected using a pre-trained hidden danger detection model, including:

[0158] S71: Divide the fused video data into several sub-images according to the input image size of the YOLOv8 network, and stack all the sub-images into a batch.

[0159] In some specific implementations, the sliding window size is set according to the input image size of the YOLOV8 network, the large-field-of-view distribution network hidden danger video frame image is evenly divided into several sub-images according to a 10% overlapping area, and all the sub-images are stacked into a batch.

[0160] S71: Input data into the hidden danger detection model according to the batch, output the detection results of the hidden danger targets of the distribution network equipment of all sub-images in the batch, merge all the detection results, and convert the position coordinates of the target in the sub-image into the position coordinates of the original image.

[0161] In some specific implementations, data is input into the trained YOLOV8 network in batches, the detection results of distribution network hidden danger targets of all sub-images in the batch are output, all detection results are merged, and the position coordinates of the target in the sub-image are converted into the position coordinates of the original image:

[0162]

[0163] in, , Respectively represent the upper left corner coordinates and lower right corner coordinates of the target box prediction area in the original image; , Respectively represent the upper left corner coordinates and lower right corner coordinates of the target box prediction area in the original image; Indicates the position coordinates of the upper left corner of the sub-image in the original image.

[0164] S71: Perform non-maximum suppression processing on all transformed detection results to obtain the final hidden danger detection result.

[0165] All the converted detection results are processed by non-maximum suppression to obtain the final hidden danger detection results, remove the hidden dangers of repeated detection, obtain the final hidden danger detection results, and complete the detection of hidden dangers in distribution network equipment with a large field of view while maintaining high pixels.

[0166] The above embodiment uses video data to construct a distribution network equipment hidden danger target detection data set, uses the data set to train and verify the target detection network, and combines the sliding block detection method to achieve distribution network equipment hidden danger detection that maintains high-pixel and large-field-of-view images.

[0167] Based on the same inventive concept, the embodiment of the present application also provides a distribution network equipment hidden danger detection device based on large field of view spliced ​​video data for realizing the distribution network equipment hidden danger detection method based on large field of view spliced ​​video data involved above. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiment of the distribution network equipment hidden danger detection device based on large field of view spliced ​​video data provided below can refer to the limitations of the distribution network equipment hidden danger detection method based on large field of view spliced ​​video data above, and will not be repeated here.

[0168] See also Figure 3 This embodiment provides a distribution network equipment hidden danger detection device based on large field of view spliced ​​video data, including:

[0169] A data acquisition module is used to acquire video data of multiple distribution network devices with synchronized time stamps from different viewing angles;

[0170] A preprocessing module, used for aligning the spatial and temporal parameters of the video data of multiple distribution network devices according to the perspective differences of the video data of multiple distribution network devices in combination with synchronization timestamps;

[0171] The video fusion and splicing module is used to extract and match the features of multiple video frame images based on the aligned video data of multiple distribution network equipment to obtain the splicing relationship between the multiple video frame images; it is also used to determine the splicing seam between two video frame images to be spliced ​​based on the splicing relationship; it is also used to fuse and splice the multiple video frame images based on the splicing relationship and splicing seam of the multiple video frame images to obtain fused video data;

[0172] The hidden danger detection module is used to detect the fused video data using a pre-trained hidden danger detection model to obtain the hidden danger detection results of the distribution network equipment.

[0173] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0174] Reference Figure 4 The embodiment of the present invention further provides a computer device 40, comprising: a memory 402 and a processor 401 and a computer program 403 stored in the memory 402. When the computer program 403 is executed on the processor 401, a distribution network equipment hidden danger detection method based on large field of view spliced ​​video data as described in any one of the above methods is implemented.

[0175] The computer device 40 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The computer device 40 may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art will appreciate that Figure 4 This is merely an example of the computer device 40 and does not constitute a limitation on the computer device 40 . The computer device 40 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0176] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0177] In some embodiments, the memory 402 may be an internal storage unit of the computer device 40, such as a hard disk or memory of the computer device 40. In other embodiments, the memory 402 may also be an external storage device of the computer device 40, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device 40. Further, the memory 402 may also include both an internal storage unit of the computer device 40 and an external storage device. The memory 402 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory 402 may also be used to temporarily store data that has been output or is to be output.

[0178] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data as described in any one of the above methods is implemented.

[0179] In this embodiment, if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, USB flash drive, mobile hard disk, disk or optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0180] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0181] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0182] In the embodiments disclosed in the present application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0183] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting hidden dangers of distribution network equipment based on large field of view spliced ​​video data, characterized in that: The steps include: Acquire multi-channel distribution network equipment video data with synchronized timestamps from different viewing angles; According to the viewing angle difference of the multi-channel power distribution network equipment video data, the spatial and temporal parameters of the multi-channel power distribution network equipment video data are aligned in combination with the synchronization timestamp; Based on the aligned multi-channel distribution network equipment video data, extract the features of the multi-channel video frame images and match them to obtain the splicing relationship between the multi-channel video frame images; Determine a stitching seam between two video frame images to be stitched based on the stitching relationship; Based on the splicing relationship and the splicing seam of the multiple video frame images, the multiple video frame images are fused and spliced ​​to obtain fused video data; The fused video data is detected using a pre-trained hidden danger detection model to obtain a hidden danger detection result for distribution network equipment.

2. The method for detecting hidden dangers of distribution network equipment based on large field of view spliced ​​video data according to claim 1 is characterized in that: Extract features of multiple video frames and match them, including: A neural network is used to extract multi-level features of multiple video frame images, and a coarse feature map that meets the small size requirement and a fine feature map that meets the large size requirement are obtained for each video frame image; Add a learnable position code to the coarse feature map of each path and convert it into a one-dimensional vector, and convert the one-dimensional vector into an easy-to-match feature representation through a Transformer-based feature matching module to obtain a feature vector; Taking any one channel as the main channel, using the inner product of pixel-by-pixel vectors to calculate the similarity matrix between the video frame images of other channels and the video frame image of the main channel, and using the dual-Softmax method to calculate the optimal match, filtering some outlier matching pairs through the mutual nearest neighbor algorithm, and obtaining the rough matching point pairs that make the multi-channel video frame images match each other, the similarity matrix is ​​used to represent the similarity between the feature vectors of the two video frame images; The coarse matching point pairs are mapped to the corresponding fine feature maps, and the mapping area is input into the Transformer-based feature matching module to calculate the matching probability between the central features of the main video frame image and all the features of the matching other video frame images. The feature position with the highest matching probability is used as the matching point position with sub-pixel precision. The matching point position with sub-pixel precision represents the splicing relationship between the two matching video frame images.

3. The method for detecting hidden dangers of distribution network equipment based on large field of view spliced ​​video data according to claim 1 is characterized in that: The stitching seam is an optimal stitching seam, and the optimal stitching seam is determined by a dynamic programming method with minimizing image differences as the goal, and the determination process includes: Determine a pixel-by-pixel energy sum function of the two spliced ​​video frame images, wherein the energy sum function is used to quantify the degree of image difference in the overlapping area of ​​the two video frame images; Initialize the stitching path, and take the first row of the overlapping area as the minimum energy and path of the first step; iterate based on the minimum energy and, traverse each row in turn, obtain the location of the minimum energy and path source, and record the minimum energy and vector reaching each row and column, as well as the location of the path source reaching each row and column; The seams are backtracked by recording the source position matrix from the last row, and the seam with the smallest total energy is selected as the best seam.

4. The method for detecting hidden dangers of distribution network equipment based on large field of view spliced ​​video data according to claim 1 is characterized in that: The multi-channel video frame images are merged and spliced, including: Performing wavelet transform on the video frame images to be spliced, and decomposing them using Mallat algorithm to obtain low-frequency components and high-frequency components; A weighted average fusion method is used for the low-frequency components to obtain fused low-frequency components; The high-frequency components are fused by using a method based on taking the maximum absolute value of window coefficients to obtain fused high-frequency components; The fused low-frequency component and the fused high-frequency component are subjected to inverse wavelet transformation to obtain a fused image.

5. The method for detecting hidden dangers of distribution network equipment based on large field of view spliced ​​video data according to claim 1 is characterized in that: According to the perspective difference of the multi-channel distribution network equipment video data, the spatial and temporal parameters of the multi-channel distribution network equipment video data are aligned in combination with the synchronization timestamp, including: Time alignment: select the main channel and obtain the video frame data of the main channel at the set time, index the time and video frame data of the frames before and after the set time of other channels, and obtain the video frame data of other channels at the set time by linear interpolation method; Spatial alignment: According to the camera calibration parameters of different viewing angles, geometric correction is performed on each channel of video data to eliminate the different degrees of distortion that occurs when each camera is imaging; Brightness alignment: after converting the frame images of each channel of video data from RGB format to HSV format, unify the brightness of the video frame images at the same moment in each channel of video data; The coordinate system is aligned, and the image reference system obtained from the main channel is used as the reference reference system. According to the camera external parameters between other channels and the main channel, the projection transformation model of the video frame image of other channels is calculated, and the multi-channel video frame data is projected to the reference reference system.

6. The method for detecting hidden dangers of distribution network equipment based on large field of view spliced ​​video data according to claim 1 is characterized in that: The hidden danger detection model is trained by using the YOLOv8 network, and the fused video data is detected by using the pre-trained hidden danger detection model, including: The fused video data is divided into a plurality of sub-images according to the input image size of the YOLOv8 network, and all the sub-images are stacked into a batch; Input data into the hidden danger detection model in batches, output the detection results of hidden danger targets of distribution network equipment in all sub-images in the batch, merge all the detection results, and convert the position coordinates of the target in the sub-image into the position coordinates of the original image; All transformed detection results are subjected to non-maximum suppression processing to obtain the final hidden danger detection results.

7. The method for detecting hidden dangers of distribution network equipment based on large field of view spliced ​​video data according to claim 1 is characterized in that: Acquire multi-channel distribution network equipment video data with synchronized timestamps from different viewing angles, including: Arranging multiple viewing angles according to a rule that adjacent viewing angles overlap by a set degree to construct a multi-view shooting system, wherein the video shooting parameters of each viewing angle in the multi-view shooting system are the same; The multi-view shooting system is placed in the distribution network equipment scene to be detected, and the video shooting mode is set to continuously acquire data.

8. A device for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data, characterized in that: include: A data acquisition module is used to acquire video data of multiple distribution network devices with synchronized time stamps from different viewing angles; A preprocessing module, configured to align the spatial and temporal parameters of the multi-channel power distribution network device video data according to the viewing angle difference of the multi-channel power distribution network device video data in combination with the synchronization timestamp; A video fusion and splicing module is used to extract and match the features of multiple video frame images based on the aligned multiple distribution network equipment video data to obtain the splicing relationship between the multiple video frame images; it is also used to determine the splicing seam between two video frame images to be spliced ​​based on the splicing relationship; it is also used to fuse and splice the multiple video frame images based on the splicing relationship and the splicing seam of the multiple video frame images to obtain fused video data; The hidden danger detection module is used to detect the fused video data using a pre-trained hidden danger detection model to obtain a hidden danger detection result of the distribution network equipment.

9. A computer device, characterized in that: The device comprises a processor and a memory: The memory is used to store a computer program and send instructions of the computer program to the processor; The processor executes a method for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data as described in any one of claims 1-7 according to the instructions of the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a method for detecting hidden dangers of distribution network equipment based on large-field-of-view spliced ​​video data according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • An intelligent image fusion method based on target feature driving

    CN109035188A

  • Monitoring video splicing method and device, equipment and storage medium

    CN119359537A

Cited By

  • Multi-mode collaborative awareness power station high-risk operation inspection method and system

    CN120236249A

  • A multimodal collaborative perception power station high-risk operation inspection method and system

    CN120236249B

  • Overhead exit ramp connection area vehicle track restoration method, device and medium

    CN121053166A