Power transmission channel abnormal target detection method and device based on multi-modal sensor and deep learning

By combining multimodal sensors with deep learning, accurate detection of abnormal targets in power transmission channels has been achieved, solving the problems of poor environmental adaptability and response delay in existing technologies, and improving detection accuracy and real-time performance.

CN121580326APending Publication Date: 2026-02-27BEIJING GUOWANG FUDA SCI & TECH DEV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511864729.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing technologies have poor environmental adaptability in power transmission channel safety monitoring, especially in complex environments such as nighttime and fog, where the accuracy of identification is low, the false alarm rate is high, and it is difficult to predict dynamic risks in real time.

Method used

By employing multimodal sensor fusion technology, real-time acquisition and preprocessing of visible light, infrared, lidar, and sound data are combined with a dual-branch CNN deep learning model for data registration and feature extraction, generating a joint feature map, and performing feature-level and decision-level data fusion to achieve accurate detection of abnormal targets in power transmission channels.

Benefits of technology

It improves the detection accuracy and real-time performance of abnormal targets in power transmission channels, and can accurately identify targets such as vehicles and smoke in complex environments, reduce the false negative rate, and improve the system's response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121580326A_ABST
    Figure CN121580326A_ABST
Patent Text Reader

Abstract

The invention discloses a power transmission channel abnormal target detection method and device based on a multi-modal sensor and deep learning, and relates to the technical field of target detection, and the method comprises the steps: obtaining multi-modal data in real time; registering the visible light data and the infrared data to obtain a registered visible light image and a registered infrared image; a double-branch CNN deep learning model is adopted, and a combined feature map is obtained based on the registered visible light image and the registered infrared image; preprocessing the laser radar data to obtain a laser radar thermodynamic diagram; obtaining a sound spectrum code based on the sound data; obtaining fusion data based on the joint feature map, the laser radar thermodynamic diagram and sound spectrum coding; inputting the fusion data into a double-task detection head to obtain a target detection result; the target detection result comprises a vehicle detection frame and a smoke and fire segmentation mask. According to the invention, the detection precision and real-time performance of the abnormal target of the power transmission channel in a complex environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of target detection technology, and in particular to a method and apparatus for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning. Background Technology

[0002] With the advancement of smart grid construction, real-time monitoring of safety hazards in power transmission channels (such as intrusion by construction vehicles and smoke from wildfires) has become crucial for ensuring grid security. Related technologies mainly rely on single visible light cameras or manual inspections, which present the following problems: ① Poor environmental adaptability: The accuracy of visible light recognition drops sharply at night and in foggy weather. ② High false negative rate: Small targets such as smoke and fire are easily interfered with by complex backgrounds. ③ Response delay: Traditional methods struggle to predict dynamic risks (such as the movement trajectory of vehicle robotic arms).

[0003] In recent years, the combination of multimodal sensor fusion and deep learning technology has provided new ideas for solving the above problems. However, in related technologies, infrared and lidar data are usually processed independently, lacking a spatiotemporally aligned fusion framework; sound signals have not been effectively used for the collaborative discrimination of fireworks and vehicle engines; and model training data augmentation methods are singular and difficult to cover extreme environments (such as strong light reflection, rain and snow).

[0004] Therefore, there is an urgent need for a method for detecting abnormal targets in power transmission channels using multimodal data fusion and adaptive deep learning, in order to improve the detection accuracy and real-time performance of potential hazards in power transmission channels under complex environments. Summary of the Invention

[0005] The purpose of this application is to provide a method and device for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning, which can improve the detection accuracy and real-time performance of abnormal targets in power transmission channels under complex environments.

[0006] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning, including: Real-time acquisition of multimodal data; the multimodal data is acquired by multimodal sensors, including visible light data, infrared data, lidar data, and sound data; The visible light data and the infrared data are registered to obtain a registered visible light image and a registered infrared image; A dual-branch CNN deep learning model is used to obtain a joint feature map based on the registered visible light image and the registered infrared image; The lidar data is preprocessed to obtain a lidar heat map; The sound spectrum encoding is obtained based on the sound data; Fusion data is obtained based on the joint feature map, the lidar heat map, and the sound spectrum encoding; The fused data is input into the dual-task detection head to obtain the target detection result; the target detection result includes: vehicle detection frame and smoke and fire segmentation mask.

[0007] Secondly, this application provides a power transmission channel abnormal target detection device based on multimodal sensors and deep learning, comprising: A multimodal sensor unit is used to collect multimodal data, including visible light data, infrared data, lidar data, and sound data. A data preprocessing unit, connected to the multimodal data acquisition unit, is used to register the visible light data and the infrared data to obtain a registered visible light image and a registered infrared image. The data preprocessing unit is also used to obtain a joint feature map based on the registered visible light image and the registered infrared image using a dual-branch CNN deep learning model. The data preprocessing unit is also used to preprocess the lidar data to obtain a lidar heatmap. The data preprocessing unit is also used to obtain sound spectrum encoding based on the sound data. A data fusion unit, connected to the data preprocessing unit, is used to obtain fused data based on the joint feature map, the lidar heat map, and the sound spectrum encoding. A detection unit, connected to the data fusion unit, is used to obtain target detection results based on the fused data; the target detection results include: vehicle detection frames and smoke / fire segmentation masks.

[0008] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a method and apparatus for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning. It registers real-time acquired visible light and infrared data, and obtains a joint feature map based on the registered visible light and infrared images, enabling the feature map to contain fused features with complementary information from both modes. Furthermore, it obtains fused data based on the joint feature map, lidar thermal map, and sound spectrum encoding, and uses this fused data for target detection. The sound signal is effectively used in the collaborative discrimination of smoke and vehicle target detection, enabling accurate target detection in complex environments, thereby improving the detection accuracy and real-time performance of abnormal targets in power transmission channels under complex conditions. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart illustrating an abnormal target detection method for power transmission channels based on multimodal sensors and deep learning, according to one embodiment of this application. Figure 2 This is a schematic diagram of the visible light data and infrared data registration process provided in an embodiment of this application; Figure 3 This is a schematic diagram of a multimodal data fusion process provided in an embodiment of this application; Figure 4 This is a schematic diagram of a power transmission channel abnormal target detection device based on multimodal sensors and deep learning in one embodiment of this application. Detailed Implementation

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0013] In one exemplary embodiment, such as Figure 1 As shown, a method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning is provided, including: Step 100: Acquire multimodal data in real time. The multimodal data is acquired by multimodal sensors, including visible light data, infrared data, lidar data, and sound data.

[0014] Step 200: Register the visible light data and infrared data to obtain the registered visible light image and the registered infrared image.

[0015] Step 300: A dual-branch CNN deep learning model is used to obtain a joint feature map based on the registered visible light image and the registered infrared image.

[0016] Step 400: Preprocess the lidar data to obtain a lidar heatmap. Obtain sound spectrum encoding based on the sound data.

[0017] Step 500: fused data is obtained based on the joint feature map, lidar heat map, and acoustic spectrum coding.

[0018] Step 600: Input the fused data into the dual-task detection head to obtain the target detection results. The target detection results include: vehicle detection boxes and smoke / fire segmentation masks.

[0019] As an optional implementation, step 100 includes: acquiring raw visible light data, raw infrared data, raw lidar data, and raw sound data in real time. Then, performing spatiotemporal alignment processing on the raw visible light data, raw infrared data, raw lidar data, and raw sound data to obtain visible light data, infrared data, lidar data, and sound data.

[0020] For example, raw visible light data, raw infrared data, raw lidar data, and raw sound data are acquired in parallel and synchronized using the PTP (Precision Time Protocol) (using the timestamp of the raw visible light data as a reference time). The timestamp and device (i.e., multimodal sensor) attitude are written in a unified manner, and spatiotemporal alignment processing is performed on the raw visible light data, raw infrared data, raw lidar data, and raw sound data.

[0021] In one embodiment, the spatiotemporal alignment of multi-source data (including raw visible light data, raw infrared data, raw lidar data, and raw sound data) includes two parts: temporal alignment and spatial alignment.

[0022] For time alignment: using the timestamp of the original visible light data as a reference time, the scan cycle closest to the reference time is retrieved in the original lidar data (radar point cloud time series). For point clouds that cross this time window (i.e., scan cycle), the scanning angle at the reference time is estimated by linear interpolation, and a lidar point cloud frame corresponding to the original visible light data is generated. The original sound data is divided into several time windows according to a fixed frame length and overlap rate. The center time of each time window is aligned with the timestamp of the original visible light data, thereby obtaining the sound data corresponding to the current image frame in the original visible light data.

[0023] In terms of spatial alignment: Extrinsic parameter matrices, including rotation matrix R and translation vector T, between the visible light camera, infrared thermal imager, and lidar were obtained beforehand through calibration experiments to describe the spatial relationship between the coordinate systems of each sensor. For lidar data, after time alignment, the point cloud coordinates were transformed from the lidar coordinate system to the visible light camera coordinate system using ([R|T]). Then, the known intrinsic parameter matrices were used to project the 3D points onto the pixel plane, and the point cloud density and height information within the pixel grid were statistically analyzed to obtain lidar data corresponding to the visible light image resolution.

[0024] As an optional implementation, to address the parallax caused by the optical axis deviation of the dual cameras (typically a deviation of 3-5 pixels) and achieve pixel-level coordinate unification between visible light and infrared data, thus providing a foundation for subsequent multimodal data fusion, step 200 includes: detecting feature points in the visible light data to obtain visible light feature points; detecting feature points in the infrared data to obtain infrared feature points; performing feature point matching using the visible light and infrared feature points to obtain feature point matching results; and performing perspective transformation on the infrared data based on the feature point matching results to obtain a registered infrared image, while using the visible light data as the registered visible light image.

[0025] In one embodiment, to improve the registration accuracy of visible light and infrared data and the robustness of subsequent feature extraction, the visible light and infrared data are processed as follows before feature point detection and registration: 1) Denoising: Gaussian filtering or bilateral filtering is applied to smooth the visible light and infrared images respectively to suppress high-frequency noise and quantization noise, while preserving edge information to provide a stable gradient for subsequent SIFT / ORB and SURF feature point detection. 2) Distortion correction: Using the camera intrinsic parameter matrix and distortion coefficients obtained during the factory or field calibration stage, geometric distortion correction is performed on the images captured by the visible light camera and infrared thermal imager to correct barrel / pincushion distortion to the standard imaging plane, ensuring accurate correspondence between pixel coordinates and physical space. 3) Brightness / Gamma Adaptive Adjustment: Based on the brightness histogram and contrast index of the current frame image, the overall brightness and gamma value of the image are adaptively adjusted. For example, when the histogram is concentrated in the low brightness range, the exposure gain is increased and the gamma index is decreased to enhance shadow details; when the histogram is concentrated in the high brightness range, the highlight areas are compressed and the gamma index is increased to suppress overexposure. Through the above brightness / gamma adaptive adjustment, images acquired at different times and under different weather conditions are unified to a brightness range suitable for feature extraction and subsequent network input.

[0026] After completing the above preprocessing, feature points are extracted from visible light data using the SIFT or ORB algorithm, and feature points are extracted from infrared data using the SURF algorithm. Cross-modal registration is then performed accordingly.

[0027] To achieve pixel-level coordinate unification between infrared and visible light data, facilitating the extraction of texture and temperature features by the subsequent dual-branch CNN deep learning model, the registration process for infrared and visible light data includes: 1) Intrinsic / Extrinsic Parameter Calibration: During the equipment installation phase before power transmission channel anomaly detection, a calibration board is used to calibrate the intrinsic parameters of the visible light camera and the infrared thermal imager, obtaining the focal length, principal point coordinates, and distortion coefficients. Simultaneously, multiple sets of synchronous images are acquired, and the rotation matrix and translation vector between the two camera coordinate systems are obtained through hand-eye calibration or common field-of-view calibration, forming the extrinsic parameter transformation from the infrared camera to the visible light camera. 2) Coarse Alignment: For the acquired infrared data, using the aforementioned intrinsic / extrinsic parameters, the infrared pixel coordinates are mapped to the visible light camera imaging plane through three-dimensional projection and coordinate transformation to obtain a coarsely registered infrared image, initially eliminating parallax and optical axis deviation. 3) Fine-grained registration of feature points: Based on the coarse registration results, for visible light data (RGB images, i.e., visible light images), SIFT (Scale-Invariant Feature Transform) or ORB (Oriented FAST and Rotated BRIEF) algorithms are used for feature point detection, taking into account the robustness to brightness / contrast differences and scale rotation changes across the spectrum (IR↔RGB). In low-computing-power scenarios, ORB can be used as a fallback for higher speed. Robustness to scale, rotation, and certain illumination changes is achieved through scale-space extremum detection, principal direction assignment, and gradient histogram description. For infrared data (IR images, i.e., infrared images), SURF (Speeded Up Robust Features) is used for feature point detection. Brute-force matching is performed on visible light feature points and infrared feature points, followed by bidirectional cross-validation to obtain matching points. Then, the RANSAC (RANdom SAmple Consensus) algorithm is used to remove outliers from the matched points, and the homography matrix H is calculated. Perspective transformation is then performed on the infrared data according to the homography matrix H to obtain the registered infrared image. Visible light data is used as the registered visible light image. For example... Figure 2 As shown.

[0028] As an optional implementation, in order to convert the disordered point cloud in the lidar data into regular data to meet the detection requirements of the power transmission scenario, step 400 preprocesses the lidar data to obtain a lidar heat map, including: performing voxelization downsampling processing on the lidar data according to a set grid size to obtain a lidar heat map.

[0029] The implementation process of obtaining the sound spectrum code based on the sound data in step 400 includes: extracting the spectrum of the sound data to obtain the sound spectrum code.

[0030] In one embodiment, the preprocessing of lidar data includes: 1) Voxel downsampling: Setting the grid size (e.g., 0.1m × 0.1m), constructing a two-dimensional or three-dimensional voxel grid within a preset region of interest, aggregating point clouds falling into the same voxel, and using the average height, average reflection intensity, and number of points within that voxel as statistical features of that voxel, thereby achieving voxel-based downsampling of the point cloud and obtaining basic lidar heatmap data in the form of a regular grid. Extreme outliers within the grid are removed using mean filtering or median filtering to effectively filter out transient interference such as birds, rain, and snow. 2) Ground segmentation: Using RANSAC plane fitting or progressive morphological filtering methods, estimating the main ground plane from the voxel point cloud, classifying points with heights close to this plane as ground points, and classifying the remaining points as non-ground points. Ground segmentation can eliminate interference from large areas of flat ground, making subsequent targets such as engineering vehicle outlines, tree obstacles, and equipment more prominent. 3) Initial Static / Dynamic Target Separation: A time-series comparison is performed on multiple consecutive frames of LiDAR point clouds. Under a unified coordinate system, the occupancy of each voxel in adjacent frames is statistically analyzed. If a voxel remains stable within a long time window with minimal shape change, it is marked as a "static background" (e.g., towers, permanent buildings). If a voxel appears within a short time or its surrounding point cloud shape changes significantly, it is marked as a "dynamic candidate target" (e.g., moving construction vehicles, swaying tree branches). Through this initial static / dynamic separation, when generating a 32×32 resolution LiDAR heatmap, a higher weight can be applied to dynamic candidate regions, providing prior information for subsequent abnormal target detection. 4) Heatmap Generation: The processed point cloud is projected onto a two-dimensional plane in a ground coordinate system. A 32×32 LiDAR heatmap is constructed based on the point cloud height, density, or reflection intensity of each grid cell. 32×32 represents the spatial grid resolution, and each pixel corresponds to a fixed-area transmission channel region.

[0031] In one embodiment, to convert continuous sound data into a compact representation that can be fused with visual and point cloud features, the process of extracting the spectrum of the sound data may include: 1) Endpoint detection: The original sound waveform is divided into frames with a fixed frame length (e.g., 50ms) and an overlap rate (e.g., 50%), and the short-time energy and zero-crossing rate of each frame are calculated; by setting dual thresholds for energy and zero-crossing rate, speech activity segments containing significant events (such as arc discharge, mechanical noise, and combustion explosion sounds) are detected, and pure noise segments are eliminated to reduce redundant calculations in subsequent feature extraction. 2) Short-time energy and Mel spectrum extraction: For valid frames that pass the endpoint detection, their short-time energy is first calculated as a temporal feature reflecting the overall sound intensity; then, each frame is windowed (e.g., Hamming window) and subjected to Fast Fourier Transform (FFT) to obtain the spectral amplitude, and after weighting and logarithmic operation by the Mel filter bank, a log-Mel spectrum is obtained, which retains the frequency band distribution characteristics close to human auditory perception. 3) Feature encoding: The log-Mel spectra of multiple consecutive frames are stacked in chronological order into a two-dimensional time-frequency map, and then the features are compressed by a one-dimensional or two-dimensional convolutional neural network to output a fixed-dimensional sound spectrum code (e.g., 64-dimensional), which serves as a vector representation describing the acoustic events within the current time window.

[0032] As an optional implementation, the dual-branch CNN deep learning model includes a first branch, a second branch, and a feature interaction layer. Based on this, step 300 includes: using the first branch to obtain visible light image features based on the registered visible light image; using the second branch to obtain infrared image features based on the registered infrared image; and using the feature interaction layer to obtain a joint feature map based on the visible light image features and the infrared image features. Specifically, the visible light image features extracted by the first branch are the texture features in the registered visible light image. The infrared image features extracted by the second branch are the temperature features in the registered infrared image. The feature interaction layer can employ a cross-attention module to calculate the mutual attention weights of the visible light image features and the infrared image features, perform spatial alignment enhancement, and the final output joint feature map includes fused features of complementary information from both visible light and infrared modes.

[0033] The implementation process of step 500 includes: weighted fusion of the joint feature map, lidar heat map and sound spectrum coding to obtain fused data.

[0034] For example, LiDAR heatmaps and acoustic spectral encodings are injected into the visual backbone (i.e., joint feature maps) using a spatially biased cross-modal attention + channel-wise Linear Modulation (FiLM) gating + dynamic weight fusion approach to obtain fused data. Figure 3 As shown, the fused data is input into the dual-task detection head, which outputs vehicle detection frames and smoke / fire segmentation masks in parallel.

[0035] The fusion of multimodal data includes two levels: feature-level fusion and decision-level fusion, both of which adopt a dynamic weight allocation mechanism.

[0036] In feature-level fusion, the lidar heatmap and sound spectrum encoding are injected into the joint feature map through cross-modal attention and channel FiLM gating: 1) A lightweight quality assessment network is used to estimate the confidence weights of each modality based on indicators such as the mean brightness, contrast, and haze index of the visible light image, the effective temperature dynamic range of the infrared image, the effective grid ratio of the lidar heatmap, and the signal-to-noise ratio of the sound spectrum. 2) The confidence weights are normalized by Softmax to obtain the feature-level fusion weights, which are then used as scaling and translation parameters for FiLM gating. Channel-by-channel linear modulation is performed on each modal feature channel, and then weighted summation is performed in the spatial dimension to obtain an environment-adaptive fusion feature map.

[0037] In decision-level fusion, for the vehicle detection bounding boxes and smoke / fire segmentation masks output by the dual-task detection heads, auxiliary evidence from LiDAR and the audio channel is superimposed while retaining the confidence level of the visual backbone output. For example, if a vehicle detection bounding box has obvious 3D structural echoes in the corresponding area of ​​the LiDAR heatmap, the final confidence level of the bounding box is increased; if a smoke / fire segmentation mask exhibits arc or deflagration frequency band energy in the audio spectrum, the risk score of the corresponding alarm is improved. The final risk score is obtained by weighted summation of the visual confidence level and the confidence weights of each modality, and is used to drive subsequent graded warnings. For engineering vehicle targets, clustering and geometric fitting are first performed in the 3D point cloud corresponding to the LiDAR heatmap to estimate the 3D bounding boxes and attitude angles of the vehicle body and robotic arm, obtaining attitude vectors representing the vehicle's spatial pose. Then, the attitude vectors are mapped to a set of "attitude query" vectors through a fully connected layer, and cross-modal attention calculation is performed with the spatial feature map of the visual backbone, thereby enhancing the response of the region consistent with the vehicle structure and suppressing background interference. For fireworks targets, the sound spectrum encoding is input into the voiceprint recognition subnetwork to extract high-dimensional voiceprint feature vectors that represent arc discharge, mechanical noise and combustion explosion sounds. The channel weights related to smoke texture and bright flame areas in the visual feature map are adjusted by channel attention to enhance the fireworks response when there is acoustic evidence and reduce false positives caused by smog or noise alone.

[0038] Cross-modal attention fusion not only uses lidar heatmaps and acoustic spectral coding to weight visual features, but also explicitly utilizes point cloud pose estimation and voiceprint features to assist in the discrimination of engineering vehicles and fireworks. Through the cross-modal attention injection of point cloud pose estimation and voiceprint spectrum, the network can still maintain robust detection of engineering vehicles and fireworks by relying on geometric and acoustic information even when visual information is insufficient or occluded.

[0039] As an optional implementation, the dual-task detection head includes a shared backbone network, a feature pyramid network, a first task branch, and a second task branch. The fused data is input into the dual-task detection head to obtain target detection results, including: using the shared backbone network to obtain multi-scale feature maps based on the fused data; using the feature pyramid network to obtain a shared feature map based on the multi-scale feature maps; using the first task branch to obtain vehicle detection boxes based on the shared feature maps; and using the second task branch to obtain a smoke / fire segmentation mask based on the shared feature maps.

[0040] For example, fused data can correspond to a set of multi-scale fused maps (e.g., spliced ​​from feature maps at three scales: 1 / 8, 1 / 16, and 1 / 32, or fused from top to bottom). In the process of obtaining target detection results from the fused data using a dual-task detection head: the fused data is input into a shared backbone network, passing through multiple convolutional layers and downsampling to obtain multi-scale feature maps. These multi-scale feature maps are then input into a feature pyramid network, where they undergo top-down feature pyramid encoding. Each scale feature map sequentially passes through a 3×3 convolutional layer, a normalization layer, and a non-linear activation layer to increase the receptive field of the context. Upsampling and element-wise addition are used to fuse high-level semantics with low-level details, enhancing the simultaneous perception capability of large-scale engineering vehicles and small-scale inspection targets, resulting in a shared feature map.

[0041] In the first task branch, for the shared feature map at each scale, two (or three) 3×3 convolutional layers are concatenated for feature refinement, and then a 1×1 convolutional layer outputs the bounding box and class prediction parameters at each grid location. Each spatial location outputs a (4+1+Cveh) dimensional vector, where 4 dimensions are the bounding box regression parameters (center coordinates, width, and height), 1 dimension is the target presence confidence, and Cveh is the classification probability for the relevant category of the engineering vehicle (e.g., ordinary vehicle, vehicle with robotic arm, etc.). After NMS (Non-Maximum Suppression), a set of vehicle detection boxes (including location and category) is obtained.

[0042] In the second task branch, shared feature maps at multiple scales are converged to a uniform resolution (e.g., 1 / 4 of the original image size) through upsampling and concatenation. Then, they are sequentially passed through multiple 3×3 convolutional layers and 1×1 convolutional layers to compress the number of channels to a 1-channel segmentation logits map. Finally, bilinear interpolation (or deconvolution upsampling) is used to obtain the same output resolution as the input image (i.e., the shared feature maps) or a preset resolution. After thresholding, the fireworks segmentation mask is obtained.

[0043] Optionally, a branch (a third task branch, i.e., a key point detection branch) can be added to the dual-task detection head, enabling the power transmission channel abnormal target detection method provided in this application to also detect key points of vehicles within the vehicle detection frame. In the third task branch, RoI alignment and local convolutional encoding are further performed on the fused features within the vehicle detection frame, and the coordinate offset of a preset number of key points is predicted through 1×1 convolution to obtain the positions of key points of the engineering vehicle's robotic arm or body.

[0044] Furthermore, the abnormal target detection method for power transmission channels based on multimodal sensors and deep learning also includes: constructing an initial dual-task detection head; acquiring sample multimodal data and corresponding sample target results; the sample multimodal data includes sample visible light data, sample infrared data, sample lidar data, and sample sound data; performing data augmentation processing on the sample visible light data, sample infrared data, sample lidar data, and sample sound data respectively to obtain augmented sample visible light data, augmented sample infrared data, augmented sample lidar data, and augmented sample sound data; using a dual-branch CNN deep learning model, obtaining a sample joint feature map based on the augmented sample visible light data and augmented sample infrared data; obtaining a sample lidar heatmap based on the augmented sample lidar data; obtaining sample sound spectral encoding based on the augmented sample sound data; obtaining sample fusion data based on the sample joint feature map, sample lidar heatmap, and sample sound spectral encoding; labeling the sample fusion data based on the sample target results corresponding to the sample multimodal data to obtain labeled targets; and constructing a training dataset based on the sample fusion data and labeled targets. The initial dual-task detection head is trained using the training dataset until the loss function value reaches the set requirement, thus obtaining the dual-task detection head.

[0045] For the first task branch (vehicle detection branch), IoU loss and classification loss are used as loss functions; for the second task branch (smoke and fire segmentation branch), cross-entropy loss and Dice loss are used as loss functions. During the data augmentation process for the multimodal sample data, parametric simulations are performed to construct 12 typical extreme weather scenarios, including strong light, backlight, rain, snow, fog, and low nighttime illumination. This allows the initial dual-task detection head to incorporate these scenarios during training, thus providing targeted robustness in practical applications. The 12 extreme weather scenarios can be categorized by combinations of illumination, visibility, and precipitation type, for example: A1~A3: direct sunlight during the day, strong reflection from glass / water surfaces, and deep shadows; B1~B3: low visibility caused by moderate / heavy fog, haze, and dust storms; C1~C4: light rain / heavy rain, light snow / heavy snow (including strong reflection from snow accumulation on the ground); D1~D2: low nighttime illumination and localized strong light at night (vehicle lights, firelight, etc.). By applying corresponding enhancement operators to visible light, infrared, lidar, and sound data respectively, and labeling them with environmental tags based on the combination of enhancement parameters, a sample set covering the above 12 extreme environments is finally obtained.

[0046] In one implementation, the following data augmentation strategies are employed for the sample visible light data and sample infrared data: 1) Strong light / backlight simulation (corresponding to A1~A3, D2 class scenes): High-brightness spots or lensflares are randomly selected in the image to increase local brightness to near saturation, used to simulate direct sunlight, headlight glare, snow reflection, etc. Nonlinear brightness stretching and contrast compression are applied to the overall image to create backlight and large-area shadow effects. Sidelight and backlight scenes in the early morning and late evening are simulated through gamma transformation and local contrast adjustment. 2) Haze / dust synthesis (corresponding to B1~B3 class scenes): Based on an atmospheric scattering model, the contrast of each pixel is approximately attenuated by depth, and atmospheric light components are superimposed to generate haze / fog effects of different concentrations. Color shift (yellowish / grayish) and low-frequency noise superposition are used to simulate the color shift and blurring of dusty weather. The target bounding box remains unchanged, enabling the network to learn to locate the target even in low-contrast backgrounds. 3) Rain and Snow Simulation (corresponding to C1~C4 scenes): Generate directional linear rain streaks, overlay them on the image, and add slight motion blur to simulate light rain / heavy rain. Generate randomly distributed bright snowflake dots or sheet-like textures, and control the density and size to simulate light snow / heavy snow, as well as snow cover on the ground. Slightly increase the overall brightness to simulate the reflective characteristics of snowy scenes. 4) Night / Low Light and Sensor Noise (corresponding to D1~D2 scenes): Reduce the overall brightness and saturation, increase high ISO noise and slight blur to simulate night and dusk scenes. For infrared images, adjust the grayscale dynamic range and gamma curve to enhance the contrast in high-temperature areas, and add a small amount of random noise to the background to simulate the drift of infrared sensors under high / low temperature environments.

[0047] One or more of the above enhancements can be randomly selected for each original sample and superimposed, and the corresponding environmental labels can be recorded so that the synthetic samples in the sample set statistically cover the above 12 extreme weather scenarios.

[0048] Enhancement of sample LiDAR point cloud data can supplement robustness in rain / snow / occlusion scenarios. In one implementation, for sample LiDAR point cloud data synchronized with images (sample visible light data and sample infrared data), the following enhancement strategies are adopted: 1) Simulation of point loss and misplacement caused by rain and snow: Randomly drop points in the point cloud according to the probability set by distance and pitch angle to simulate echo attenuation caused by raindrops / snowflakes. Randomly add a small number of isolated points in local areas to simulate rain and snow noise and small falling objects. 2) Local occlusion and viewpoint changes: Randomly hollow out / occlude some voxels in the area near the ground to simulate the occlusion of the line of sight by vegetation, vehicles, snow, etc. Apply a small random rotation and translation to the overall point cloud to simulate changes in device attitude and wind shaking. 3) Random jitter at the heatmap level: Introduce a small amount of Gaussian noise and local smoothing on the 32×32 heatmap generated from the point cloud to simulate changes in echo intensity and density distribution under different weather conditions.

[0049] With the above enhancements, even in real-world scenarios with limited samples in extreme rain and snow, the dual-task detection head can still learn robust geometric feature representations during the training phase by simulating point loss and occlusion.

[0050] In one embodiment, for synchronously acquired sample sound data, data enhancement processing is performed before extracting short-time energy and Mel spectrum (i.e., obtaining sample sound spectrum encoding): 1) Background noise mixing: Randomly select several noises from a pre-recorded library of wind noise, rain noise, insect chirping, and traffic noise, and mix them with the original signal according to a set signal-to-noise ratio to simulate complex sound environments such as strong winds, heavy rain, and mountainous environments. 2) Gain and frequency band filtering: Randomly adjust the overall volume gain and the attenuation / enhancement of certain frequency bands to simulate different distances, different device sensitivities, and obstruction conditions. 3) Time stretching and slight distortion: Perform small-scale time stretching or phase perturbation on the waveform to simulate sampling clock errors and device jitter.

[0051] The above enhancements ensure that acoustic features such as arc discharge, engine noise, and combustion sound can still be distinguished by the model under various noise backgrounds, and supplement extreme environmental scenarios such as strong winds and rain, and complex background noise.

[0052] When constructing the sample set (including enhanced visible light data, enhanced infrared data, enhanced lidar data, and enhanced sound data), the proportion of various enhancement combinations is controlled to ensure that each type of extreme environment scenario has a sufficient number of synthetic samples in the training set to participate in parameter learning, thereby technically achieving coverage and robustness improvement for 12 typical extreme weather scenarios.

[0053] Based on the same inventive concept, this application also provides a device for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning, used to implement the aforementioned method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations in the embodiments of the device for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning provided below can be found in the limitations of the method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning described above, and will not be repeated here.

[0054] In one exemplary embodiment, such as Figure 4 As shown, a power transmission channel abnormal target detection device based on multimodal sensors and deep learning is provided, including: a multimodal sensor unit, a data preprocessing unit, a data fusion unit, and a detection unit.

[0055] The multimodal sensor unit is used to acquire multimodal data. Multimodal data includes visible light data, infrared data, lidar data, and sound data.

[0056] The data preprocessing unit is connected to the multimodal data acquisition unit. The data preprocessing unit is used to register visible light and infrared data, obtaining registered visible light and infrared images. It also uses a two-branch CNN deep learning model to obtain a joint feature map based on the registered visible light and infrared images. Furthermore, the data preprocessing unit preprocesses LiDAR data to obtain a LiDAR heatmap. Finally, it obtains sound spectrum encoding based on sound data.

[0057] The data fusion unit is connected to the data preprocessing unit. The data fusion unit is used to obtain fused data based on joint feature maps, lidar heat maps, and acoustic spectral coding.

[0058] The detection unit is connected to the data fusion unit. The detection unit is used to obtain target detection results based on the fused data. The target detection results include: vehicle detection boxes and smoke / fire segmentation masks.

[0059] As an optional implementation, the multimodal sensor unit includes: a visible light camera, an infrared thermal imager, a lidar, a directional microphone array, and a spatiotemporal alignment module.

[0060] The visible light camera, infrared thermal imager, lidar, and directional microphone array are all connected to the spatiotemporal alignment module. The spatiotemporal alignment module is connected to the data preprocessing unit.

[0061] A visible light camera is used to acquire raw visible light data. An infrared thermal imager is used to acquire raw infrared data. A lidar is used to acquire raw lidar data. A directional microphone array is used to acquire raw sound data. A spatiotemporal alignment module is used to perform spatiotemporal alignment processing on the raw visible light data, raw infrared data, raw lidar data, and raw sound data to obtain visible light data, infrared data, lidar data, and sound data.

[0062] In one embodiment, before collecting multimodal data, a visible light camera (in this embodiment, a camera with at least 2 megapixels), an infrared thermal imager (temperature measurement range -20℃ to 550℃), a lidar (in this embodiment, a 16-line lidar), and a directional microphone array are deployed. 1. Visible Light Camera. Acquisition Location: Deployed on the crossarm of the transmission tower, with a downward viewing angle of 30°~45°, covering a 100-meter channel range on both sides of the conductor. Data Type: RGB three-channel video stream. Target Objects: Engineering vehicles (license plates, robotic arm posture); fireworks (flame color, smoke pattern); foreign objects on the conductor (kite strings, plastic film). Data Transmission Method: Directly connected to the spatiotemporal alignment module via gigabit Ethernet; synchronous trigger signals are received via the PTP protocol. 2. Infrared Thermal Imager. Acquisition Location: Coaxially mounted with the visible light camera, optical axis deviation <0.1°. Data Type: Thermal radiation matrix (640×512 pixels, temperature measurement accuracy ±2℃). Target objects: Fireworks (temperature anomaly area, alarm triggered >150℃); equipment overheating (local temperature rise of insulators >10℃ / min); biological activity (human body temperature characteristics). Data transmission method: Transmitted to the spatiotemporal alignment module via coaxial cable + Ethernet; spatiotemporally aligned with visible light data. 3. 16-line LiDAR. Acquisition location: Horizontally installed on the top of the transmission tower, scanning plane parallel to the conductor direction. Data type: 3D point cloud (16-layer scan, angular resolution 0.1°, maximum ranging 200m). Target objects: Engineering vehicles (3D contour, boom extension distance); conductor sag (point cloud fitting sag); tree obstacles (vegetation intrusion distance). Data transmission method: Connected to an industrial switch (PTP master clock) via RS-422 interface; the switch forwards to the spatiotemporal alignment module. 4. Directional microphone array. Acquisition location: 1.5 meters below and to the side of the conductor, 60° elevation angle pointing towards the conductor. Data type: Sound pressure waveform (sampling rate 48kHz, dynamic range 96dB). Target objects: electric arc discharge (3-15kHz high-frequency pulse); mechanical noise (crane engine acoustic spectrum characteristics); combustion crackling sound (low-frequency transient signal). Data transmission method: connected to the audio acquisition card via shielded audio cable (anti-electromagnetic interference); transmitted to the time-space alignment module via USB 3.0.

[0063] In addition, the data preprocessing unit, data fusion unit, and detection unit can be deployed inside the waterproof box of the transmission tower maintenance platform.

[0064] As an optional implementation, the power transmission channel anomaly target detection device based on multimodal sensors and deep learning further includes a cloud platform. The cloud platform is connected to the detection unit. The cloud platform is used for visualizing the detection results.

[0065] The cloud platform can display real-time visualizations: vehicle detection bounding boxes and key points, smoke / fire segmentation masks can be overlaid on video / IR images. It can also display target trajectories and speed / direction arrows.

[0066] In one embodiment, the power transmission channel abnormal target detection device based on multimodal sensors and deep learning further includes a local adjudication unit. The local adjudication unit is connected to both the cloud platform and the detection unit. The local adjudication unit performs event structuring and local adjudication based on the target detection results, including: 1) Frame-by-frame candidate generation: The target detection results output by the detection unit are traversed frame by frame. Vehicle detection boxes and smoke / fire segmentation masks with a confidence level exceeding a first threshold (0.3 in this embodiment) are recorded as candidate targets, and their corresponding LiDAR geometric quantities (minimum distance, velocity), infrared temperature peak, sound evidence score, and other multimodal features are associated. 2) Spatiotemporal association and trajectory generation (spatiotemporal continuity): For multiple consecutive frames of candidate targets, association is performed based on the IoU of the detection boxes in the image, their spatial location in the geographic coordinate system, and the time interval. Candidates belonging to the same engineering vehicle or the same smoke / fire area are merged into a single trajectory, resulting in a target trajectory containing start and end times, movement paths, and a multimodal feature time series. 3) Multimodal Consistency Assessment: Along each target trajectory, visual confidence, LiDAR distance, infrared temperature, and acoustic evidence are aggregated to calculate a multimodal consistency score. For example, when the engineering vehicle has high detection confidence in the visible light image, the corresponding LiDAR point cloud shows a clear three-dimensional structure, and engine noise characteristics are present in the acoustic signal, the multimodal consistency score is improved. If a low-confidence response only occurs briefly in a single modality, the consistency score is low. 4) Event Triggering and Hysteresis Decision: For each target trajectory, a risk score is calculated by combining model confidence, multimodal consistency score, minimum distance to the conductor / tower, and smoke / fire area growth rate. When the risk score exceeds a high threshold (set according to actual conditions) for several consecutive frames, the event is triggered to "start"; when the risk score is below a low threshold (set according to actual conditions) for several consecutive frames, the event is determined to "end." This hysteresis mechanism, using high and low thresholds and the number of consecutive frames, suppresses jitter alarms caused by single-frame fluctuations. 5) Event Deduplication and Merging: Merge events that are similar in time and space, of the same type, and with a time interval less than a set value. For example, if the same vehicle moves between adjacent towers in a short period of time, only one "Engineering Vehicle Intrusion Event" is generated to avoid duplicate reporting by multiple towers along the line. 6) Event Schema Generation: For each event that passes the adjudication, generate a structured event record (event card) containing: event ID, event type (vehicle / intrusion / tree obstruction / smoke / fire / electric arc, etc.), occurrence line and tower number, latitude and longitude / mileage marker location, start / end time, multimodal confidence and risk score and risk level, minimum distance to conductor / tower, peak temperature, acoustic evidence score, and reference addresses of evidence such as keyframe screenshots, segmentation masks, point cloud slices, and spectrograms.

[0067] Through steps 1) to 6) above, the continuous frame-by-frame detection results are transformed into deduplicated structured events in the cloud. The detection bounding box, segmentation mask, and multimodal features of each frame are aggregated into a small number of "event cards" (e.g., a construction vehicle staying near a pole on a certain line for 30 seconds becomes an construction vehicle intrusion event). Combining the confidence level, multimodal consistency, and spatiotemporal continuity of the target detection results, it is first determined whether it is a "real hidden danger event." Hysteresis / jitter suppression is used to avoid "flickering alarms" caused by single-frame fluctuations, reducing false alarms and frequent reporting, and significantly reducing duplicate alarms and data volume. A preliminary risk assessment is completed, and only the screened events and necessary evidence are uploaded to the cloud, reducing bandwidth consumption and cloud storage pressure, which helps to reduce false alarm rate and data redundancy, and improve alarm stability and engineering availability.

[0068] In one implementation, after structuring and adjudicating events locally based on target detection results in the local adjudication unit, the structured event records are securely reported and cached to the access layer network management of the cloud platform. This includes the following steps: 1) Evidence Packaging: When the local adjudication unit generates a new structured event, it extracts visible light / infrared image frames, corresponding smoke and fire segmentation masks, lidar thermal maps and point cloud slices, and sound waveforms / spectral maps from the local circular buffer several seconds before and after the event. These are then compressed and encoded (e.g., H.264 short videos, PNG screenshots, binary point cloud files, and spectrum matrices), stored in the local file system or object storage, and the path or hash value of the evidence file is written into the event record. 2) Secure Encapsulation and Encrypted Transmission: The local adjudication unit can generate a unique event ID and device ID for each event, encapsulate the event record and evidence summary into a JSON / Protobuf message according to a predefined schema, and send it to the cloud access gateway via a TLS-encrypted HTTPS or MQTT channel, along with device authentication information (such as certificates, signatures, or tokens) to prevent data from being eavesdropped on or tampered with. 3) Event Triggering Conditions: Reporting is triggered immediately only when the event risk level reaches a preset threshold (e.g., medium risk or above) or the event type belongs to a high-risk category (e.g., intrusion of engineering vehicles, open flame, arc discharge). Low-risk events can be batched or reported at certain time intervals to balance real-time performance and bandwidth consumption. 4) Local Caching and Retransmission: When there is a network anomaly or the cloud access layer returns a failure, the edge node records the unreported events and evidence in a local cache queue, including the event ID, retry count, and last attempt time. The background thread periodically checks the network recovery status and the cloud ACK status, retransmitting unconfirmed events until successful or the maximum number of retryes is exceeded. When the cache space is close to its limit, high-risk events are prioritized, while low-risk or expired data is eliminated in chronological order. Through the above secure reporting and caching mechanisms, it can be ensured that abnormal events of important power transmission channels are not lost in weak network or outage scenarios, and the reporting process meets encryption and authentication requirements.

[0069] In one implementation, the structured events uploaded are standardized at the access layer of the cloud platform, including the following steps: 1) Gateway authentication and connection management: The cloud access gateway maintains the device IDs, certificates, or keys of all legitimate edge nodes, performs bidirectional TLS handshakes and device identity authentication for new connections, and rejects connections from unregistered or invalid devices to prevent unauthorized devices from forging alarms. 2) Format verification and field integrity check: For each reported message, it is parsed according to the predefined event schema to check whether the required fields (event ID, device ID, event type, time, location, risk level, etc.) exist and whether the field types and value ranges conform to the specifications; messages lacking key fields or with incorrect formats are rejected or recorded in the alarm log to prevent dirty data from entering the subsequent decision-making chain. 3) Clock alignment and time standardization: The local timestamp carried by the local decision unit is converted to a unified UTC or power grid standard time by correcting the difference between the local timestamp and the cloud time; for delayed reported events caused by cached retransmission, the "event occurrence time" and "receive time" fields are retained to ensure that the actual occurrence time is used in subsequent time series analysis. 4) Idempotent Deduplication: To avoid duplicate data entry caused by repeated reporting or retries by local adjudication units, the access layer uses an idempotent key formed by combining event ID, device ID, and time window to detect duplicate messages. If the same event ID has already been successfully written to disk, subsequent duplicate messages are discarded, and only necessary status fields (such as the number of retries) are updated. This ensures that the cloud retains only one standard record for each real event. 5) Disk Writing and Message Queue Inbound: For events that pass verification, they are written to the cloud event database and file / object storage according to a standardized schema (evidence information is stored only as references and hashes). At the same time, the event summary is delivered to a message queue (such as Kafka) for the upper-layer "Multi-Source Association and Risk Assessment" module to subscribe to and process. Through standardized processing at the access layer, unified access and format standardization of data reported by different local adjudication units and different collection sites can be achieved, avoiding duplicate records and data inconsistencies, and providing the subsequent decision-making layer with a high-quality, time-consistent data source for abnormal power transmission channel events.

[0070] Based on the standardized structured alarm events at the access layer, decisions are made at the decision layer in the cloud platform. This forms a closed-loop optimization mechanism of "detection → adjudication → handling → feedback → retraining" with the training of the dual-task detection head. The dual-task detection head is continuously optimized using the feedback results, further reducing the false alarm rate and false alarm rate of power transmission channel anomaly detection.

[0071] In one implementation, the decision-making layer and closed-loop processing include the following steps: 1) Multi-source association: Clustering and associating events from different towers, different equipment, and even different systems in the time and space dimensions. For example, associating "fire events" from multiple towers along the same line in the same time period with fire information from external forest fire monitoring platforms to determine that they are the same wildfire; connecting multiple "engineering vehicle events" on the same route according to their trajectories to identify vehicle movement trajectories and operational trends. 2) Risk assessment: Based on multimodal physical quantities such as minimum distance to conductors / towers, vehicle speed, robotic arm posture angle, fire area and growth trend, infrared temperature peak value, and sound evidence score in the structured records of events, combined with line voltage level, importance, and sensitivity to the surrounding environment (such as whether it crosses residential areas or forest areas), risk scores are calculated through rule scoring or statistical learning models, and events are classified into high, medium, and low risk levels. 3) Rule Engine Adjudication: Triggers corresponding handling strategies based on risk level and event type. For example, when the minimum distance of an "engineering vehicle intrusion event" is less than the safety threshold and the risk level is high, the rule engine automatically generates handling suggestions such as "lockdown / load reduction / emergency dispatch." When a "fire event" shows a continuous increase in temperature and an expansion of the smoke area, it triggers the "forest fire prevention early warning" and "on-site verification" processes. When the event may be related to planned maintenance or construction, it lowers the risk level or converts it to a "patrol inclusion" suggestion. 4) Command and Dispatch and Work Order Assignment: The adjudication results are pushed to the dispatch, security, and operation and maintenance work order systems via interfaces, automatically generating work orders, specifying responsible teams and handling deadlines, and recording work order status (accepted / in transit / completed) and closed-loop results. In the digital twin scenario, the spatial distribution and progress of various alarms and work orders are displayed using colors and icons. 5) Visualized Alarms and Human-Machine Collaborative Review: On the cloud platform interface, event cards, target trajectories, speed / direction arrows, and smoke / fire separation masks are overlaid onto a 3D line model or GIS map. Maintenance personnel can click to view the evidence chain and short video playback, perform operations such as "confirmation / rejection / reclassification" of events, and generate manual review results to correct handling decisions and record false alarms. 6) Feedback and Model Iteration: Manual review results, work order closed-loop results, and report statistics (false alarm rate, response time, completion rate, and other KPIs) are fed back into the model (i.e., dual-task detection head) training and strategy configuration to update data augmentation strategies, adjust multimodal weights and risk thresholds, thereby forming a closed-loop optimization mechanism of "detection → adjudication → handling → feedback → retraining".

[0072] Through the aforementioned decision-making layer and closed-loop processing, this application can not only identify abnormal targets in power transmission channels under complex environments, but also transform the identification results into specific operation and maintenance decisions and on-site handling actions, thereby achieving closed-loop management of power transmission channel safety risks.

[0073] Based on the above embodiments, the final results obtainable on the cloud platform include: 1) Spatial and Physical Quantities: ① Spatial Positioning and Coordinate Unification: For each vehicle detection frame and smoke / fire segmentation mask, the image pixel coordinates are back-projected onto the LiDAR coordinate system using camera calibration parameters. Combined with tower coordinates, route path, and mileage information, the latitude and longitude position of the target in the geographic coordinate system and its mileage along the route are calculated, while also associating it with the route and tower number. ② Geometric Quantity Calculation: Under the unified coordinate system, the minimum distance between the target and the nearest conductor, tower, and passage boundary is calculated using LiDAR point cloud or heat map. The vehicle's length, width, height, and ground clearance are calculated, as well as its displacement between consecutive frames to obtain velocity and direction of motion. For smoke / fire targets, the smoke / fire coverage area and its rate of change over time are obtained based on the projected area of ​​the segmentation mask in the geographic coordinate system. ③ Physical Quantity Extraction: Combining infrared images, thermophysical quantities such as peak temperature, average temperature, and temperature gradient are statistically analyzed within the target area. Combining sound spectrum coding, feature scores such as arc discharge, mechanical noise, and combustion explosion sounds are extracted as acoustic evidence. Through steps ① to ③, the target detection results, which only include vehicle detection frames and smoke and fire segmentation masks, are transformed into target information with spatial location and physical quantity descriptions, providing basic data for subsequent event structuring and risk assessment.

[0074] 2) Structured Alarm Events: Based on the aforementioned spatial and physical quantities, the detection results of consecutive frames are spatiotemporally aggregated to form structured alarm events, including: ① Candidate Target Screening: Vehicles and fireworks targets with a confidence level higher than the first threshold in each frame's detection results are marked as "candidate targets" along with their spatial location and physical quantities. ② Spatiotemporal Association and Trajectory Generation: Based on the overlap (IoU), spatial distance, time interval, and direction of movement of candidate targets in adjacent frames, candidate targets belonging to the same engineering vehicle or the same fireworks area are associated as a target trajectory, and the start time, end time, and path are recorded. ③ Multimodal Consistency Assessment and Risk Scoring: Combining visual confidence level, minimum distance to conductors / towers, peak temperature, fireworks area growth rate, acoustic evidence, etc., a risk score and multimodal consistency score are calculated for each trajectory. ④ Hysteresis Threshold + Jitter Suppression Adjudication: Two threshold levels are set: when the risk score of a trajectory is higher than the high threshold for several consecutive frames, the event is considered to have started; only when the risk score is lower than the low threshold for several consecutive frames is the event considered to have ended, thus suppressing alarm jitter caused by single-frame fluctuations. ⑤ Event Deduplication and Merging: Events that are close in time, spatially adjacent, and of the same type are merged. For example, if the same vehicle moves between adjacent towers, only one "construction vehicle intrusion event" is retained. ⑥ Structured Event Record Generation: For each adjudicated event, a structured alarm event record is generated, including: event ID, event type (construction vehicle intrusion, road obstruction, smoke, etc.), line / tower number, latitude / longitude / mileage location, event start and end time, spatial and physical quantities (minimum distance, speed, area, temperature, acoustic score, etc.), risk score, and risk level metadata.

[0075] 3) Evidence Chain and Evidence Package: To support evidence collection and post-event tracing, an evidence chain and evidence package are constructed based on structured events. Specifically, this includes: ① Key Fragment Selection: Based on the start and end times of the event, visible light / infrared images, LiDAR point clouds or heat maps, and sound waveforms or spectrograms within a preset time window before and after the event are extracted from locally cached multimodal data. Detection boxes, segmentation masks, and key points are overlaid on the images, and local regions containing the target are cut out from the point cloud. ② Multimodal Evidence Packaging: The above multimodal fragments are compressed into short videos, images, point cloud slice files, and acoustic feature files, generating corresponding file identifiers and integrity verification information (such as hash values ​​or signatures), and packaged into evidence packages according to a unified structure. ③ Evidence Chain Association: The unique identifier of the evidence package and its contained multimodal evidence list are stored in the structured event record, realizing a chain index of "event → evidence," ensuring that the event process can be completely reconstructed during subsequent manual review, responsibility allocation, and legal evidence collection.

[0076] 4) Handling Recommendations and Work Order Status: The cloud platform performs risk assessments and makes decisions on structured alarm events, linking with the work order system to generate handling recommendations and work order statuses: ① Cloud-based Risk Assessment and Decision-making: Based on spatial and physical quantities in the event (such as minimum distance, speed, attitude, smoke and fire area and growth rate, temperature, etc.) and static information such as line voltage level, importance, and surrounding environment, the cloud uses a rule engine or risk assessment model to calculate the comprehensive risk level and selects appropriate handling strategies (such as lockdown, load limiting, on-site verification, arranging patrols, monitoring and tracking, etc.). ② Handling Recommendation Generation: The risk assessment results are transformed into structured handling recommendations, including: recommended measure type, recommended completion deadline, recommended responsible unit, etc., and displayed to dispatch / maintenance personnel in the form of alarm cards. ③ Work Order Generation and Issuance: Disposal suggestions are pushed to the operation and maintenance work order system via an interface, automatically generating work orders. The work order number, associated event ID, task content, responsible team, planned start / end time, etc., are recorded, and the work order status (generated, accepted, in progress, completed, closed, etc.) is linked to the event records. ④ Work Order Status Write-back and Closed Loop: As on-site operations progress, status changes in the work order system and on-site feedback (such as "on-site verification confirms normal construction" or "tree obstruction cleared") are written back to the cloud platform, updating the corresponding event handling status and results, achieving closed-loop management from detection to handling.

[0077] 5) Reporting and Statistics: The cloud platform regularly performs statistical analysis on accumulated structured events, work orders, and handling results, generating reports and visualized statistical results, including: ① Data Summary: Based on structured alarm events and work order records, fields such as event type, risk level, line / tower distribution, occurrence time, handling duration, handling result, and false alarm / missed alarm flags are written into the statistical database. ② Indicator Calculation: The frequency of events such as engineering vehicle intrusion, smoke, and tree obstruction is statistically analyzed by line, region, and time dimensions. Average response time, handling completion rate, false alarm rate, and distribution by level are calculated to form multi-dimensional statistical results. ③ Reporting and Visualization: The statistical results are generated into regular reports (such as monthly, quarterly, and annual reports) and displayed in the visualization interface in the form of line charts, bar charts, heat maps, etc., enabling operation and maintenance personnel to grasp the distribution of abnormal situations in the transmission channel and the effectiveness of management from a macro perspective. ④ Statistical results are used for model optimization: events related to false positives / false negatives are categorized and analyzed, serving as the basis for subsequent threshold adjustments, optimization of multimodal weights, updating of data augmentation strategies, and retraining of the model, thereby achieving data-driven continuous optimization.

[0078] The above outputs are displayed visually on the cloud platform and also provided externally via API / MQ, facilitating integration with scheduling / security / work order systems.

[0079] In conjunction with the above embodiments, the method and apparatus for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning provided in this application have the following advantages: 1. Improved detection accuracy: The accuracy of smoke and fire recognition in smog environment is improved by 32% (actual F1-score from 0.68 to 0.90); the posture estimation error of the robotic arm of engineering vehicle is ≤5°.

[0080] 2. Real-time optimization: Data preprocessing unit processing latency <200ms (1080P video stream); sound signal assistance reduces the false alarm rate of fireworks by 45%.

[0081] 3. Enhanced adaptability: Supports an ambient temperature range of -30℃ to 60℃; data enhancement covers 12 types of extreme weather scenarios.

[0082] 4. Application scenarios: It can be applied to unmanned inspection of high-voltage transmission lines, perimeter security protection of substations, and forest fire early warning systems.

[0083] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0084] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning, characterized in that, include: Real-time acquisition of multimodal data; the multimodal data is acquired by multimodal sensors, including visible light data, infrared data, lidar data, and sound data; The visible light data and the infrared data are registered to obtain a registered visible light image and a registered infrared image; A dual-branch CNN deep learning model is used to obtain a joint feature map based on the registered visible light image and the registered infrared image; The lidar data is preprocessed to obtain a lidar heat map; The sound spectrum encoding is obtained based on the sound data; Fusion data is obtained based on the joint feature map, the lidar heat map, and the sound spectrum encoding; The fused data is input into the dual-task detection head to obtain the target detection result; The target detection results include: vehicle detection frames and smoke / fire separation masks.

2. The method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning according to claim 1, characterized in that, The process of acquiring multimodal data in real time includes: Real-time acquisition of raw visible light data, raw infrared data, raw lidar data, and raw sound data; The original visible light data, the original infrared data, the original lidar data, and the original sound data are spatiotemporally aligned to obtain the visible light data, the infrared data, the lidar data, and the sound data.

3. The method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning according to claim 1, characterized in that, The visible light data and infrared data are registered to obtain a registered visible light image and a registered infrared image, including: Detect the feature points of the visible light data to obtain visible light feature points; Detect the feature points of the infrared data to obtain infrared feature points; Feature point matching is performed using the visible light feature points and the infrared feature points to obtain the feature point matching result; Based on the feature point matching results, the infrared data is subjected to perspective transformation to obtain a registered infrared image, and the visible light data is used as the registered visible light image.

4. The method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning according to claim 1, characterized in that, The lidar data is preprocessed to obtain a lidar heatmap, including: The lidar data is downsampled by voxelization according to a set grid size to obtain the lidar heat map.

5. The method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning according to claim 1, characterized in that, The power transmission channel abnormal target detection method based on multimodal sensors and deep learning also includes: Construct the initial dual-task detection head; Acquire sample multimodal data and the corresponding sample target results; the sample multimodal data includes sample visible light data, sample infrared data, sample lidar data, and sample sound data; Data enhancement processing is performed on the sample visible light data, the sample infrared data, the sample lidar data, and the sample sound data respectively to obtain enhanced sample visible light data, enhanced sample infrared data, enhanced sample lidar data, and enhanced sample sound data; Using the dual-branch CNN deep learning model, a joint feature map of the samples is obtained based on the enhanced sample visible light data and the enhanced sample infrared data; A sample lidar heat map is obtained based on the enhanced sample lidar data; The sample sound spectrum encoding is obtained based on the enhanced sample sound data; Sample fusion data is obtained based on the sample joint feature map, the sample lidar heat map, and the sample sound spectrum encoding; the sample fusion data is labeled based on the sample target results corresponding to the sample multimodal data to obtain labeled targets; A training dataset is constructed based on the sample fusion data and labeled targets; The initial dual-task detection head is trained using a training dataset until the loss function value reaches the set requirement, thus obtaining the dual-task detection head.

6. The method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning according to claim 1, characterized in that, The dual-branch CNN deep learning model includes a first branch, a second branch, and a feature interaction layer; A dual-branch CNN deep learning model is used to obtain a joint feature map based on the registered visible light image and the registered infrared image, including: Using the first branch, visible light image features are obtained based on the registered visible light image; Using the second branch, infrared image features are obtained based on the registered infrared image; The joint feature map is obtained by using the feature interaction layer based on the visible light image features and the infrared image features.

7. The method for detecting abnormal targets in power transmission channels based on multimodal sensors and deep learning according to claim 1, characterized in that, The dual-task detection head includes a shared backbone network, a feature pyramid network, a first task branch, and a second task branch. The fused data is input into the dual-task detection head to obtain the target detection results, including: Using the shared backbone network, multi-scale feature maps are obtained based on the fused data; Using the aforementioned feature pyramid network, a shared feature map is obtained based on the multi-scale feature map; Using the first task branch, a vehicle detection box is obtained based on the shared feature map; Using the second task branch, a smoke segmentation mask is obtained based on the shared feature map.

8. A power transmission channel abnormal target detection device based on multimodal sensors and deep learning, characterized in that, include: A multimodal sensor unit is used to collect multimodal data, including visible light data, infrared data, lidar data, and sound data. A data preprocessing unit, connected to the multimodal data acquisition unit, is used to register the visible light data and the infrared data to obtain registered visible light images and registered infrared images. The data preprocessing unit is also used to obtain a joint feature map based on the registered visible light images and the registered infrared images using a dual-branch CNN deep learning model. The data preprocessing unit is also used to preprocess the lidar data to obtain a lidar heatmap. The data preprocessing unit is also used to obtain sound spectrum encoding based on the sound data. A data fusion unit, connected to the data preprocessing unit, is used to obtain fused data based on the joint feature map, the lidar heat map, and the sound spectrum encoding. A detection unit, connected to the data fusion unit, is used to obtain target detection results based on the fused data; The target detection results include: vehicle detection frames and smoke / fire separation masks.

9. The power transmission channel abnormal target detection device based on multimodal sensors and deep learning according to claim 8, characterized in that, The multimodal sensor unit includes: a visible light camera, an infrared thermal imager, a lidar, a directional microphone array, and a spatiotemporal alignment module; The visible light camera, the infrared thermal imager, the lidar, and the directional microphone array are all connected to the spatiotemporal alignment module; the spatiotemporal alignment module is connected to the data preprocessing unit. The visible light camera is used to collect raw visible light data; The infrared thermal imager is used to collect raw infrared data; The lidar is used to collect raw lidar data; The directional microphone array is used to collect raw sound data; The spatiotemporal alignment module is used to perform spatiotemporal alignment processing on the original visible light data, the original infrared data, the original lidar data, and the original sound data to obtain the visible light data, the infrared data, the lidar data, and the sound data.

10. The power transmission channel abnormal target detection device based on multimodal sensors and deep learning according to claim 8, characterized in that, The power transmission channel abnormal target detection device based on multimodal sensors and deep learning also includes: a cloud platform; The cloud platform is connected to the detection unit; the cloud platform is used to visualize the detection results.