A method for detecting illegal behaviors of power workers based on dual-light image fusion

By fusing visible light images and infrared images with dual-light images, and using semantic perception fusion model and violation detection model, the problems of missed detection and insufficient imaging clarity in detecting violations of power workers in the prior art are solved, and the full-day monitoring and accurate detection of the behavior of power workers are achieved.

CN119418278BActive Publication Date: 2025-05-16HUI CHUANGCHUANG (SHANDONG) NEW ENERGY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411558229.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-05-16
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

When detecting violations of power operators, the prior art is prone to missed detection and missed detection based on visible light images, and the imaging clarity of infrared images is poor, making it impossible to effectively capture the details of the working environment.

Method used

Using a method based on dual-light image fusion, combining the imaging characteristics of visible light images and infrared images, image registration is performed through the R2RIFT algorithm, feature extraction and fusion is used to generate dual-light fusion images, and detection is performed using violation detection models.

Benefits of technology

It realizes full-day monitoring of the behavior of power workers, which can not only capture behavior information clearly and completely, but also retain detailed information of the operation scenario, improving the accuracy and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418278B_ABST
    Figure CN119418278B_ABST
Patent Text Reader

Abstract

The present application provides a method for detecting violations of power workers based on dual-light image fusion, including: obtaining visible light images and infrared images of power workers during power operations; using the R2RIFT algorithm to perform image registration on the visible light image and infrared image; using a semantic perception fusion model to extract features and fuse the registered visible light image and infrared image to obtain a dual-light fusion image; using a violation detection model to detect the dual-light fusion image to obtain the violation detection result of the power worker. The method combines the imaging characteristics of visible light images and infrared images, and the obtained dual-light fusion image can comprehensively utilize the unique features of both, achieve complementary advantages and reduce redundant information, and can not only clearly and completely capture the behavior information of power workers, but also retain the detailed information of the operation scene, and realize all-day monitoring of the behavior of power workers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of target detection technology, and in particular, relates to a method for detecting illegal behaviors of power workers based on dual-light image fusion. Background Art

[0002] Smart grid is an important trend in the development of power grid production. The corresponding safety production supervision mode based on computer vision technology has also become a new direction for the development of power grid management. At present, visible light images or infrared images are usually used to detect illegal behaviors of power workers. Among them, visible light images are rich in color information, and detailed information such as texture and edge can be accurately captured, and they are more in line with the visual habits of the human eye. However, the detection of illegal behaviors of power workers based on visible light images is prone to missed detection and false detection. This is because the weather in outdoor operations is changeable and the equipment is complex and numerous, which brings great interference to the behavior images of power workers. Especially in bad weather conditions, the actions of power workers in visible light images are not clear, and the detection effect is not good. Infrared images have the characteristics of strong anti-interference ability and sensitivity to thermal targets. The use of infrared images can capture the illegal actions of power workers more comprehensively and accurately, and improve the target detection effect. However, due to the poor expression ability of infrared images for details and poor imaging clarity, it is impossible to judge the working environment when detecting illegal behaviors of power workers based on infrared images. Summary of the invention

[0003] In view of this, the purpose of this application is to provide a method for detecting illegal behaviors of power workers based on dual-light image fusion, which combines the imaging characteristics of visible light images and infrared images. The obtained dual-light fusion image can comprehensively utilize the unique characteristics of both, achieve complementary advantages while reducing redundant information, and can not only clearly and completely capture the behavioral information of power workers, but also retain the detailed information of the working scene, thereby realizing all-day monitoring of the behavior of power workers.

[0004] The present application provides a method for detecting illegal behavior of power workers based on dual-light image fusion, the method comprising:

[0005] Obtain visible light images and infrared images of power workers during power operations;

[0006] Using the R2RIFT algorithm, the visible light image and the infrared image are registered;

[0007] Using the semantic perception fusion model, feature extraction and fusion are performed on the registered visible light image and infrared image to obtain a dual-light fusion image.

[0008] The dual-light fusion image is detected using a violation detection model to obtain a violation detection result of the power operator.

[0009] Furthermore, the use of the R2RIFT algorithm to perform image registration on the visible light image and the infrared image includes:

[0010] By using phase congruency, PC images of the visible light image and the infrared image are respectively constructed;

[0011] Using the FAST feature point detection algorithm, feature points in the PC images of the visible light image and the infrared image are detected respectively;

[0012] A convolution sequence is obtained by using a method for constructing PC maps of the visible light image and the infrared image, and maximum index maps of the visible light image and the infrared image are respectively constructed based on the convolution sequence;

[0013] For each detected feature point, based on the maximum index map, the feature descriptors of the visible light image and the infrared image are generated respectively by using the BRIEF algorithm; wherein the feature descriptors are binary code strings of fixed length;

[0014] Based on the feature descriptors of the visible light image and the infrared image, the visible light image and the infrared image are matched using a RANSAC algorithm.

[0015] Furthermore, the utilizing phase congruency to construct PC images of the visible light image and the infrared image respectively includes:

[0016] Apply a two-dimensional Gaussian filter to the image I(x, y) and get the response component and

[0017] Based on the response component and The response vector of the orthogonal filter pair is obtained by convolution operation [e no (x,y),o no (x, y)];

[0018] Based on the response vector [e no (x,y),o no (x, y)], and get the amplitude component A no (x, y) and the phase component φ no (x, y);

[0019] Based on the amplitude component A no (x, y) and the phase component φ no(x, y), and obtain the PC image PC(x, y) of the visible light image and the infrared image.

[0020] Furthermore, the semantic perception fusion model is used to extract and fuse the registered visible light image and infrared image to obtain a dual-light fusion image, including:

[0021] Based on the registered visible light image and infrared image, the semantic perception network is used to extract multi-scale visible light feature images and infrared feature images respectively.

[0022] For the visible light feature image and the infrared feature image, a multi-scale visible light feature enhanced image and an infrared feature enhanced image are generated based on a channel attention mechanism;

[0023] For the visible light feature enhanced image and the infrared feature enhanced image, a shallow dual-light fusion image containing detail and structure information is generated based on a channel attention mechanism and a spatial attention mechanism;

[0024] Based on the cross attention mechanism, the deep features of the visible light feature enhanced image and the infrared feature enhanced image are integrated to generate a deep dual-light fusion feature containing context information;

[0025] The deep dual-light fusion feature is injected into the shallow dual-light fusion image to obtain the dual-light fusion image.

[0026] Furthermore, the generation of multi-scale visible light feature enhanced images and infrared feature enhanced images based on the channel attention mechanism includes:

[0027] After splicing the visible light feature image and infrared feature image at the same scale in the channel dimension, the channel attention layer is used to generate the attention weights corresponding to the visible light feature image and the infrared feature image respectively;

[0028] The attention weights corresponding to the visible light feature image and the infrared feature image are respectively applied to the visible light feature image and the infrared feature image at this scale to obtain a multi-scale visible light feature enhanced image and an infrared feature enhanced image.

[0029] Furthermore, the shallow bi-light fusion image containing detail and structure information is generated based on the channel attention mechanism and the spatial attention mechanism, including:

[0030] After connecting the multi-scale visible light feature enhanced images and infrared feature enhanced images in series in the channel dimension, the fusion weights corresponding to the visible light feature enhanced images and infrared feature enhanced images are generated using the parallel channel attention layer and spatial attention layer.

[0031] The shallow dual-light fusion image is generated based on the fusion weights corresponding to the visible light feature enhanced image and the infrared feature enhanced image respectively.

[0032] Furthermore, based on the cross attention mechanism, the deep features of the visible light feature enhanced image and the infrared feature enhanced image are integrated to generate a deep dual-light fusion feature containing context information, including:

[0033] Using a projection function, respectively, convolution and shaping operations are performed on the visible light feature enhanced image and the infrared feature enhanced image in sequence to convert the visible light feature enhanced image and the infrared feature enhanced image into corresponding keys and values;

[0034] The keys and values ​​of the obtained visible light feature enhanced image and infrared feature enhanced image are combined to obtain the deep dual-light fusion feature.

[0035] Furthermore, the violations include: leaning or relying on, not wearing a safety helmet, crossing the safety fence, smoking in the work area, and no one using the elevator during the operation.

[0036] Furthermore, the violation detection model is pre-trained in the following manner:

[0037] Collect multiple groups of visible light images and infrared images simulating violations by power workers to generate dual-light fusion images of multiple violations;

[0038] Marking the violation behavior on the dual-light fusion image of the violation behavior;

[0039] The dual-light fusion image of the violation is taken as input, the corresponding violation annotation is taken as output, and the Yolov10n model is trained to obtain a violation detection model.

[0040] The method for detecting illegal behaviors of power operators based on dual-light images provided in the present application combines the imaging characteristics of visible light images and infrared images. The obtained dual-light fusion images can comprehensively utilize the unique characteristics of both, achieve complementary advantages while reducing redundant information. It can not only clearly and completely capture the behavioral information of power operators, but also retain the detailed information of the working scene, thereby realizing all-day monitoring of the behavior of power operators. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A flow chart of a method for detecting illegal behavior of electric power workers based on dual-light image fusion provided in an embodiment of the present application is shown;

[0042] Figure 2 The visible light image and infrared image provided by the embodiment of the present application when the electric power operator commits illegal behavior are shown;

[0043] Figure 3 The schematic diagram of the FAST feature point detection algorithm provided in the embodiment of the present application is shown;

[0044] Figure 4 A schematic diagram showing matching of a visible light image and an infrared image using the R2RIFT algorithm provided in an embodiment of the present application is shown;

[0045] Figure 5 A schematic diagram of the structure of the YOLOv10n model provided in an embodiment of the present application is shown;

[0046] Figure 6 A comparison chart of target detection effects obtained by using a dual-light fusion image and a violation detection model provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the technical solution more clear, the technical solution is further described in detail below in conjunction with specific implementation methods. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the technical solution.

[0048] Infrared thermal imaging mainly captures the temperature radiation intensity information of the target, while visible light imaging mainly presents the texture and contour information of the target. The information they obtain is naturally complementary. The dual-light image fusion method can extract more complex and comprehensive behavioral characteristics of power workers, while retaining detailed information of the operation scene to achieve accurate identification of the target and effectively assist safety supervisors in supervising the situation at the power operation site.

[0049] In view of this, the present application embodiment proposes a method for detecting illegal behavior of power workers based on dual-light images. For details, please refer to Figure 1 , Figure 1 FIG. 1 is a flow chart of a method for detecting illegal behavior of electric power workers based on dual-light image fusion provided in an embodiment of the present application. Figure 1 As shown: The method includes:

[0050] S101. Obtain visible light images and infrared images of power workers during power operations.

[0051] In this step, in the past, when computer vision technology was used to study the detection of violations by power workers, only visible light images were usually used as input data for violations. However, when visible light images are interfered with by external light, their effective image information will be quickly lost, causing the accuracy of violation detection to drop sharply. The use of visible light and infrared image fusion can achieve the complementarity of the two types of image information, enhance the anti-interference ability, and effectively retain the basic information of the image. Therefore, the embodiment of the present application collects the fused image of the visible light and infrared images when the power workers commit violations as the input data for violations.

[0052] As an example, the present application embodiment uses Hikvision DS-2TD2166T-25 camera and Hikvision DS-2CD3T10D-15 camera to collect visible light images and infrared images of power workers during power operations. Figure 2 , Figure 2 The following are the visible light images and infrared images of the electric power workers in violation of regulations provided by the embodiment of the present application. Figure 2 As shown, there are visible light images and infrared images of a power worker crossing a safety fence, and visible light images and infrared images of a power worker smoking.

[0053] S102: Using the R2RIFT algorithm, perform image registration on the visible light image and the infrared image.

[0054] In specific implementation, the visible light image and the infrared image can be registered by the following steps:

[0055] Step 1021: Utilize phase consistency to construct PC images of the visible light image and infrared image respectively.

[0056] In this step, traditional image registration techniques mainly rely on the intensity value and gradient information of pixels, but the effectiveness of these methods is greatly reduced when faced with images with significant nonlinear radiation distortion (NRD). In contrast, the phase information obtained by Fourier transform or wavelet transform can highlight the key features in the image, and this feature is insensitive to changes in image brightness and contrast, and can effectively identify image features even in uneven lighting environments. Therefore, this application uses phase consistency to construct PC maps for visible light images and infrared images respectively, and the specific methods are as follows:

[0057] Step 201: Apply a two-dimensional Gaussian filter to the image I(x, y) to obtain the response component and

[0058] In this step, the two-dimensional Gaussian filter (2D-LGF) can generally be obtained by Gaussian diffusion in the vertical direction of the Log-Gabor filter (LGF). Therefore, the definition of the 2D-LGF function is as follows:

[0059] (1)

[0061] Where (ρ, θ) represents logarithmic polar coordinates; n and o are the scale and direction of 2D-LGF respectively; (ρ n ,θ no ) is the center frequency of 2D-LGF; σ ρ and σ θ are the bandwidths in ρ and θ respectively.

[0062] Among them, LGF is a frequency domain filter, and its corresponding spatial domain filter can be obtained by inverse Fourier transform, that is, in the spatial domain, 2D-LGF can be expressed as:

[0063] L(x, y, n, o) = L even (x,y,n,o)+iL odd (x, y, n, o); (2)

[0065] In the formula, the real part L even (x, y, n, o) and the imaginary part L odd (x, y, n, o) represent even-symmetric and odd-symmetric Log-Gabor wavelets respectively.

[0066] Furthermore, using the response component and They represent even symmetric filters L with scale n and direction o respectively. even (x, y, n, o) and odd symmetric filter L odd (x, y, n, o).

[0067] Step 202: Based on the response component and The response vector of the orthogonal filter pair is obtained by convolution operation [e no (x,y),o no (x, y)].

[0068] In this step, the response vector [e no (x,y),o no (x,y)]:

[0069]

[0070] Where * is the convolution operator.

[0071] Step 203: Based on the response vector [e no (x,y),o no (x, y)], and obtain the amplitude component A no (x, y) and the phase component φ no (x, y).

[0072] In this step, the amplitude component A is obtained by the following formula (4): no (x, y):

[0073]

[0074] The phase component φ is obtained by the following formula (5): no (x, y):

[0075] φ no (x, y) = arctan(e no (x,y),o no (x, y)). (5)

[0077] Step 204: Based on the amplitude component A no (x, y) and the phase component φ no (x, y), and obtain the PC image PC(x, y) of the visible light image and the infrared image.

[0078] In this step, considering the analysis results of various directions and various directions, based on the amplitude component A no (x, y) and the phase component φ no (x, y), and introduce the noise compensation term T0, and obtain the two-dimensional PC image PC(x, y) through the following formula (6):

[0079]

[0080] Step 1022: Detect feature points in the PC images of the visible light image and the infrared image respectively using the FAST feature point detection algorithm.

[0081] In this step, compared with SIFT and SURF feature point detection algorithms, FAST feature point detection algorithm has the advantages of fast operation speed and small amount of calculation. Figure 3 , Figure 3 FIG. 2 is a schematic diagram of the FAST feature point detection algorithm provided in an embodiment of the present application. Figure 3 As shown in the figure, the principle of the FAST feature point detection algorithm is to select a defined pixel point as the center point P, and determine whether the center point P is a feature point by comparing the grayscale values ​​of the pixels adjacent to the center point P. Specifically:

[0082] 1) With a pixel point P as the center, usually the distance of 3 pixels is selected as the radius to construct a circular area. This circular area will pass through and include 16 pixels. These passed pixels are regarded as the neighboring pixels of the center point P. 2) The 16 neighboring pixels are marked as 1 to 16 in circular order, and a judgment threshold T is set. 3) The absolute difference between the grayscale values ​​of the 16 neighboring pixels and the center point P is calculated. If N absolute differences are greater than or equal to the threshold T, the center point P is judged to be a feature point.

[0083] Here, since this judgment process is time-consuming, in order to improve the judgment efficiency, the absolute difference between the grayscale values ​​of the center point P and the four equally spaced pixel points numbered 1, 5, 9, and 13 (these four points are evenly distributed on the circumference, one every 90 degrees) is first calculated. If at least three of the four absolute differences are greater than or equal to the set judgment threshold T, then the center point P is regarded as a potential feature point. If the four differences do not meet the conditions, the center point P is directly excluded. This acceleration strategy improves the computational efficiency of feature point search. Among them, the judgment formula is as follows:

[0084]

[0085] In the formula, I p is the gray value of the center point P; I x is the gray value of the pixels to be judged around the center point P; T is the judgment threshold.

[0086] As an example, FAST feature point detection is performed on the constructed PC image. The results show that when FAST feature point detection is performed on visible light images and infrared images, the number of feature points obtained is relatively small, while when FAST feature point detection is performed on the constructed PC image, the number of edge corner point features obtained increases, which lays a stable foundation for subsequent feature point matching.

[0087] Step 1023: obtain a convolution sequence using the method of constructing the PC map of the visible light image and the infrared image, and construct the maximum index map of the visible light image and the infrared image respectively based on the convolution sequence.

[0088] In this step, after the feature point detection stage, in order to enhance the distinction between feature points, it is usually necessary to describe the features of these feature points. Traditional feature descriptors mostly rely on the intensity information or gradient distribution of the image to construct feature vectors. However, this type of descriptor is very sensitive to nonlinear radiation distortion (NRD), so it is not applicable when dealing with the registration task of visible light and infrared images. In addition, phase congruency (PC) images are also not suitable for feature description. There are two main reasons: on the one hand, the amount of information in PC images is relatively small, because most of its pixel values ​​are close to zero, which makes the feature description lack sufficient robustness; on the other hand, PC images are highly sensitive to noise, because it mainly focuses on edge information, which often leads to inaccurate feature description. Therefore, the embodiment of the present application introduces a MIM metric to construct a maximum index map, specifically:

[0089] Step 301: All M in direction o n Amplitude of scale A no (x, y) is summed to obtain the total amplitude A(x, y).

[0090] In this step, the total amplitude A(x, y) is obtained by the following formula (8):

[0091]

[0092] Step 302: Among all directions n, the direction with the maximum amplitude A(x, y) is taken as the maximum index map.

[0093] Step 1024: for each detected feature point, based on the maximum index map, use the BRIEF algorithm to generate feature descriptors for the visible light image and the infrared image respectively; wherein the feature descriptors are binary code strings of fixed length.

[0094] In this step, the principle of generating feature descriptors using the BRIEF algorithm based on the maximum index map is to generate a binary code string based on the grayscale value comparison of pixels in the neighborhood of the feature point. Specifically:

[0095] Step 401: For each detected feature point, based on the maximum index map, a square neighborhood window with a side length of a is established with the feature point as the center; wherein a is a preset parameter.

[0096] Step 402: Use a Gaussian kernel with a variance of σ and a convolution kernel with a size of N×N to perform Gaussian filtering on the pixel values ​​in the neighborhood window.

[0097] In this step, the pixel values ​​in the neighborhood window can be weighted averaged through convolution operation to obtain smooth pixel values. Furthermore, Gaussian filtering is performed on the pixels in the neighborhood window mainly to eliminate the influence of noise on feature description.

[0098] Step 403: randomly generate a number of pairs of points (x, y) in the processed neighborhood window, and compare the grayscale values ​​of each pair of points (x, y).

[0099] In this step, the randomly generated point pair (x, y) can fully reflect the grayscale changes within the neighborhood window.

[0100] If the grayscale value of point x is less than the grayscale value of point y, step 404 is executed to mark the corresponding position of the binary code string as 1.

[0101] Otherwise, execute step 405 and mark the corresponding position of the binary code string as 0.

[0102] Step 406: Repeat the above steps 403-405 until a binary code string of a fixed length is generated.

[0103] In this step, the binary code string is the feature descriptor of the feature point generated by using the BRIEF algorithm.

[0104] Step 1025: Based on the feature descriptors of the visible light image and the infrared image, use the RANSAC algorithm to match the visible light image and the infrared image.

[0105] In this step, feature points are randomly selected from the feature points of the visible light image and the infrared image, and the feature descriptors of the feature points are extracted to form data pairs. The data pairs can be correct data, that is, correctly matched feature points, called inliers, or abnormal data, that is, incorrectly matched feature points, called outliers. The purpose of the RANSAC algorithm is to find as many inliers as possible and remove outliers. Specifically:

[0106] Step 501: Select multiple groups of data pairs to form a minimum data set.

[0107] In this step, the selected data pairs are data pairs for which an affine model can be estimated.

[0108] Step 502: Use the minimum data set to calculate the parameters of the affine model.

[0109] Step 503: bring the remaining data pairs into the affine model after setting parameters, and calculate and record the number of inliers in the remaining data pairs.

[0110] Step 504: Compare the number of inliers recorded this time with the number of inliers recorded last time, and take the affine model with more inliers as the optimal affine model.

[0111] Steps 501 to 504 are repeated until the maximum number of iterations is reached or the number of inliers meets the requirement, so as to output the optimal affine model and the corresponding inliers.

[0112] From the above disclosure, it can be known that the RANSAC algorithm finds the inliers in the data pair by iterating the optimal affine model, and the final set of inliers is the registration result of the visible light image and the infrared image.

[0113] As an example, see Figure 4 , Figure 4 FIG. 2 is a schematic diagram of matching a visible light image and an infrared image using the R2RIFT algorithm provided in an embodiment of the present application. Figure 4 As shown, the registration method of the present application has a higher registration accuracy.

[0114] S103: Using a semantic perception fusion model, extract features and fuse the registered visible light image and infrared image to obtain a dual-light fusion image.

[0115] In specific implementation, the dual-light fusion image can be obtained through the following steps:

[0116] Step 1031: Based on the registered visible light image and infrared image, a semantic perception network is used to extract multi-scale visible light feature images and infrared feature images respectively.

[0117] In this step, multiple shallow feature extraction blocks are used to extract multi-scale visible light feature images and infrared feature images respectively; the shallow feature extraction blocks include: 3×3 convolution layer, batch normalization layer and ReLU activation function layer.

[0118] Step 1032: for the visible light feature image and the infrared feature image, generate a multi-scale visible light feature enhanced image and an infrared feature enhanced image based on a channel attention mechanism.

[0119] In specific implementation, the following steps can be used to generate multi-scale visible light feature enhanced images and infrared feature enhanced images:

[0120] Step 601: After splicing the visible light feature image and the infrared feature image at the same scale in the channel dimension, the channel attention layer is used to generate the attention weights corresponding to the visible light feature image and the infrared feature image respectively.

[0121] In this step, considering that the shallow feature extraction block can extract rich details and structural information, the shallow semantic features are integrated based on the channel-space attention mechanism. Specifically, the visible light feature image and infrared feature image at the same scale are spliced ​​in the channel dimension, and then fed into the channel attention layer composed of convolution and pooling operations to generate the attention weights corresponding to the visible light feature image and the infrared feature image respectively.

[0122] Step 602: Apply the attention weights corresponding to the visible light feature image and the infrared feature image respectively to the visible light feature image and the infrared feature image at the scale to obtain a multi-scale visible light feature enhanced image and an infrared feature enhanced image.

[0123] In this step, the attention weights corresponding to the visible light feature image and the infrared feature image are applied to the original features to generate weighted original features Pi, and the weighted original features Pi are added to the original features from another branch to enhance their representation. The generation process of the feature enhanced image can be expressed as follows:

[0124]

[0125]

[0126]

[0127] In the formula, represents the infrared characteristic image, represents the visible light feature image, represents the infrared feature enhanced image, represents the visible light feature enhanced image, represents element-wise summation, Represents element-wise multiplication, Conv n (·) represents n 1*1 convolutional layers, Contact(·) represents the connection operation in the channel dimension, σ(·) and GAP(·) represent the sigmoid function and global average pooling, respectively.

[0128] Step 1033: for the visible light feature enhanced image and the infrared feature enhanced image, generate a shallow dual-light fusion image containing detail and structure information based on a channel attention mechanism and a spatial attention mechanism.

[0129] In specific implementation, a shallow dual-light fusion image containing detail and structure information can be generated through the following steps:

[0130] Step 701: After connecting the multi-scale visible light feature enhanced image and the infrared feature enhanced image in series in the channel dimension, a parallel channel attention layer and a spatial attention layer are used to generate the fusion weights corresponding to the visible light feature enhanced image and the infrared feature enhanced image, respectively.

[0131] In this step, the generation process of fusion weights can be expressed as follows:

[0132]

[0133]

[0134]

[0135] Here, since the visible light feature and the infrared feature are complementary, the generated weight W i Applied to one of the modalities, the fusion weight of the other modality is expressed as W r , then

[0136] W r =1-W i (15)

[0137] Step 702: Generate the shallow dual-light fusion image based on the fusion weights corresponding to the visible light feature enhanced image and the infrared feature enhanced image respectively.

[0138] In this step, the shallow dual-light fusion image is generated using the following formula (16):

[0139]

[0140] Here, the generated shallow bi-light fusion image contains details and structural information.

[0141] Step 1034: Based on the cross-attention mechanism, the deep features of the visible light feature enhanced image and the infrared feature enhanced image are integrated to generate a deep dual-light fusion feature containing context information.

[0142] In specific implementation, deep dual-light fusion features containing contextual information can be generated through the following steps:

[0143] Step 801: Use a projection function to perform convolution and shaping operations on the visible light feature enhanced image and the infrared feature enhanced image in sequence, so as to convert the visible light feature enhanced image and the infrared feature enhanced image into corresponding keys and values.

[0144] In this step, since high-level visual tasks usually require rich contextual information for comprehensive understanding, deep semantic features are integrated based on the cross-attention mechanism. Specifically, a projection function is deployed, which includes convolution and reshaping operations. The convolution and reshaping operation process can be expressed as:

[0145]

[0146]

[0147] In the formula, x∈{ir,vi} represents the mode, represents a key, denotes the value, Conv(·) and R(·) correspond to the convolutional layer with a kernel size of 3×3 and the reshaping operation, respectively.

[0148] The above formulas (17) and (18) are applied to the visible light feature enhanced image and the infrared feature enhanced image to convert them into corresponding keys and values ​​respectively.

[0149] Step 802: Combine the keys and values ​​of the obtained visible light feature enhanced image and infrared feature enhanced image to obtain the deep dual-light fusion feature.

[0150] In this step, the deep dual-light fusion feature is obtained using the following formula (19):

[0151]

[0152] Here, the features of the visible light feature enhanced image and the infrared feature enhanced image are combined together, so that the deep dual-light fusion features have complementary properties in multimodal features.

[0153] Step 1035: inject the deep dual-light fusion feature into the shallow dual-light fusion image to obtain the dual-light fusion image.

[0154] In this step, the deep semantic information of the visible light image and the infrared image is injected into the dual-light shallow fusion image, so that the dual-light fusion image contains both the rich details and structural information carried by the shallow features and the contextual information carried by the deep features, laying the foundation for the next step of violation detection.

[0155] It should be noted that the semantic perception fusion model is pre-trained, and multiple groups of visible light images and infrared images simulating illegal behaviors of power workers are collected. The visible light images and infrared images are used as inputs of the semantic perception fusion model. By setting the loss function, the output dual-light fusion image can fully integrate the complementary information of the visible light image and the infrared image to ensure the visual fidelity of the dual-light fusion image. Specifically, the loss function includes:

[0156] 1) Content loss function

[0157] In order to promote the semantic perception fusion model to integrate more meaningful information and improve visual quality and quantitative indicators, the content loss function shown in the following formula (20) is set:

[0158]

[0159] In the formula, is the strength loss, is the structural similarity loss, For texture loss.

[0160] Among them, the strength loss It is used to measure the difference between the dual-light fusion image and the source image (i.e., the visible light image and the infrared image) at the pixel level. Specifically, the pixel intensity distribution of the infrared and visible light images is integrated through the maximum selection strategy, and then the integral distribution is used to constrain the pixel intensity distribution of the fusion image. It can be expressed as the following formula (21):

[0161]

[0162] Where H and W are the height and width of the image, respectively, ||·||1 represents the l1-norm, and max(·) denotes the element-wise maximum selection.

[0163] Structural Similarity Loss It is used to make the dual-light fusion image retain more brightness and contrast of the source image, and its expression is as follows (22)(23)(24):

[0164]

[0165]

[0166] L ssim =1-(w vi ×SSIM(I f , I vi )+w ir ×SSIM(I f , I ir )); (twenty four)

[0168] In the formula, and Respectively represent the average values ​​of visible light image, infrared image and dual light fusion image, σ vi and σ ir and σ f Respectively represent the standard deviation of visible light image, infrared image and dual light fusion image, σvif and σ irf They represent the covariance between the visible light image and the dual-light fusion image and the covariance between the infrared image and the dual-light fusion image respectively. c1 and c2 represent constants for stable calculation.

[0169] Texture Loss It is used to force the dual-light fusion image to contain finer texture information, and its expression is as follows (25):

[0170]

[0171] In the formula, Represents the Sobel gradient operator, which is used to measure the fine-grained texture information of an image.

[0172] |·| represents the absolute value operation. Here, it is assumed that the optimal texture of the dual-light fusion image is the maximum aggregation of the infrared image and the visible light image texture.

[0173] 2) Semantic loss function

[0174] Use the segmentation network to output the visible light image segmentation results And infrared image segmentation results The loss function is defined as follows:

[0175]

[0176] S104: Detect the dual-light fusion image using a violation detection model to obtain a violation detection result of the power operator.

[0177] In this step, select YOLOv10n as the detection model. The YOLOv10n algorithm architecture consists of three parts: the backbone network (Backbone), the neck network (Neck) and the detection head (Head). Here, please refer to Figure 5 , Figure 5 The figure shows a schematic diagram of the structure of the YOLOv10n model provided in the embodiment of the present application. Figure 5As shown in the figure, the backbone network inherits the C2f and SPPF structures of YOLOv8, and introduces three new structures: SCDown, C2fUIB and PSA. Among them, SCDown uses the series connection of point convolution and deep convolution to achieve spatial and channel separation of downsampling. As the basic building block of YOLOv10, C2fUIB adopts a compact inverted block structure, combined with deep convolution for spatial mixing and point convolution for channel mixing. In order to solve the high computational complexity problem brought by the self-attention mechanism, YOLOv10n proposes an efficient partial self-attention design called position-sensitive spatial attention (PSA) to improve model performance. The neck network retains the FPN and PANet structures to achieve effective feature fusion. The detection head part uses two lightweight detection heads, One-to-one Head and One-to-manyHead, for training.

[0178] Specifically, the violation detection model is pre-trained in the following manner:

[0179] Step 901: Collect multiple groups of visible light images and infrared images of simulated power workers performing illegal behaviors to generate dual-light fusion images of multiple illegal behaviors.

[0180] Step 902: marking the violation on the dual-light fusion image of the violation.

[0181] In this step, the violations include: leaning or relying on, not wearing a safety helmet, crossing the safety fence, smoking in the work area, and no one using the elevator during the operation.

[0182] Step 903: Take the dual-light fusion image of the violation as input, take the corresponding violation annotation as output, and train the Yolov10n model to obtain a violation detection model.

[0183] As an example, see Figure 6 , Figure 6 Shown is a comparison chart of target detection effects obtained by using a dual-light fusion image and a violation detection model provided in an embodiment of the present application. Figure 6 (a) is the behavior detection effect diagram obtained using visible light images. Figure 6 (b) is the behavior detection effect diagram obtained using infrared images. Figure 6 (c) is the behavior detection effect diagram obtained by using dual-light fusion images. Figure 6 The dual-light fusion image obtained by the embodiment of the present application has significantly improved the confidence of the detection result, which proves that the dual-light fusion image can provide richer semantic information for the target detection task, thereby improving the detection performance.

[0184] The above contents are only preferred embodiments of the present invention. For ordinary technicians in this field, many changes can be made in the specific implementation methods and application scopes based on the ideas of the present technical content. As long as these changes do not deviate from the concept of the present invention, they all fall within the scope of protection of this patent.

Claims

1. A method for detecting illegal behavior of power workers based on dual-light image fusion, characterized in that: The method comprises: Obtain visible light images and infrared images of power workers during power operations; Using the R2RIFT algorithm, the visible light image and the infrared image are registered; Using the semantic perception fusion model, feature extraction and fusion are performed on the registered visible light image and infrared image to obtain a dual-light fusion image. Using a violation detection model, the dual-light fusion image is detected to obtain a violation detection result of the power operator; The method of using the R2RIFT algorithm to perform image registration on the visible light image and the infrared image includes: By using phase congruency, PC images of the visible light image and the infrared image are respectively constructed; Using the FAST feature point detection algorithm, feature points in the PC images of the visible light image and the infrared image are detected respectively; A convolution sequence is obtained by using a method for constructing PC maps of the visible light image and the infrared image, and maximum index maps of the visible light image and the infrared image are respectively constructed based on the convolution sequence; For each detected feature point, based on the maximum index map, the feature descriptors of the visible light image and the infrared image are generated respectively by using the BRIEF algorithm; wherein the feature descriptors are binary code strings of fixed length; Based on the feature descriptors of the visible light image and the infrared image, the visible light image and the infrared image are matched using a RANSAC algorithm; The semantic perception fusion model is used to extract features and fuse the registered visible light image and infrared image to obtain a dual-light fusion image, including: Based on the registered visible light image and infrared image, the semantic perception network is used to extract multi-scale visible light feature images and infrared feature images respectively. For the visible light feature image and the infrared feature image, a multi-scale visible light feature enhanced image and an infrared feature enhanced image are generated based on a channel attention mechanism; For the visible light feature enhanced image and the infrared feature enhanced image, a shallow dual-light fusion image containing detail and structure information is generated based on a channel attention mechanism and a spatial attention mechanism; Based on the cross attention mechanism, the deep features of the visible light feature enhanced image and the infrared feature enhanced image are integrated to generate a deep dual-light fusion feature containing context information; The deep dual-light fusion feature is injected into the shallow dual-light fusion image to obtain the dual-light fusion image.

2. The method according to claim 1, characterized in that The method of utilizing phase congruency to construct PC images of the visible light image and the infrared image respectively includes: Apply a two-dimensional Gaussian filter to the image I(x,y) to get the response component and Based on the response component and The response vector of the orthogonal filter pair is obtained by convolution operation [e no (x,y),o no (x,y)]; Based on the response vector [e no (x,y),o no (x,y)], and get the amplitude component A on (x,y) and phase component φ no (x,y); Based on the amplitude component A no (x,y) and phase component φ no (x, y), and obtain the PC image PC(x, y) of the visible light image and the infrared image.

3. The method according to claim 1, characterized in that The method of generating a multi-scale visible light feature enhanced image and an infrared feature enhanced image based on a channel attention mechanism includes: After splicing the visible light feature image and infrared feature image at the same scale in the channel dimension, the channel attention layer is used to generate the attention weights corresponding to the visible light feature image and the infrared feature image respectively; The attention weights corresponding to the visible light feature image and the infrared feature image are respectively applied to the visible light feature image and the infrared feature image at this scale to obtain a multi-scale visible light feature enhanced image and an infrared feature enhanced image.

4. The method according to claim 1, characterized in that The method of generating a shallow dual-light fusion image containing details and structural information based on a channel attention mechanism and a spatial attention mechanism includes: After connecting the multi-scale visible light feature enhanced images and infrared feature enhanced images in series in the channel dimension, the fusion weights corresponding to the visible light feature enhanced images and infrared feature enhanced images are generated using the parallel channel attention layer and spatial attention layer. The shallow dual-light fusion image is generated based on the fusion weights corresponding to the visible light feature enhanced image and the infrared feature enhanced image respectively.

5. The method according to claim 1, characterized in that Based on the cross attention mechanism, the deep features of the visible light feature enhanced image and the infrared feature enhanced image are integrated to generate a deep dual-light fusion feature containing context information, including: Using a projection function, respectively, convolution and shaping operations are performed on the visible light feature enhanced image and the infrared feature enhanced image in sequence to convert the visible light feature enhanced image and the infrared feature enhanced image into corresponding keys and values; The keys and values ​​of the obtained visible light feature enhanced image and infrared feature enhanced image are combined to obtain the deep dual-light fusion feature.

6. The method according to claim 1, characterized in that The violations include: leaning or relying on, not wearing a safety helmet, crossing the safety fence, smoking in the work area, and no one using the ladder during the operation.

7. The method according to claim 1, characterized in that The violation detection model is pre-trained by: Collect multiple groups of visible light images and infrared images simulating violations by power workers to generate dual-light fusion images of multiple violations; Marking the violation behavior on the dual-light fusion image of the violation behavior; The dual-light fusion image of the violation is taken as input, the corresponding violation annotation is taken as output, and the Yolov10n model is trained to obtain a violation detection model.

Citation Information

Patent Citations

  • Drowning person detection method and system based on visible light and thermal imaging data fusion

    CN111986240A

  • Power grid field operation integrated management and control method based on infrared-visible light image fusion

    CN116363748A