Power distribution network equipment inspection method and platform based on full-life-cycle multi-modal data fusion
Through the multi-modal data fusion method, the image, temperature and prior data of distribution network equipment are obtained and analyzed, and the problem of difficulty in timely detection of faults in traditional inspection methods is solved, efficient and accurate fault detection of distribution network equipment is achieved, and the safe and stable operation of the power grid is ensured.
Patent Information
- Application Number
- CN202510301659.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-08-01
AI Technical Summary
The manual inspection methods of traditional distribution networks are inefficient and difficult to detect equipment failures in a timely and accurate manner, which affects the safety and stability of the power grid operation.
The multimodal data fusion method based on the whole life cycle is adopted to obtain image data, temperature data and prior data of distribution network equipment. Through convolution processing, feature fusion and quantitative analysis, time, position, visual and temperature characteristics are extracted, equipment operation status is quantified, and fault detection is realized.
It realizes timely and accurate fault detection of distribution network equipment, improves the detection level, and ensures the safe and stable operation of the power grid.
Smart Images

Figure CN120409994A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and in particular, to a distribution network equipment inspection method and platform based on multi-modal data fusion throughout the life cycle. Background Art
[0002] Equipment failures may be caused by reasons such as equipment aging, overloaded operation, and environmental factors. For example, cable aging leads to a decline in insulation performance, and short circuits or leakage accidents may occur; overloaded operation of transformers may cause excessive temperatures, which in turn affect the service life; environmental factors, such as natural disasters like heavy rain and wind disasters, can damage equipment. When equipment fails, the normal operation of the power grid will be affected, and even the power system may collapse. Therefore, it is necessary to regularly inspect and maintain equipment, and replace faulty equipment in a timely manner when equipment failures are detected to ensure the normal operation of the distribution network.
[0003] With the vigorous development of the new round of energy revolution, the form of the power grid has changed greatly, becoming more and more complex, and the requirements for power supply reliability are also getting higher and higher. The traditional manual inspection method for distribution networks is inefficient, difficult to comprehensively cover various types of equipment with complex distributions, and may not be able to detect faulty equipment in a timely manner. Moreover, manual inspections are greatly affected by subjective factors of personnel, prone to missed detections of faults, and the data obtained from manual inspections is limited. Data sorting and analysis are time-consuming and laborious, and cannot accurately reflect the operating status of equipment. These limitations pose extremely high risks to the safe and stable operation of the distribution network. Summary of the Invention
[0004] Embodiments of the present invention provide a distribution network equipment inspection method and platform based on multi-modal data fusion throughout the life cycle to solve the problem of being unable to detect distribution network equipment failures in a timely and accurate manner.
[0005] In a first aspect, embodiments of the present invention provide a distribution network equipment inspection method based on multi-modal data fusion throughout the life cycle, including:
[0006] Obtaining image data, temperature data, and prior data corresponding to the distribution network equipment; wherein, the temperature data includes the real-time voltage and real-time operating temperature of each component in the distribution network equipment, and the prior data is obtained based on the inspection reports and maintenance specifications of the distribution network equipment;
[0007] According to the image data, extracting time feature data, position feature data, and visual feature data of the distribution network equipment; wherein, the image data is marked with shooting time information and equipment position information of the distribution network equipment;
[0008] Quantitatively analyze the operating state of the distribution network equipment by using the time feature data, the location feature data, the visual feature data, the temperature data, and the prior data to obtain corresponding quantitative values;
[0009] Determine whether the distribution network equipment has a fault according to the quantitative value.
[0010] In a possible implementation manner, the image data includes visible light images and infrared images;
[0011] Extract the visual feature data of the distribution network equipment according to the image data, including:
[0012] Perform convolution processing on the visible light image to obtain a first background feature map and a first detail feature map corresponding to the visible light image, and perform convolution processing on the infrared image to obtain a second background feature map and a second detail feature map corresponding to the infrared image;
[0013] Perform feature fusion on the first background feature map and the second background feature map to obtain a comprehensive background feature map, and perform feature fusion on the first detail feature map and the second detail feature map to obtain a comprehensive detail feature map;
[0014] Obtain the visual feature data based on the comprehensive background feature map and the comprehensive detail feature map.
[0015] In a possible implementation manner, the obtaining the visual feature data based on the comprehensive background feature map and the comprehensive detail feature map includes:
[0016] Perform image stitching on the comprehensive background feature map and the comprehensive detail feature map to obtain a stitched image, and perform image reconstruction on the stitched image to obtain a fused image;
[0017] Perform wavelet downsampling processing and wavelet upsampling processing on the fused image in sequence to obtain a processed fused image;
[0018] Obtain an attention weight matrix corresponding to the processed fused image, and perform weighted calculation on the processed fused image by using the attention weight matrix to obtain the visual feature data.
[0019] In a possible implementation manner, the obtaining the attention weight matrix corresponding to the processed fused image includes:
[0020] Perform channel compression processing on the processed fused image to obtain corresponding channel attention information, and perform spatial compression processing on the processed fused image to obtain corresponding spatial attention information; wherein, the channel attention information represents the average response characteristics of each channel in the processed fused image over the entire space, and the spatial attention information represents the response characteristics of each spatial position in the processed fused image.
[0021] Superimpose the channel attention vector and the spatial attention vector to obtain the attention weight matrix.
[0022] In a possible implementation, the using the time feature data, position feature data, visual feature data, the temperature data, and the prior data to perform quantitative analysis on the operating state of the distribution network device to obtain corresponding quantitative values includes:
[0023] Obtain the preset time feature data, preset position feature data, preset visual feature data, preset temperature data, and preset prior data when the distribution network device is operating normally.
[0024] Obtain the quantitative value according to the difference between the preset time feature data and the time feature data, the difference between the preset position feature data and the position feature data, the difference between the preset visual feature data and the visual feature data, the difference between the preset temperature data and the temperature data, and the difference between the preset prior data and the prior data.
[0025] In a possible implementation, the obtaining the prior data of the distribution network device includes:
[0026] Obtain the prior text of the distribution network device; wherein, the prior text includes inspection reports and maintenance specifications.
[0027] According to the pre-trained word embedding matrix, perform word vector transformation on each text word in the prior text to determine the word vector corresponding to each text word, and obtain a corresponding vector sequence based on each word vector; wherein, the pre-trained word embedding matrix is trained based on the inspection reports and maintenance specifications of multiple distribution network devices and power system professional knowledge.
[0028] Calculate the importance weights of the word vectors at different positions in the vector sequence, and obtain the prior data according to each word vector and the importance weight corresponding to the word vector.
[0029] In a possible implementation, according to the quantitative value, determining whether the distribution network device has a fault includes:
[0030] If the quantization value is greater than a preset threshold, it is determined that the distribution network device has a fault; otherwise, the distribution network device has no fault.
[0031] In a possible implementation, according to the image data, time feature data and position feature data of the distribution network device are extracted, including:
[0032] The time feature data is obtained by performing convolution processing on the image data in the time dimension, and the position feature data is obtained by performing convolution processing on the image data in the spatial dimension.
[0033] In a second aspect, an embodiment of the present invention provides a distribution network device inspection device based on multi-modal data fusion in the whole life cycle. The device includes:
[0034] An acquisition unit, configured to acquire image data, temperature data, and prior data corresponding to the distribution network device; wherein, the temperature data includes the real-time voltage and real-time operating temperature of each component in the distribution network device, and the prior data is obtained based on the inspection report and maintenance specifications of the distribution network device;
[0035] A first processing unit, configured to extract time feature data, position feature data, and visual feature data of the distribution network device according to the image data; wherein, the image data is marked with shooting time information and device position information of the distribution network device;
[0036] A second processing unit, configured to perform quantitative analysis on the operating state of the distribution network device by using the time feature data, the position feature data, the visual feature data, the temperature data, and the prior data to obtain a corresponding quantization value;
[0037] A discrimination unit, configured to determine whether the distribution network device has a fault according to the quantization value.
[0038] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method in the first aspect or any possible implementation manner of the first aspect is implemented.
[0039] An embodiment of the present invention provides a method and platform for inspecting distribution network equipment based on multi-modal data fusion throughout the life cycle. First, multi-modal data corresponding to the distribution network equipment is obtained, including image data, temperature data, and prior data, and based on the image data among them, time feature data, position feature data, and visual feature data of the distribution network equipment are extracted. Then, the time feature data, position feature data, visual feature data, temperature data, and prior data are used to quantitatively analyze the operating state of the distribution network equipment to obtain corresponding quantitative values. Finally, according to the quantitative values, it is determined whether the distribution network equipment fails. The present invention extracts features by comprehensively considering various aspects of image data, temperature data, and prior data, and quantitatively analyzes the operating state of the distribution network equipment based on the extracted comprehensive feature data, making the analysis results more objective and accurate, and avoiding the limitations of subjective judgment. Moreover, the quantitative values can accurately reflect the operating state of the equipment, can detect subtle changes in the operating state of the equipment in a timely manner, and can be applied to the entire life cycle of the distribution network equipment to realize fault detection and inspection of the distribution network equipment throughout the life cycle. Therefore, through comprehensive data acquisition, in-depth feature extraction, accurate quantitative analysis, and fault judgment, this embodiment can effectively improve the detection level of distribution network equipment, detect distribution network equipment faults in a timely and accurate manner, and ensure the safe and stable operation of the distribution network. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of the implementation of a method for inspecting distribution network equipment based on multi-modal data fusion throughout the life cycle provided by an embodiment of the present invention;
[0041] Figure 2 is a flowchart of the implementation of another method for inspecting distribution network equipment based on multi-modal data fusion throughout the life cycle provided by an embodiment of the present invention;
[0042] Figure 3 is a schematic structural diagram of a device for inspecting distribution network equipment based on multi-modal data fusion throughout the life cycle provided by an embodiment of the present invention;
[0043] Figure 4 is a schematic structural diagram of a platform for inspecting distribution network equipment based on multi-modal data fusion throughout the life cycle provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0045] See Figure 1 , which shows a flowchart of the implementation of a method for inspecting distribution network equipment based on multi-modal data fusion throughout the life cycle provided by an embodiment of the present invention, and is described in detail as follows:
[0046] Step 101: Obtain the image data, temperature data, and prior data corresponding to the distribution network equipment. Among them, the temperature data includes the real-time voltage and real-time operating temperature of each component in the distribution network equipment, and the prior data is obtained based on the inspection reports and maintenance specifications of the distribution network equipment.
[0047] Exemplarily, in this embodiment, the image data, temperature data, and prior data corresponding to the distribution network equipment in real time are obtained, or the image data, temperature data, and prior data corresponding to the distribution network equipment are obtained at preset time intervals.
[0048] Among them, the distribution network equipment may include equipment such as medium-voltage switchgear, ring main unit, low-voltage switchgear, and transformer.
[0049] The image data may include pictures and videos of the distribution network equipment obtained based on imaging devices, and the image data is used to characterize the operating environment of the distribution network equipment, the appearance of the distribution network equipment, the connection relationship between various components, etc.
[0050] The temperature data may include the real-time operating temperature of each component in the distribution network equipment. Since excessive real-time voltage will increase heat generation according to Joule's law, it will also cause an increase in hysteresis eddy current, insulation loss, etc.; if the voltage is too low, the current of equipment such as motors will increase, all of which will cause the temperature of the components to rise. Therefore, the temperature data may also include real-time voltage data.
[0051] The prior data can be obtained by converting the inspection reports and maintenance specification texts of the distribution network equipment. In addition, the prior data may also include data obtained by converting basic corpus texts such as power monographs and relevant professional papers.
[0052] Step 102: According to the image data, extract the time feature data, location feature data, and visual feature data of the distribution network equipment. Among them, the image data is marked with the shooting time information and the equipment location information of the distribution network equipment.
[0053] Among them, the time feature data is used to characterize features such as the change trend of the equipment operating state of the distribution network equipment over time, the time law of fault occurrence, and the timeliness change of equipment performance; the location feature data characterizes features such as the location distribution, topological relationship of the distribution network equipment in the power grid, the relative position with other equipment, and the operating characteristics affected by the geographical location.
[0054] The visual feature data characterizes features such as the appearance form, structural integrity, and surface condition of the distribution network equipment, such as whether there is wear, discoloration, deformation, damage, and the connection state of components.
[0055] Step 103: Use the time feature data, location feature data, visual feature data, temperature data, and prior data to quantitatively analyze the operating state of the distribution network equipment to obtain the corresponding quantitative value.
[0056] In a feasible implementation manner, in this embodiment, the obtained time feature data, location feature data, visual feature data, the obtained temperature data, and the prior data are input into a pre-constructed quantization model to obtain the quantization value of the corresponding operation state of the distribution network equipment output by the quantization model. The model outputs a quantization score of the operation state of the distribution network equipment, such as a failure probability value between 0 and 1. Among them, the quantization model can be trained with various types of feature data of the distribution network equipment as the training data set and the known equipment operation states (normal or faulty) as the labels.
[0057] Among them, corresponding weights can also be assigned according to the influence degree of each feature on the equipment operation state, and the quantization values of each feature are weighted and summed to obtain the comprehensive quantization value of the operation state of the distribution network equipment.
[0058] Step 104: Determine whether the distribution network equipment fails according to the quantization value.
[0059] Exemplarily, in this embodiment, according to the preset discrimination rule, the obtained quantization value is judged to determine the operation state of the distribution network.
[0060] In one example, in this embodiment, when the quantization value is greater than or equal to the first preset threshold, it is determined that the distribution network has failed, and relevant personnel are prompted to perform processing such as equipment repair or replacement; when the quantization value is greater than or equal to the second preset threshold and less than the first preset threshold, it is determined that the operation state of the distribution network may be abnormal, and corresponding early warnings are given to prompt relevant personnel to pay attention; when the quantization value is less than the second preset threshold, it is determined that the operation state of the distribution network is normal. Among them, the first preset threshold is greater than the second preset threshold.
[0061] In summary, the present invention obtains multi-modal data corresponding to the distribution network equipment, including image data, temperature data, and prior data, then comprehensively extracts features from the image data, temperature data, and prior data, and performs quantization analysis on the operation state of the distribution network equipment based on the extracted comprehensive feature data, making the analysis result more objective and accurate, and avoiding the limitations of subjective judgment. Moreover, the quantization value can accurately reflect the operation state of the equipment, and can timely detect the subtle changes in the operation state of the equipment. The present invention can be applied to the entire life cycle of the distribution network equipment to realize fault detection and inspection of the distribution network equipment throughout the life cycle. Therefore, through comprehensive data acquisition, in-depth feature extraction, accurate quantization analysis, and fault judgment, this embodiment can effectively improve the detection level of the distribution network equipment, timely and accurately detect faults of the distribution network equipment, and ensure the safe and stable operation of the distribution network.
[0062] In order to more fully exploit the data value and further improve the accuracy and fineness of the inspection of distribution network equipment, in terms of the processing of prior data and the feature extraction of image data, etc., by performing operations such as converting the prior text into word vectors and calculating the importance weights of word vectors, the key information contained in the prior data can be extracted more accurately; for image data, not only the visible light images and infrared images are distinguished, but also through a series of complex operations such as convolution processing, feature fusion, image reconstruction, and attention mechanism, the background and detail information in the images are deeply mined, so as to obtain more comprehensive and accurate visual feature data. At the same time, the image data is convolved in the time and space dimensions to obtain time feature data and position feature data respectively, further improving the feature data. And by obtaining the preset feature data and preset prior data when the equipment is operating normally, and calculating the quantization value based on the difference between these preset data and the actually obtained data, the quantization analysis method is made more scientific and reasonable, and can more accurately reflect the deviation degree between the operating state and the normal state of the equipment. This series of optimization operations enable the distribution network equipment inspection method based on multi-modal data fusion in the whole life cycle of the present invention to adapt to diverse detection requirements in different implementation manners, effectively improving the performance of the distribution network equipment detection and providing more reliable technical support for ensuring the safe and stable operation of the distribution network.
[0063] See Figure 2 , which shows the implementation flowchart of another distribution network equipment inspection method based on multi-modal data fusion in the whole life cycle provided by the embodiment of the present invention, and is described in detail as follows:
[0064] Step 201, obtain the image data, temperature data, and prior data corresponding to the distribution network equipment; wherein, the temperature data includes the real-time voltage and real-time operating temperature of each component in the distribution network equipment, and the prior data is obtained based on the inspection reports and maintenance specifications of the distribution network equipment.
[0065] In one example, step 201 includes the following steps:
[0066] Obtain the prior text of the distribution network equipment; wherein, the prior text includes inspection reports and maintenance specifications.
[0067] According to the pre-trained word embedding matrix, perform word vector conversion on each text word in the prior text, determine the word vector corresponding to each text word, and obtain the corresponding vector sequence based on each word vector; wherein, the pre-trained word embedding matrix is trained based on the inspection reports and maintenance specifications of multiple distribution network equipment and the professional knowledge of the power system.
[0068] Calculate the importance weights of the word vectors at different positions in the vector sequence, and obtain the prior data according to each word vector and the importance weight corresponding to the word vector.
[0069] For example, this embodiment converts the acquired prior text, such as sentences in the distribution network equipment inspection report, into an initial vector representation through the word embedding layer. Assuming that the vocabulary size is V and the embedding vector dimension is d, for each word w in the text, the embedding matrix E∈R V×d , and get its corresponding embedding vector x w , the embedded vector is input into the Transformer text encoder, and after the above Transformer calculation process, the output vector z at each position is obtained i , and then use the multi-head attention layer in the model to process the output after Transformer encoding. The multi-head attention layer has h attention heads, and the output vector calculated by each attention head is Concatenate these output vectors column by column Then the final word vector y is obtained through linear transformation w =ZW O , combine each word vector to get the prior data y.
[0070] In one example, this embodiment pre-embeds the word matrix E∈R V×d For training, let the window size be n, and for the target word w t , whose context word is w t-n ,...,w t-1 ,w t+1 ,...,w t+n The input of the model is the vector representation of the context word, which is obtained by averaging these vectors. Right now:
[0071]
[0072] Among them, v wi is the context word w i The goal of the model is to predict the target word w t To learn these embedding vectors. When predicting, the Softmax function is used to calculate the probability distribution of the target word:
[0073]
[0074] Among them, u wt is the target word w t The output vector of , V is the vocabulary. During the training process, the embedding matrix and output vector are updated by minimizing the cross entropy loss between the predicted results and the true labels. The cross entropy loss function L is:
[0075]
[0076] Among them, D is all the words in the training dataset. Through iterative training, the word embedding matrix is continuously updated, and finally a trained word embedding matrix is obtained.
[0077] After obtaining the trained word embedding matrix, in this embodiment, the trained word embedding matrix is used to perform word vector transformation on each text word in the obtained prior text of the distribution network equipment, including texts such as inspection reports and maintenance specifications, determine the word vector corresponding to each text word, and obtain a corresponding vector sequence based on each word vector.
[0078] In this embodiment, the vector sequence corresponding to the distribution network equipment is input into the text encoder, where the text encoder is composed of a Transformer text encoder and a multi-head attention layer. Its core lies in calculating the degree of association between each position in the input sequence and other positions.
[0079] Let the input vector sequence be x = [x1, x2,...., xn1], where x i represents the embedding vector of the i-th word, and n1 is the sequence length. For each word i, calculate its attention score e ij The formula for is:
[0080]
[0081] Among them, Q i 、K j are the query vector and key vector of word i respectively, and n1 is the sequence length. The query vector Q i 、key vector K j and value vector V j are obtained through linear transformation of the input vector x i , that is:
[0082] Q i =W Q xi, K j =W K x i , V j =W V x i
[0083] Among them, W Q 、W K 、W V are learnable weight matrices. The attention scores are normalized through the Softmax function to obtain the attention weights α ij :
[0084]
[0085] Finally, the output vector z at position i iis the weighted sum of value vectors
[0086] The multi-head attention mechanism uses multiple attention heads in parallel. Assuming h attention heads are used, each attention head has its own independent W Q 、W K 、W V matrix. The outputs of the h attention heads are concatenated and then passed through a linear transformation to obtain the final output.
[0087] To reduce the dependence on a single image data, temperature data is acquired based on the image data in this embodiment. The device temperature text information T0 obtained by the thermocouple temperature sensor is input into the temperature text encoding module of the model. According to the different component temperatures of each device, a temperature vector T λ is generated, where the temperature text vector T λ is an m1-dimensional vector, and m1 represents the number of components. For each component of the λ-th device, the temperature text vector includes the real-time voltage, real-time operating temperature, etc. of each component. That is, after encoding the obtained above data, T λ is formed, T λ =[T λ1 ,T λ2 ,....,T λm1 .
[0088] Step 202, the image data includes visible light images and infrared images; the visible light images are subjected to convolution processing to obtain a first background feature map and a first detail feature map corresponding to the visible light images, and the infrared images are subjected to convolution processing to obtain a second background feature map and a second detail feature map corresponding to the infrared images.
[0089] Among them, in this embodiment, a high-definition camera is preset to capture visible light images, and an online infrared thermal imager is used to obtain infrared images.
[0090] In a feasible implementation manner, in this embodiment, the visible light image or the infrared image is input into the encoder of the model. Among them, the encoder includes a preset number of convolutional layers, such as three convolutional layers. Convolutional layer 1 and convolutional layer 2 perform specific convolution operations, use their respective convolutional kernels to extract the preliminary features of the input image, and then pass these features to the subsequent layers. Convolutional layer 3 further refines the features and uses an activation function to generate a background feature map and a detail feature map.
[0091] Among them, the background feature map is used to represent the common features of the source image, and the detail feature map is used to represent the unique features of the infrared or visible light image, that is, the stable thermal information features of the infrared image and the rich texture detail features of the visible light image.
[0092] To minimize the difference in background feature maps and maximize the difference in detailed feature maps between visible light images and infrared images, the loss function of the encoder is defined as follows in this embodiment:
[0093] L1 = Φ(||B V - B I ||2 2 ) - α1Φ(||D V - D I ||2 2 )
[0094] Among them, B V , D V are the first background feature map and the first detailed feature map of the visible light image V respectively, B I , D I are the second background feature map and the second detailed feature map of the infrared image, Φ(·) is an activation function used to limit the difference between (-1, 1), and α1 is a constant with a value of 0.5.
[0095] Step 203: Perform feature fusion on the first background feature map and the second background feature map to obtain a comprehensive background feature map, and perform feature fusion on the first detailed feature map and the second detailed feature map to obtain a comprehensive detailed feature map.
[0096] Exemplarily, after feature extraction in this embodiment, a fusion layer is connected after the encoder, and a specific fusion strategy is used to fuse the background and detailed feature maps of infrared and visible light images. The fusion strategy can be a summation strategy or a weighted average strategy.
[0097] Among them, the summation method is to use element addition when fusing the background and detailed feature maps, and the expression is:
[0098] B F = B I ⊕ B V , D F = D I ⊕ D V
[0099] Among them. B F and D F respectively represent the fused comprehensive background feature map and comprehensive detailed feature map; B I and B V respectively represent the second background feature map of the infrared image and the first background feature map of the visible light image, D I , D V are respectively the second detailed feature map of the infrared image and the first detailed feature map of the visible light image, and ⊕ represents direct addition of elements at corresponding positions.
[0100] The weighted average method realizes fusion by assigning weights to the feature maps of different images and then performing element-wise addition. The expression is as follows:
[0101] B F = γ1B I ⊕ γ2B V , D F = γ3D I ⊕ γ4D V
[0102] where γ1 + γ2 = γ3 + γ4 = 1, and by default γ i (i = 1,..., 4) are all equal to 0.5. Here, γ1 and γ2 are the weighting coefficients of the background feature maps corresponding to the infrared image and the visible light image respectively, and γ3 and γ4 are the weighting coefficients of the detail feature maps corresponding to the infrared image and the visible light image respectively. ⊕ still represents element-wise addition at the corresponding positions.
[0103] Step 204: Obtain visual feature data based on the comprehensive background feature map and the comprehensive detail feature map.
[0104] In a feasible implementation, Step 204 includes the following steps:
[0105] Perform image stitching on the comprehensive background feature map and the comprehensive detail feature map to obtain a stitched image, and perform image reconstruction on the stitched image to obtain a fused image; sequentially perform wavelet downsampling processing and wavelet upsampling processing on the fused image to obtain a processed fused image; obtain the attention weight matrix corresponding to the processed fused image, and use the attention weight matrix to perform weighted calculation on the processed fused image to obtain visual feature data.
[0106] In addition, Step 204 further includes:
[0107] Perform channel compression processing on the processed fused image to obtain the corresponding channel attention information, and perform spatial compression processing on the processed fused image to obtain the corresponding spatial attention information; where the channel attention information represents the average response feature of each channel in the entire space of the processed fused image, and the spatial attention information represents the response feature of each spatial position in the processed fused image. Superimpose the channel attention vector and the spatial attention vector to obtain the attention weight matrix.
[0108] Exemplarily, in this embodiment, image stitching is performed on the comprehensive background feature map and the comprehensive detail feature map to obtain a stitched image, and the stitched image is input into a decoder for image reconstruction to obtain a fused image (i.e., I hybrid), and then the fused image is input into the Haar Wavelet Downsampling (HWD) and Haar Wavelet Upsampling (HWU) modules. The HWD module first uses continuous wavelet transform to decompose the fused image into a low-frequency wavelet component yL and a high-frequency wavelet component yH. The high-frequency component contains the detailed information of the image. For a given function (signal) f(t), the wavelet transform formula is as follows:
[0109]
[0110] where a is the scale parameter, which controls the stretching degree of the wavelet function, b is the translation parameter, and by changing b, the signal can be analyzed at different time positions. Ψ(t) is the basic wavelet function, and the formula is as follows:
[0111]
[0112] where this formula is the basic function obtained by stretching and translating the basic wavelet function. The symbol * represents the conjugate complex number. And W f (a, b) reflects the characteristic information of the signal at different scales and positions.
[0113] Then, the high-frequency components in the horizontal (HL1), vertical (LH1), and diagonal (HH1) directions are extracted from yH1. Finally, the extracted high-frequency components and the low-frequency wavelet component yL1 are concatenated along the channel dimension to form a new tensor Γ. The tensor expression is as follows:
[0114] Γ = [yL1, HL1, LH1, HH1]
[0115] HWD effectively retains the detailed information of the image and provides richer data for subsequent feature extraction.
[0116] The HWU module designed based on the wavelet inverse operation is used to upsample the input feature map F, that is, the new tensor obtained by the HWD module. In the HWU module, first, the input feature map is processed by convolution, batch normalization, and activation (Convolution Batch normalization and Swish, CBS). The CBS process consists of a convolutional layer, a batch normalization layer, and an activation function. For the convolutional layer, let the input feature map be F, with a size of H × W × C, and the convolutional kernel be K, with a size of k h × k w × C × C out , and the convolution operation can be expressed as:
[0117]
[0118] Among them, J is the output feature map of the convolutional layer, (i, j) is the spatial position index of the element in the output feature map, c is the output channel index, represents the value at the spatial position (i, j) and channel c in the output feature map; b c is the bias term, m, n are the spatial indices of the convolutional kernel, m is used to index the position of the convolutional kernel in the height direction, n is used to index the position of the convolutional kernel in the width direction, d is the channel index of the input feature map, indicating traversal in the channel, represents the value of the input map F corresponding to the output map position i and channel d. K is the convolutional kernel, m, n are the spatial dimension indices of the convolutional kernel, d is the channel index of the input feature map, and c is the output channel index.
[0119] For the batch normalization layer, batch normalization is performed on the output J of the convolutional layer. Suppose it operates on a dataset of size N. For each channel c, the batch normalization operation is as follows:
[0120]
[0121] Among them, μ c is the mean of channel c, is the variance of channel c, ε is a very small number, defaulting to 10 -3 , used to prevent the denominator from being 0, is the result after normalization and is the final output.
[0122] The ReLU activation function is used as the activation function, and the function expression is as follows:
[0123] f(x) = max(0, x) (1)
[0124] The activation function is applied to the normalization result, and the expression is as follows:
[0125]
[0126] Among them, is the result after being processed by the activation function, that is, the result after CBS processing.
[0127] Then, the input channel is divided into four parts: horizontal (HL2), vertical (LH2), diagonal (HH2), and low-frequency wavelet component yL2 using wavelet transform. Then, the three high-frequency components in different directions are stacked into a high-frequency wavelet component. Finally, the inverse wavelet transform (IWT) is used to reconstruct the high-frequency and low-frequency wavelet components to obtain the output of wavelet upsampling. The inverse wavelet transform formula is as follows:
[0128]
[0129] While restoring the image resolution, the HWU module preserves the details of the image as much as possible.
[0130] In this embodiment, the Graph Channel Attention (GCA) module is used to add attention information to the fused image processed by the HWD module and the HWU module to obtain visual feature data.
[0131] In a feasible implementation manner, first, this embodiment performs channel compression processing on the fused image processed by the HWD module and the HWU module to obtain corresponding channel attention information, and performs spatial compression processing on the processed fused image to obtain corresponding spatial attention information; wherein, the channel attention information represents the average response characteristics of each channel in the processed fused image over the entire space, and the spatial attention information represents the response characteristics of each spatial position in the processed fused image. The channel attention vector and the spatial attention vector are superimposed to obtain an attention weight matrix.
[0132] In an example, this embodiment is based on the Squeeze and Excitation (SE) attention mechanism. First, the fused image processed by the HWD module and the HWU module is divided into g sub-feature groups, and then the Global Squeeze Attention (GSA) module is used to perform feature extraction on each sub-feature group simultaneously to achieve collaborative attention. The GSA module filters the input features in both the spatial and channel dimensions, effectively improving the feature extraction ability for small target defects. When processing the input feature map X with dimensions H×W×C, that is, the fused image processed by the HWD module and the HWU module, channel compression and spatial compression are first performed. Channel compression obtains the channel attention vector through operations such as global average pooling, and spatial compression obtains the spatial attention weight through operations such as 1×1×C convolution. The sum of the outputs of the two is the final output of the GSA module.
[0133] In an example, the GSA module can be divided into the following process. Through global average pooling, the information in the spatial dimensions (H and W) is aggregated to obtain the channel attention vector Z c .
[0134]
[0135] where c = 1, 2,..., C, X ijc represents the element value of the c-th channel at the (i, j) position of the feature map X. Z c is a 1×1×C vector that summarizes the average response characteristics of each channel over the entire space.
[0136] Perform a convolution operation on the input feature map X using a 1×1×C convolution operation to obtain the spatial attention weight Z s .
[0137] Z s = σ(Conv1 × 1 ×c (X)+b s )
[0138] where σ is the Sigmoid activation function, and Conv 1×1×C is the convolution operation. The dimension of Z s is H×W×1, which assigns an attention weight to each spatial position of the feature map, and b s is the bias value.
[0139] Adding Z s and Z c will result in the final output A of the GSA, that is, the attention weight matrix.
[0140] The dimension of A is H×W×C, which fuses the attention information of the channels and the space. Then, multiplying A element-wise with X, the weighted visual feature data X out of the GSA module is obtained:
[0141] X out = X×A.
[0142] Step 205, the image data is marked with shooting time information and the device location information of the distribution network equipment; perform a convolution process on the image data in the time dimension to obtain time feature data, and perform a convolution process on the image data in the spatial dimension to obtain location feature data.
[0143] Exemplarily, after obtaining the image data corresponding to the distribution network equipment in step 201, in this embodiment, the image data is input into a Temporal Convolutional Network (TCN) model, and the TCN model is used to extract the time features of the distribution network equipment included in the image data to obtain time feature data; and the image data is input into a Spatial Convolutional Neural Network (SCNN) model, and the SCNN model is used to extract the spatial location features of the distribution network equipment included in the image data to obtain location feature data.
[0144] In a feasible implementation, the image data includes visible light images and infrared images. The visible light images are marked with the device position information of the distribution network equipment, and the infrared images are marked with the shooting time information of the images. In this embodiment, the visible light images are input into the SCNN model to obtain position feature data, and the infrared images are input into the TCN model to obtain time feature data.
[0145] In one example, when the TCN model processes the time features of the inspection images of the distribution network equipment, it mainly processes the images marked with the shooting time information through the convolutional layer and the activation function, so as to obtain the time feature data. Suppose the input picture data with time marks is I infrared , and its dimension is [N, C, H, W], where N represents the number of images, C is the number of channels, H is the picture height, and W is the data width. Let the convolutional kernel be K1, and its dimension is [C in1 , C out1 , k h1 , k h2 , where C in1 is the number of input channels, C out1 is the number of output channels, k h1 and k h2 are the height and width of the convolutional kernel respectively. The convolution operation can be expressed as
[0146] Y1 = I infrared * K1 + b1
[0147] where * is the convolution operation and b1 is the bias term. The specific calculation method of the convolution operation is:
[0148]
[0149] where i represents the sample index, j represents the output channel index, m and n respectively represent the position indices on the height and width of the output feature map, c represents the channel index, h represents the index of the convolutional kernel in the height direction, and w represents the index of the convolutional kernel in the width direction. Apply the ReLU activation function to Y1, and the expression of the ReLU activation function is as shown in (1).
[0150] Use the activation function for the result to obtain:
[0151] Z1[i, j, m, n] = max(0, Y1[i, j, m, n])
[0152] After the convolution process, the obtained result Z1 is the time feature data related to time extracted by the TCN model.
[0153] Input the visible light image (i.e., Z2) marked with the device position information into the SCNN model. Suppose the input high-definition visible light picture with position marks is I visible, with dimensions [N, C, H, W], where N represents the number of samples, C is the number of channels, H is the height of the image, and W is the width of the data. Let the convolutional kernel be K2, with dimensions [C in2 , C out2 , k h3 , k h4 , where C in2 is the number of input channels, C out2 is the number of output channels, k h3 and k h4 are the height and width of the convolutional kernel respectively. The convolution operation can be expressed as:
[0154] Y2 = I visible * K2 + b2
[0155] where * is the convolution operation and b2 is the bias term. The formula for the convolution operation is shown in (2).
[0156] Apply the Sigmoid activation function to Y1, and the functional expression of Sigmoid is:
[0157]
[0158] Apply the Sigmoid activation function to the result, and the output is:
[0159]
[0160] After convolution processing, the resulting Z2 is the position feature data related to space extracted by the SCNN model.
[0161] Step 206, obtain the preset time feature data, preset position feature data, preset visual feature data, preset temperature data, and preset prior data when the distribution network equipment is operating normally.
[0162] Exemplarily, in this embodiment, image data, temperature data, and prior data when the distribution network equipment is operating normally are pre - obtained, and corresponding feature data, including preset time feature data, preset position feature data, and preset visual feature data, are obtained based on these data.
[0163] Step 207, obtain quantization values according to the differences between the preset time feature data and the time feature data, the preset position feature data and the position feature data, the preset visual feature data and the visual feature data, the preset temperature data and the temperature data, and the preset prior data and the prior data.
[0164] In a feasible implementation, this embodiment compares the real-time characteristic data, temperature data and prior data of the distribution network equipment with the characteristic data, temperature data and prior data of the distribution network equipment during normal operation, determines the gap information between the two, and based on the gap information, quantifies the real-time operating status of the distribution network equipment to obtain a quantified value.
[0165] In an example, the preset temperature data of the distribution network equipment during normal operation is T', the preset time characteristic data is Z1', the preset position characteristic data is Z2', the preset prior data is y', and the preset visual characteristic data is X out '; The real-time temperature data of the distribution network is T, the preset time feature data is Z1, the preset location feature data is Z2, the preset prior data is y, and the preset visual feature data is X out The quantized value is calculated according to the following formula:
[0166] Y=W(|T'-T| 2 +|Z1'-Z1| 2 +|Z2'-Z2| 2 +|X out '-X out | 2 )+b f
[0167] Where W is the weight matrix, b f is the bias value, and random initialization is used. The elements in the weight matrix are usually randomly selected within a small range. During the training process, the process of forward propagation, loss calculation, backpropagation gradient calculation and weight matrix update is repeated. As the training progresses, the weight matrix is gradually adjusted, making the model's prediction results closer and closer to the true label, the loss continues to decrease, and finally the appropriate weight matrix is determined.
[0168] Step 208: If the quantized value is greater than the preset threshold, it is determined that the distribution network equipment has a fault; otherwise, the distribution network equipment has not a fault.
[0169] In a feasible implementation, the T, Z1, Z2, and X out The values are all n×1 dimensional, and the weight matrix W is a 1×n dimensional matrix. After the above operation, the Y value is obtained. If the Y value is greater than the preset threshold, it indicates that there is a fault in the distribution network equipment. The preset threshold value can be 1.
[0170] If there is a fault in the distribution network equipment, Z2 is used to perform the inverse convolution operation, and the formula is as follows:
[0171]
[0172] Among them, the input is y, the convolution kernel is k×k, the default dimension is 1×1, the stride is s, and the padding is p. After the above operations, the original position information, that is, the position of the faulty device, will be obtained.
[0173] In summary, in this embodiment, by performing operations such as word vector transformation on the prior text and calculating the importance weights of word vectors, the key information contained in the prior data can be extracted more accurately, the data value can be mined more fully, and the accuracy and fineness of the distribution network equipment detection can be further improved; for image data, through a series of complex operations such as convolution processing, feature fusion, image reconstruction, and attention mechanism, the background and detail information in the image are deeply mined to obtain more comprehensive and accurate visual feature data. At the same time, the image data is convolved in the time and space dimensions to obtain time feature data and position feature data respectively, further improving the feature data. And by obtaining the preset feature data and preset prior data when the equipment is operating normally, and calculating the quantization value based on the difference between these preset data and the actually obtained data, the quantization analysis method is made more scientific and reasonable, and it can more accurately reflect the deviation degree between the operating state and the normal state of the equipment. This series of optimization operations enable the distribution network equipment inspection method based on multi-modal data fusion in the whole life cycle of the present invention to adapt to diverse detection requirements under different implementation modes, effectively improve the performance of the distribution network equipment detection, and provide more reliable technical support for ensuring the safe and stable operation of the distribution network.
[0174] See Figure 3 , which shows the structural schematic diagram of the distribution network equipment inspection device based on multi-modal data fusion in the whole life cycle provided by the embodiment of the present invention. The distribution network equipment inspection device based on multi-modal data fusion in the whole life cycle includes:
[0175] An acquisition unit 31, configured to acquire image data, temperature data, and prior data corresponding to the distribution network equipment; among them, the temperature data includes the real-time voltage and real-time operating temperature of each component in the distribution network equipment, and the prior data is obtained based on the inspection report and maintenance specifications of the distribution network equipment.
[0176] A first processing unit 32, configured to extract time feature data, position feature data, and visual feature data of the distribution network equipment according to the image data; among them, the image data is marked with shooting time information and equipment position information of the distribution network equipment.
[0177] A second processing unit 33, configured to perform quantization analysis on the operating state of the distribution network equipment by using the time feature data, position feature data, visual feature data, temperature data, and prior data to obtain a corresponding quantization value.
[0178] A discrimination unit 34, configured to determine whether the distribution network equipment fails according to the quantization value.
[0179] In a possible implementation, the image data includes a visible light image and an infrared image; the first processing unit 32 is specifically configured to:
[0180] Perform convolution processing on the visible light image to obtain a first background feature map and a first detail feature map corresponding to the visible light image, and perform convolution processing on the infrared image to obtain a second background feature map and a second detail feature map corresponding to the infrared image.
[0181] Perform feature fusion on the first background feature map and the second background feature map to obtain a comprehensive background feature map, and perform feature fusion on the first detail feature map and the second detail feature map to obtain a comprehensive detail feature map.
[0182] Based on the comprehensive background feature map and the comprehensive detail feature map, obtain visual feature data.
[0183] In a possible implementation, the first processing unit 32 is specifically configured to:
[0184] Perform image stitching on the comprehensive background feature map and the comprehensive detail feature map to obtain a stitched image, and perform image reconstruction on the stitched image to obtain a fused image.
[0185] Perform wavelet downsampling processing and wavelet upsampling processing on the fused image in sequence to obtain a processed fused image.
[0186] Obtain an attention weight matrix corresponding to the processed fused image, and use the attention weight matrix to perform weighted calculation on the processed fused image to obtain visual feature data.
[0187] In a possible implementation, the first processing unit 32 is further specifically configured to:
[0188] Perform channel compression processing on the processed fused image to obtain corresponding channel attention information, and perform spatial compression processing on the processed fused image to obtain corresponding spatial attention information; wherein, the channel attention information represents the average response feature of each channel in the processed fused image over the entire space, and the spatial attention information represents the response feature of each spatial position in the processed fused image.
[0189] Superimpose the channel attention vector and the spatial attention vector to obtain an attention weight matrix.
[0190] In a possible implementation, the second processing unit 33 is specifically configured to:
[0191] Obtain preset time feature data, preset position feature data, preset visual feature data, preset temperature data, and preset prior data when the distribution network equipment is operating normally.
[0192] A quantization value is obtained based on the differences between the preset time feature data and the time feature data, the preset position feature data and the position feature data, the preset visual feature data and the visual feature data, the preset temperature data and the temperature data, and the preset prior data and the prior data.
[0193] In a possible implementation manner, the obtaining unit 31 is specifically configured to:
[0194] Obtain the prior text of the distribution network equipment; wherein, the prior text includes inspection reports and maintenance specifications.
[0195] According to the pre-trained word embedding matrix, perform word vector transformation on each text word in the prior text to determine the word vector corresponding to each text word, and obtain the corresponding vector sequence based on each word vector; wherein, the pre-trained word embedding matrix is trained based on the inspection reports and maintenance specifications of multiple distribution network equipment and power system professional knowledge.
[0196] Calculate the importance weights of the word vectors at different positions in the vector sequence, and obtain the prior data according to each word vector and the importance weight corresponding to the word vector.
[0197] In a possible implementation manner, the discrimination unit 34 is specifically configured to:
[0198] If the quantization value is greater than the preset threshold, it is determined that the distribution network equipment has a fault; otherwise, the distribution network equipment has no fault.
[0199] In a possible implementation manner, the first processing unit 32 is further specifically configured to:
[0200] Perform convolution processing on the image data in the time dimension to obtain time feature data, and perform convolution processing on the image data in the space dimension to obtain position feature data.
[0201] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0202] Figure 4 It is a schematic structural diagram of a distribution network equipment inspection platform based on multi-modal data fusion in the whole life cycle provided by the embodiments of the present invention. As Figure 4As shown, the distribution network equipment inspection platform 4 for multi-modal data fusion based on the full life cycle includes: a processor 40 and a memory 41. The memory 41 stores a computer program 42. When the processor 40 executes the computer program 42, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 40 executes the computer program 42, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0203] Exemplarily, the computer program 42 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 41 and executed by the processor 40 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 42 in the distribution network equipment inspection platform 4 for multi-modal data fusion based on the full life cycle.
[0204] The distribution network equipment inspection platform 4 for multi-modal data fusion based on the full life cycle may include, but is not limited to, a processor 40 and a memory 41. Those skilled in the art can understand that Figure 4 merely examples of the distribution network equipment inspection platform 4 for multi-modal data fusion based on the full life cycle, and do not constitute a limitation on the distribution network equipment inspection platform 4 for multi-modal data fusion based on the full life cycle. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the distribution network equipment inspection platform 4 for multi-modal data fusion based on the full life cycle may further include input / output devices, network access devices, buses, etc.
[0205] The processor 40 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0206] The memory 41 can be an internal storage unit of the power distribution network equipment inspection platform 4 based on full-life-cycle multi-modal data fusion, such as the hard disk or memory of the power distribution network equipment inspection platform 4 based on full-life-cycle multi-modal data fusion. The memory 41 can also be an external storage device of the power distribution network equipment inspection platform 4 based on full-life-cycle multi-modal data fusion, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the power distribution network equipment inspection platform 4 based on full-life-cycle multi-modal data fusion. Further, the memory 41 can also include both the internal storage unit of the power distribution network equipment inspection platform 4 based on full-life-cycle multi-modal data fusion and the external storage device. The memory 41 is used to store the computer program 42 and other programs and data required by the power distribution network equipment inspection platform 4 based on full-life-cycle multi-modal data fusion. The memory 41 can also be used to temporarily store the data that has been output or will be output.
[0207] For the convenience and simplicity of description, only the above division of each functional module / unit is used as an example for illustration. In practical applications, the above functions can be assigned to different functional modules / units according to needs. The above modules / units can be implemented in the form of hardware, or in the form of software, or in the form of a combination of hardware and software.
[0208] The embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented.
[0209] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented.
[0210] Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0211] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described in detail or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Without special instructions and logical conflicts, the terms and / or descriptions among different embodiments are consistent and can be mutually referred to. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0212] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the respective embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A power distribution network equipment inspection method based on multi-modal data fusion in the whole life cycle, characterized in that, The method includes: Obtaining image data, temperature data, and prior data corresponding to the distribution network equipment; wherein, the temperature data includes the real-time voltage and real-time operating temperature of each component in the distribution network equipment, and the prior data is obtained based on the inspection reports and maintenance specifications of the distribution network equipment; Extracting time feature data, location feature data, and visual feature data of the distribution network equipment according to the image data; wherein, the image data is marked with shooting time information and equipment location information of the distribution network equipment; Using the time feature data, the location feature data, the visual feature data, the temperature data, and the prior data to perform quantitative analysis on the operating state of the distribution network equipment to obtain corresponding quantitative values; Determining whether the distribution network equipment fails according to the quantitative values.
2. The inspection method for distribution network equipment based on multi-modal data fusion in the whole life cycle according to claim 1, wherein The image data includes visible light images and infrared images; According to the image data, extracting the visual feature data of the distribution network equipment includes: Performing convolution processing on the visible light image to obtain a first background feature map and a first detail feature map corresponding to the visible light image, and performing convolution processing on the infrared image to obtain a second background feature map and a second detail feature map corresponding to the infrared image; Performing feature fusion on the first background feature map and the second background feature map to obtain a comprehensive background feature map, and performing feature fusion on the first detail feature map and the second detail feature map to obtain a comprehensive detail feature map; Based on the comprehensive background feature map and the comprehensive detail feature map, obtaining the visual feature data.
3. The inspection method for distribution network equipment based on multi-modal data fusion in the whole life cycle according to claim 2, wherein, The obtaining the visual feature data based on the comprehensive background feature map and the comprehensive detail feature map includes: Performing image stitching on the comprehensive background feature map and the comprehensive detail feature map to obtain a stitched image, and performing image reconstruction on the stitched image to obtain a fused image; Successively performing wavelet downsampling processing and wavelet upsampling processing on the fused image to obtain a processed fused image; Obtaining an attention weight matrix corresponding to the processed fused image, and using the attention weight matrix to perform weighted calculation on the processed fused image to obtain the visual feature data.
4. The inspection method for distribution network equipment based on multi-modal data fusion throughout the life cycle according to claim 3, characterized in that The obtaining the attention weight matrix corresponding to the processed fused image includes: Performing channel compression processing on the processed fused image to obtain corresponding channel attention information, and performing spatial compression processing on the processed fused image to obtain corresponding spatial attention information; wherein, the channel attention information represents the average response feature of each channel in the processed fused image over the entire space, and the spatial attention information represents the response feature of each spatial position in the processed fused image; Superimposing the channel attention vector and the spatial attention vector to obtain the attention weight matrix.
5. The inspection method for distribution network equipment based on multi-modal data fusion throughout the life cycle according to any one of claims 1-4, characterized in that, The performing quantitative analysis on the operating state of the distribution network equipment using the time feature data, location feature data, visual feature data, the temperature data, and the prior data to obtain corresponding quantitative values includes: Obtain the preset time feature data, preset position feature data, preset visual feature data, preset temperature data, and preset prior data when the distribution network equipment operates normally; Obtain the quantization value according to the difference between the preset time feature data and the time feature data, the difference between the preset position feature data and the position feature data, the difference between the preset visual feature data and the visual feature data, the difference between the preset temperature data and the temperature data, and the difference between the preset prior data and the prior data.
6. The inspection method for distribution network equipment based on multi-modal data fusion throughout the life cycle according to any one of claims 1-4, characterized in that, The obtaining of the prior data of the distribution network equipment includes: Obtain the prior text of the distribution network equipment; wherein, the prior text includes inspection reports and maintenance specifications; According to the pre-trained word embedding matrix, perform word vector transformation on each text word in the prior text to determine the word vector corresponding to each text word, and obtain the corresponding vector sequence based on each word vector; wherein, the pre-trained word embedding matrix is trained based on the inspection reports and maintenance specifications of multiple distribution network equipment and power system professional knowledge; Calculate the importance weights of the word vectors at different positions in the vector sequence, and obtain the prior data according to each word vector and the importance weight corresponding to the word vector.
7. The method for inspecting and patroling distribution network equipment based on multi-modal data fusion in the whole life cycle according to any one of claims 1-4, characterized in that Determine whether the distribution network equipment has a fault according to the quantization value, including: If the quantization value is greater than the preset threshold, determine that the distribution network equipment has a fault; otherwise, the distribution network equipment has no fault.
8. The method for inspecting and maintaining distribution network equipment based on multi-modal data fusion in the whole life cycle according to any one of claims 1-4, characterized in that, Extract the time feature data and position feature data of the distribution network equipment according to the image data, including: Perform convolution processing on the image data in the time dimension to obtain the time feature data, and perform convolution processing on the image data in the spatial dimension to obtain the position feature data.
9. A power distribution network equipment inspection device based on multi-modal data fusion throughout the life cycle, characterized in that, The device includes: An obtaining unit, configured to obtain the image data, temperature data, and prior data corresponding to the distribution network equipment; wherein, the temperature data includes the real-time voltage and real-time operating temperature of each component in the distribution network equipment, and the prior data is obtained based on the inspection reports and maintenance specifications of the distribution network equipment; A first processing unit, configured to extract the time feature data, position feature data, and visual feature data of the distribution network equipment according to the image data; wherein, the image data is marked with shooting time information and equipment position information of the distribution network equipment; A second processing unit, configured to perform quantitative analysis on the operating state of the distribution network equipment by using the time feature data, the position feature data, the visual feature data, the temperature data, and the prior data, and obtain the corresponding quantization value; A discrimination unit, configured to determine whether the distribution network equipment has a fault according to the quantization value.
10. A power distribution network equipment inspection platform based on multi-modal data fusion throughout the life cycle, characterized in that, It includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in any one of claims 1 to 8 is implemented.