A power transmission line detection method based on multi-modal data fusion and related equipment

By performing regional division and graph structure feature fusion of multimodal data on transmission lines, the problem of low accuracy and reliability of transmission line status detection under the traditional manual inspection mode is solved, and more accurate line status monitoring is achieved.

CN120766155BActive Publication Date: 2026-01-23FIBRLINK NETWORKS +4
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511280072.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-01-23
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Traditional manual inspection methods are insufficient to meet the dynamic needs of line perception, resulting in low accuracy and reliability in detecting abnormal conditions of transmission lines.

Method used

The transmission line detection method based on multimodal data fusion divides the transmission line into regions, filters multimodal data of the target region, constructs a graph structure for feature extraction and fusion, and determines the line's operating status.

Benefits of technology

It enables more comprehensive and accurate transmission line status monitoring, allowing for timely detection of potential problems and ensuring the safe and stable operation of the lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766155B_ABST
    Figure CN120766155B_ABST
Patent Text Reader

Abstract

The present disclosure provides a power transmission line detection method based on multi-modal data fusion and related equipment. The method comprises: performing scene segmentation on the spatial region of the power transmission line based on geographic information and remote sensing image data to obtain a plurality of scene regions with spatial continuity; obtaining original multi-modal data of the power transmission line; filtering target multi-modal data of a target region in a target time period from the original multi-modal data based on a preset distance condition; obtaining a graph structure representing the relationship between the target multi-modal data based on the target multi-modal data and the corresponding position information; performing feature extraction on the target multi-modal data to obtain target multi-modal features; performing feature fusion on the target multi-modal features based on the graph structure to obtain global fusion features; and determining the operating state of the target position of the power transmission line at the target time based on the global fusion features. The operating state of the power transmission line can be accurately and reliably determined.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of line detection, and particularly relates to a power transmission line detection method based on multi-modal data fusion and related equipment. BACKGROUND

[0002] With the exponential growth of the power grid scale, the power transmission network has developed into a three-dimensional spatial system spanning complex geographical environments. The traditional manual inspection mode is limited by the spatiotemporal coverage and subjective experience, and it is difficult to meet the dynamic changing line perception demand. The abnormal state of the power transmission line cannot be captured in time due to the low efficiency of manual inspection, resulting in low accuracy and reliability of power transmission line detection. SUMMARY

[0003] The present disclosure proposes a power transmission line detection method based on multi-modal data fusion and related equipment to solve the above technical problems to some extent.

[0004] In a first aspect, the present disclosure provides a power transmission line detection method based on multi-modal data fusion, comprising:

[0005] performing scene segmentation on a spatial region of the power transmission line based on geographic information and remote sensing image data to obtain a plurality of scene regions with spatial continuity;

[0006] obtain original multi-modal data of the power transmission line; wherein the original multi-modal data includes sensor data and unmanned aerial vehicle data with position information of the power transmission line;

[0007] obtain target multi-modal data of a target region in a target time period from the original multi-modal data based on a preset distance condition; wherein the target region includes a first scene region where a target position in the scene region is located and a second scene region adjacent to the first scene region, and the target time period includes a target time;

[0008] obtain a graph structure representing the relationship between the target multi-modal data based on the target multi-modal data and the corresponding position information;

[0009] perform feature extraction on the target multi-modal data to obtain target multi-modal features;

[0010] perform feature fusion on the target multi-modal features based on the graph structure to obtain global fusion features;

[0011] determine the operating state of the target position of the power transmission line at the target time based on the global fusion features.

[0012] In some embodiments, obtaining a graph structure representing the relationship between the target multi-modal data based on the target multi-modal data and the corresponding position information comprises:

[0013] determine a node of the graph structure based on each of the target multi-modal data;

[0014] in response to a spatial distance between the nodes being less than a preset threshold, construct an edge between the nodes, and determine an edge weight of the edge based on a data type of the nodes and the spatial distance;

[0015] traverse the nodes and the edges based on a preset weight rule to update a node weight of the nodes and / or the edge weight of the edges.

[0016] In some embodiments, the edge weight is:

[0017] wherein w ij is an edge weight of a node i and a neighbor node j , is a data type relationship of a node i and a neighbor node j , ; is a spatial distance between a node i and a neighbor node j , and p is a parameter for controlling a weight decay speed.

[0018] In some embodiments, the target multi-modal features are fused based on the graph structure to obtain global fusion features, including:

[0019] determining an attention coefficient of the node based on the node weight, the edge weight, and the target multi-modal features;

[0020] fusing the target multi-modal features of the node based on the attention coefficient to obtain node fusion features;

[0021] optimizing the node fusion features globally to obtain the global fusion features;

[0022] wherein the attention coefficient includes:

[0023] ,

[0024] ,

[0025] ,

[0026] wherein is an attention coefficient of a node i and a neighbor node j, and exp is an exponential function, is a set of neighbor nodes of node i, j and k are serial numbers of neighbor nodes, is an activation function, is an attention vector, and W is a weight matrix, is a node weight of node i and a weighting of the target multi-modal feature, is a node weight of neighbor node j and a weighting of the target multi-modal feature, h k weighted is a node weight of neighbor node k and a weighting of the target multi-modal feature, is a feature concatenation operation, is an edge weight between node i and neighbor node j, w ik is an edge weight between node i and neighbor node k.

[0027] In some embodiments, the feature fusion of the target multi-modal feature of the node based on the attention coefficient obtains a node fusion feature, including:

[0028] ,

[0029] wherein, is a node fusion feature of node i, is a nonlinear activation function.

[0030] In some embodiments, the global optimization of the node fusion feature obtains the global fusion feature, including:

[0031] initializing a label distribution of the node;

[0032] updating the label distribution based on message passing between the node and neighbor nodes until a preset condition is met, to obtain the global fusion feature;

[0033] wherein, updating the label distribution based on message passing between the node and neighbor nodes until a preset condition is met, to obtain the global fusion feature further includes:

[0034] the node and neighbor nodes pass messages, wherein the message , is a label of node i, is a label of neighbor node j, is a transition term of node i and neighbor node j, is a label distribution of node i;

[0035] updating the label distribution based on the message includes:

[0036] wherein, is an updated label distribution of node i, Z is a node item of node i i is a normalization constant;

[0037] obtaining the global fusion feature based on the updated label distribution, comprising:

[0038] wherein, is a global fusion feature of node i.

[0039] In some embodiments, scene segmentation is performed on the spatial region of the power transmission line based on geographic information and remote sensing image data to obtain a plurality of scene regions with spatial continuity, comprising:

[0040] determining a first penalty value based on the vertical distance of two pixel points in the remote sensing image data to the corresponding nearest power transmission line and whether the two pixel points correspond to the same power transmission line section of the nearest power transmission line;

[0041] determining a second penalty value based on the vertical distance of two pixel points in the remote sensing image data to the corresponding nearest tower and whether the two pixel points correspond to the same tower of the nearest tower;

[0042] determining a prior information distance based on the sum of the first penalty value and the second penalty value;

[0043] obtaining a distance measure based on the prior information distance, the color distance of the two pixel points and the spatial distance, wherein the spatial distance is obtained based on the geographic information of the two pixel points;

[0044] determining pixel points belonging to the same scene region based on the distance measure.

[0045] The second aspect of the present disclosure provides a power transmission line detection device based on multi-modal data fusion, comprising:

[0046] a region division module for performing scene segmentation on the spatial region of the power transmission line based on geographic information and remote sensing image data to obtain a plurality of scene regions with spatial continuity;

[0047] a data acquisition module for acquiring original multi-modal data of the power transmission line, wherein the original multi-modal data includes sensor data and unmanned aerial vehicle data with location information of the power transmission line;

[0048] a data filtering module configured to filter target multi-modal data of a target region in a target time period from the original multi-modal data based on a preset distance condition, wherein the target region includes a first scene region in which a target position in the scene region is located and a second scene region adjacent to the first scene region, and the target time period includes a target time point;

[0049] a graph structure module configured to obtain a graph structure representing relationships between the target multi-modal data based on the target multi-modal data and corresponding position information;

[0050] a feature extraction module configured to perform feature extraction on the target multi-modal data to obtain target multi-modal features;

[0051] a feature fusion module configured to perform feature fusion on the target multi-modal features based on the graph structure to obtain global fusion features;

[0052] a state detection module configured to determine an operating state of the target position of the power transmission line at the target time point based on the global fusion features.

[0053] In some embodiments, obtaining the graph structure representing relationships between the target multi-modal data based on the target multi-modal data and corresponding position information includes:

[0054] determining nodes of the graph structure based on each of the target multi-modal data;

[0055] in response to a spatial distance between the nodes being less than a preset threshold, constructing edges between the nodes and determining edge weights of the edges based on data types of the nodes and the spatial distance;

[0056] traversing the nodes and the edges based on a preset weight rule to update node weights of the nodes and / or edge weights of the edges.

[0057] In some embodiments, the edge weight is:

[0058] wherein w ij is an edge weight of a node i and a neighbor node j , is a data type relationship of a node i and a neighbor node j , ; is a spatial distance between a node i and a neighbor node j , and p is a parameter for controlling a weight decay speed.

[0059] In some embodiments, the feature fusion is performed on the target multi-modal features based on the graph structure to obtain global fusion features, including:

[0060] The attention coefficient of the node is determined based on the node weight, the edge weight, and the target multi-modal feature;

[0061] The feature fusion is performed on the target multi-modal feature of the node based on the attention coefficient to obtain a node fusion feature;

[0062] The global fusion feature is obtained by globally optimizing the node fusion feature;

[0063] The attention coefficient includes:

[0064] ,

[0065] ,

[0066] ,

[0067] wherein, is the attention coefficient of the node i and the neighbor node j, exp is an exponential function, is a set of neighbor nodes of the node i, j and k are serial numbers of the neighbor nodes, is an activation function, is an attention vector, W is a weight matrix, is a weighting of the node weight and the target multi-modal feature of the node i, is a weighting of the node weight and the target multi-modal feature of the neighbor node j, h k weighted is a weighting of the node weight and the target multi-modal feature of the neighbor node k, is a feature concatenation operation, is an edge weight between the node i and the neighbor node j, w ik is an edge weight between the node i and the neighbor node k.

[0068] In some embodiments, the feature fusion is performed on the target multi-modal features based on the graph structure to obtain global fusion features, including:

[0069] ,

[0070] wherein, is a node fusion feature of the node i, is a nonlinear activation function.

[0071] In some embodiments, the global fusion feature is obtained by globally optimizing the node fusion feature, including:

[0072] initializing a label distribution of the node;

[0073] updating the label distribution based on message passing between the node and neighbor nodes until a preset condition is met, to obtain the global fusion feature;

[0074] wherein updating the label distribution based on message passing between the node and neighbor nodes until a preset condition is met, to obtain the global fusion feature further comprises:

[0075] the node and the neighbor node pass messages, wherein the message of the node i and the neighbor node j is , is the label of the node i, is the label of the neighbor node j, is the transition term of the node i and the neighbor node j, is the label distribution of the node i;

[0076] updating the label distribution based on the message comprises:

[0077] wherein, is the updated label distribution of the node i, is the node term of the node i, Z i is a normalization constant;

[0078] obtaining the global fusion feature based on the updated label distribution comprises:

[0079] wherein, is the global fusion feature of the node i.

[0080] In some embodiments, scene segmentation is performed on a spatial region of the power transmission line based on geographic information and remote sensing image data, to obtain a plurality of scene regions with spatial continuity, comprising:

[0081] determining a first penalty value based on a vertical distance from two pixel points in the remote sensing image data to the nearest power transmission line corresponding to the two pixel points, and whether the nearest power transmission lines corresponding to the two pixel points belong to the same power transmission line network section;

[0082] determining a second penalty value based on a vertical distance from two pixel points in the remote sensing image data to the nearest tower corresponding to the two pixel points, and whether the nearest towers corresponding to the two pixel points belong to the same tower;

[0083] determining a prior information distance based on a sum of the first penalty value and the second penalty value;

[0084] A distance metric is obtained based on the prior information distance, the color distance between the two pixels, and the spatial distance; wherein, the spatial distance is obtained based on the geographic information of the two pixels.

[0085] Pixels belonging to the same scene region are determined based on the distance metric.

[0086] A third aspect of this disclosure provides an electronic device including one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, the programs including instructions for performing the method according to the first aspect.

[0087] A fourth aspect of this disclosure provides a non-volatile computer-readable storage medium containing a computer program that, when executed by one or more processors, causes the processors to perform the method described in the first aspect.

[0088] A fifth aspect of this disclosure provides a computer program product including computer program instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect.

[0089] As described above, this disclosure provides a transmission line detection method and related equipment based on multimodal data fusion. The method divides the transmission line into regions and filters target multimodal data for the target region within a target time period from the original multimodal data based on preset distance conditions. This data includes sensor data and UAV data with location information. Next, a graph structure representing the relationships between the multimodal data is constructed based on the target multimodal data and its corresponding location information. Then, features are extracted from the multimodal data to obtain multimodal features. These multimodal features are then fused using the graph structure to obtain global fused features. Finally, the operating status of the transmission line target location at the target time is determined based on the global fused features. By fusing multimodal data, the advantages of different data sources can be comprehensively utilized, overcoming the limitations of a single data source, and reflecting the operating status of the transmission line more comprehensively and accurately. Using a graph structure for feature fusion can effectively mine the correlation information between multimodal data, improving the effect of feature fusion. The operating status of the transmission line determined based on the global fused features is more accurate and reliable, helping to promptly detect potential problems in the transmission line and ensuring its safe and stable operation. Attached Figure Description

[0090] In order to more clearly illustrate the technical solutions in the present disclosure or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative effort based on these drawings.

[0091] Figure 1 The schematic diagram of the power transmission line detection method based on multi-modal data fusion of the embodiment of the present disclosure.

[0092] Figure 2 The schematic diagram of the hardware structure of the exemplary electronic device of the embodiment of the present disclosure.

[0093] Figure 3 The flowchart of the power transmission line detection method based on multi-modal data fusion of the embodiment of the present disclosure.

[0094] Figure 4 The schematic diagram of the power transmission line detection method based on multi-modal data fusion of the embodiment of the present disclosure.

[0095] Figure 5 The schematic diagram of the feature extraction network of the sensor data according to the embodiment of the present disclosure.

[0096] Figure 6 The schematic diagram of the feature extraction network of the image data according to the embodiment of the present disclosure.

[0097] Figure 7 The schematic diagram of the feature extraction network of the radar data according to the embodiment of the present disclosure.

[0098] Figure 8 The accuracy curve of the power transmission line detection method based on multi-modal data fusion of the embodiment of the present disclosure.

[0099] Figure 9 The accuracy curve of the power transmission line detection method based on multi-modal data fusion of the embodiment of the present disclosure.

[0100] Figure 10 The schematic diagram of the power transmission line detection device based on multi-modal data fusion of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0101] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below in combination with specific embodiments and with reference to the drawings.

[0102] It should be noted that, unless otherwise defined, technical terms or scientific terms used in the embodiments of the present disclosure shall have their ordinary meanings to those skilled in the art to which the embodiments of the present disclosure belong. The terms "first", "second", and similar terms used in the embodiments of the present disclosure do not denote any order, quantity, or importance, but are used to distinguish different components. The terms "include", "contain", and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", and the like are used only to represent relative positional relationships, and when the absolute positions of the described objects change, the relative positional relationships can also change accordingly.

[0103] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, and the like of personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.

[0104] For example, in response to receiving a user's active request, a prompt message is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server or storage medium, and the like software or hardware that performs the operation of the technical solutions of the present disclosure according to the prompt message.

[0105] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0106] Figure 1 A schematic diagram of a power transmission line detection method architecture based on multi-modal data fusion of an embodiment of the present disclosure is shown. Referring to Figure 1 The power transmission line detection method architecture 100 based on multi-modal data fusion can include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 can be connected through the wired or wireless network 130. The server 110 can be a stand-alone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, security services, CDN, and other basic cloud computing services.

[0107] The terminal 120 can be implemented in hardware or software. For example, when the terminal 120 is implemented in hardware, the terminal 120 can be various electronic devices with a display screen and supporting page display, including but not limited to a smart phone, a tablet computer, an electronic book reader, a laptop computer, a desktop computer, and the like. When the terminal 120 is implemented in software, the terminal 120 can be installed in the above-listed electronic devices; the terminal 120 can be implemented as a plurality of software or software modules (for example, software or software modules for providing a distributed service) or as a single software or software module, which is not specifically limited herein.

[0108] It should be noted that the power transmission line detection method based on multi-modal data fusion provided in the embodiments of the present application can be executed by the terminal 120 or the server 110. It should be understood that, Figure 1 The number of terminals, networks, and servers in the above description is only illustrative and is not intended to limit the same. According to the implementation needs, there can be any number of terminals, networks, and servers.

[0109] Figure 2 A hardware structure schematic diagram of an exemplary electronic device 200 provided by the embodiments of the present disclosure is shown. As shown in Figure 2 The electronic device 200 can include a processor 202, a memory 204, a network module 206, a peripheral interface 208, and a bus 210. The processor 202, the memory 204, the network module 206, and the peripheral interface 208 are communicatively connected to each other within the electronic device 200 through the bus 210.

[0110] The processor 202 can be a central processing unit (CPU), a neural network processor (NPU), a microcontroller unit (MCU), a programmable logic device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or one or more integrated circuits. The processor 202 can be used to perform functions related to the techniques described in the present disclosure. In some embodiments, the processor 202 can also include multiple processors integrated as a single logical component. For example, as shown in Figure 2 The processor 202 can include multiple processors, i.e., a first processor 202a, a second processor 202b, and a third processor 202c.

[0111] The memory 204 can be configured to store data (e.g., instructions, computer code, etc.). As Figure 2As shown, the data stored by the memory 204 can include program instructions (e.g., program instructions for implementing the power line detection method based on multi-modal data fusion of the embodiments of the present disclosure) and data to be processed (e.g., the memory can store configuration files of other modules, etc.). The processor 202 can also access the program instructions and data stored by the memory 204, and execute the program instructions to operate on the data to be processed. The memory 204 can include volatile storage or non-volatile storage. In some embodiments, the memory 204 can include random access memory (RAM), read-only memory (ROM), optical disk, magnetic disk, hard disk, solid state disk (SSD), flash memory, memory stick, etc.

[0112] The network module 206 can be configured to provide communication with other external devices to the electronic device 200 via a network. The network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, near field communication (NFC), etc.), a cellular network, the Internet, or a combination thereof. It can be understood that the type of network is not limited to the specific examples described above. In some embodiments, the network module 206 can include any combination of any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc.

[0113] The peripheral interface 208 can be configured to connect the electronic device 200 with one or more peripheral devices to enable information input and output. For example, the peripheral devices can include input devices such as keyboards, mice, touchpads, touchscreens, microphones, various sensors, etc., and output devices such as displays, speakers, vibrators, indicator lights, etc.

[0114] The bus 210 can be configured to transmit information between various components (e.g., the processor 202, the memory 204, the network module 206, and the peripheral interface 208) of the electronic device 200, such as internal buses (e.g., processor-memory buses), external buses (USB ports, PCI-E buses), etc.

[0115] It should be noted that although the architecture of the electronic device 200 described above only shows the processor 202, the memory 204, the network module 206, the peripheral interface 208, and the bus 210, in the specific implementation process, the architecture of the electronic device 200 can also include other components necessary for normal execution. In addition, those skilled in the art can understand that the architecture of the electronic device 200 described above can also only include components necessary for implementing the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.

[0116] The state perception and detection of the power transmission line mainly uses single-modal data for judgment, such as using only sensor data to construct a monitoring model, or only using images collected by a drone for state perception. However, there are some problems in using single-modal data for state perception: first, using only single-modal data will result in incomplete information dimensions, such as sensors can obtain conductor sag changes but cannot identify surrounding external damage (such as tree falling), and drone images can obtain surrounding obstacle information but lack conductor load parameters, so relying only on single-modal data can easily lead to incomplete or biased state perception, and it is difficult to reflect the real operating state of the power transmission line. Second, various sensors and drone-mounted equipment can fail, which can lead to the collection of incorrect data or data loss, thereby seriously affecting the perception results. In addition, using single-modal data for state perception is difficult to fully capture the complex operating state of the power transmission line.

[0117] Although there are power transmission line state detection technologies based on multi-modal data, there are still difficulties in realizing the fusion of multi-modal data: first, multi-modal data differs in time dimension, sensors sample at a second or minute level, while drones perform hourly inspection, and there is a time-frequency mismatch between the two. Second, multi-modal data differs greatly in structure and has a complex relationship, and using only simple fusion methods cannot reflect the relationship between data and the influence of data on monitoring results, so it is difficult to directly fuse multi-modal data for joint state perception. For example, some methods establish a three-dimensional finite element model of the tower line system, and based on the tower leg stress numerical value, the inclination angle or ice thickness is back calculated to perceive the state of the power transmission line. However, since the established model may deviate from the actual situation, the perception result is inaccurate, and only relying on tower leg stress data makes it difficult to fully capture the complex operating state of the power transmission line. Some methods integrate waveform and time data into a coordinate system to calculate the specific location of the power transmission line fault. However, only relying on waveform data can determine a limited number of fault types, and misjudgment can easily occur. If there is inaccurate data during waveform collection, or the waveform collection equipment fails, fault diagnosis cannot be performed. Some methods rely on image recognition methods to diagnose faults of the power transmission line, but since image recognition is not sensitive to tower tilt and other faults, and the image collection frequency is low, it cannot perform continuous and high-precision fault detection.

[0118] Therefore, how to improve the accuracy and efficiency of power transmission line detection has become a technical problem to be solved.

[0119] In view of this, the embodiments of the present disclosure provide a power transmission line detection method based on multi-modal data fusion and related equipment. The power transmission line is divided into regions, and target multi-modal data of a target region in a target time period is selected from original multi-modal data based on a preset distance condition. These data cover sensor data with location information and unmanned aerial vehicle data. Then, a graph structure representing the relationship between multi-modal data is constructed according to the target multi-modal data and the corresponding location information. Next, multi-modal features are obtained by extracting features from the multi-modal data. The graph structure is used to fuse the multi-modal features to obtain global fusion features. Finally, the running state of the target position of the power transmission line at the target time is determined based on the global fusion features. By fusing multi-modal data, the advantages of different data sources can be utilized comprehensively, and the limitations of a single data source can be overcome, so that the running state of the power transmission line can be more accurately and comprehensively reflected. The use of the graph structure for feature fusion can effectively mine the correlation information between multi-modal data and improve the effect of feature fusion. The power transmission line running state determined based on the global fusion features is more accurate and reliable, which helps to discover potential problems of the power transmission line in time and ensure the safe and stable operation of the power transmission line.

[0120] Referring to Figure 3 , Figure 3 A schematic flowchart of a power transmission line detection method based on multi-modal data fusion according to an embodiment of the present disclosure is shown. The power transmission line detection method based on multi-modal data fusion according to an embodiment of the present disclosure can be deployed on a terminal or a server. Figure 3 In the power transmission line detection method based on multi-modal data fusion 300, the method can further include the following steps.

[0121] In step S310, the spatial region of the power transmission line is scene segmented based on geographic information and remote sensing image data, and a plurality of scene regions with spatial continuity are obtained.

[0122] The power transmission line can be a conductor and its auxiliary facilities in a power system for transmitting electric energy. The geographic information can refer to the basic geographic features of the area where the power transmission line is located, such as the terrain conditions, the locations of surrounding cities and villages, etc. The remote sensing image data is information records about the earth's surface or atmosphere obtained through remote sensing technology. Specifically, sensors can be used to detect the earth's surface from the air or space, obtain electromagnetic wave information reflected or emitted by the earth's surface objects, and convert it into image data, which can reflect the shape, size, color, texture, etc. of the earth's surface objects. For the power transmission line, remote sensing image data can clearly present the direction of the power transmission line, the surrounding terrain, the vegetation coverage, whether there are obstacles (such as buildings, trees, etc.), etc. to provide intuitive and rich data sources for scene segmentation. The spatial area of the power transmission line refers to the range occupied by the power transmission line in geographical space, which not only includes the area where the power transmission line (such as conductors, towers, etc.) is located, but also covers the area within a certain range around the power transmission line that may affect its operation and safety. Scene segmentation can be the process of dividing a scene into several sub-regions with specific meanings or characteristics. In the scene segmentation of the power transmission line, the spatial area of the power transmission line is divided into different parts according to the characteristics reflected by the geographic information and remote sensing image data. These parts can be divided based on different criteria, such as dividing into mountainous areas, plain areas, hilly areas, etc. according to the terrain; dividing into farmland areas, forest areas, residential areas, etc. according to land use types; or dividing into safe areas, potential danger areas, etc. according to the relationship between the power transmission line and surrounding objects. Spatial continuity can refer to the fact that the segmented scene regions are connected to each other in geographical space without interruption. That is, there is no obvious gap or jump between adjacent scene regions, and they together constitute a complete power transmission line spatial area.

[0123] In some embodiments, scene segmentation is performed on the spatial area of the power transmission line based on geographic information and remote sensing image data to obtain a plurality of scene regions with spatial continuity, including:

[0124] Based on the vertical distance from the two pixel points in the remote sensing image data to the corresponding nearest power transmission line and whether the two pixel points correspond to the same power transmission line section of the nearest power transmission line, a first penalty value is determined;

[0125] Based on the vertical distance from the two pixel points in the remote sensing image data to the corresponding nearest tower and whether the two pixel points correspond to the same tower of the nearest tower, a second penalty value is determined;

[0126] The sum of the first penalty value and the second penalty value is used to determine the prior information distance;

[0127] The distance metric is obtained based on the prior information distance, the color distance and the spatial distance of the two pixel points, wherein the spatial distance is obtained based on geographical information of the two pixel points.

[0128] The pixel points belonging to the same scene region are determined based on the distance metric.

[0129] In the remote sensing image data, the power transmission line presents a linear feature with a certain trend and distribution. By calculating the vertical distance of the pixel points to the nearest power transmission line, the closeness of the pixel points to the power transmission line can be measured. If the vertical distance of the pixel points to the nearest power transmission line is smaller, the correlation between the pixel points and the nearest power transmission line is stronger. Further, if the nearest power transmission lines corresponding to the two pixel points belong to the same power transmission line network segment, the first penalty value can be set to 0. If the nearest power transmission lines corresponding to the two pixel points do not belong to the same power transmission line network segment, the first penalty value can be determined based on the vertical distance of the two pixel points to the nearest power transmission line. The closer the pixel points are to the nearest power transmission line, the stronger the relationship between the pixel points and the respective line network segment, and the pixel points should be divided into different network segments. At this time, a larger penalty value is given to reduce the possibility of being divided into the same scene region; otherwise, the penalty value is smaller, and the probability of belonging to the same scene region is increased. For example, two pixel points A and B, the vertical distance of A to the nearest power transmission line L_A is 5 meters, and the vertical distance of B to the nearest power transmission line L_B is 10 meters. The nearest power transmission line L_A and the nearest power transmission line L_B are different line network segments. While two pixel points C and D, the vertical distance of C to the nearest power transmission line L_C is 20 meters, and the vertical distance of D to the nearest power transmission line L_D is 40 meters. The nearest power transmission line L_C and the nearest power transmission line L_D are different line network segments. According to the preset rule, the first penalty value of the pixel points A and B will be larger than the first penalty value of the pixel points C and D.

[0130] The tower is an important supporting structure of the power transmission line, and its distribution and position have an important influence on the spatial area division of the power transmission line. The vertical distance from two pixel points to their respective nearest tower reflects their relative position relationship with the tower. The smaller the vertical distance from the pixel point to the nearest tower, the stronger the correlation between the pixel point and the nearest tower. If the nearest tower corresponding to the two pixel points belongs to the same tower, the second penalty value can be set to 0. If the nearest tower corresponding to the two pixel points does not belong to the same tower, the second penalty value can be determined based on the vertical distance from the two pixel points to the nearest tower. The closer the pixel point is to the nearest tower, the stronger the relationship between them and the respective tower, and they should be divided into different network segments. At this time, a larger penalty value is given to reduce the possibility of being divided into the same scene area; otherwise, the penalty value is smaller, increasing the probability of belonging to the same scene area. For example, two pixel points A' and B', the vertical distance from A' to the nearest tower T_A is 1 meter, and the vertical distance from B' to the nearest tower T_B is 5 meters. The nearest tower T_A and the nearest tower T_B are different towers. And two pixel points C' and D', the vertical distance from C' to the nearest tower T_C is 10 meters, and the vertical distance from D' to the nearest tower T_D is 15 meters. The nearest tower T_C and the nearest tower T_D are different towers. According to the preset rule, the second penalty value of pixel points A' and B' will be greater than that of pixel points C' and D'.

[0131] The sum of the first penalty value and the second penalty value is used to determine the prior information distance. The prior information distance comprehensively considers the position relationship between the pixel points and the power transmission line and the tower. The sum of the first penalty value and the second penalty value is obtained. The greater the value, the greater the difference between the two pixel points in terms of power transmission line network segment and tower association, and the less likely they belong to the same scene area; otherwise, the smaller the prior information distance, the greater the possibility that they belong to the same scene area. For example, if the first penalty value of the pixel points A and B mentioned earlier is 10, and the second penalty value is 8, then the prior information distance of A and B is 18; the first penalty value of the pixel points C and D is 3, and the second penalty value is 2, and the prior information distance is 5. Obviously, the prior information distance of A and B is greater, and the possibility that they belong to the same scene area is lower.

[0132] The distance metric can be obtained based on the prior information distance, color distance and spatial distance of the two pixel points. Among them, the color distance can be different color characteristics of different ground objects in remote sensing image data. The color distance of two pixel points reflects their similarity in color. The smaller the color distance, the closer their colors are, and they may belong to the same ground object type or scene area; the larger the color distance, the less likely they belong to the same scene area. The spatial distance can be calculated according to the geographic information (such as latitude and longitude coordinates) of the two pixel points. A small spatial distance means that they are adjacent in geographic space and are more likely to belong to the same scene area; a large spatial distance is the opposite. Comprehensive distance metric: the prior information distance, color distance and spatial distance are comprehensively calculated to obtain a distance metric value. This value comprehensively considers the positional relationship between the pixel points and the power transmission line and tower, color characteristics and geographic spatial position, and can more comprehensively measure whether two pixel points belong to the same scene area. For example, assuming that pixel points E and F have a prior information distance of 4, a color distance of 2, and a spatial distance of 3, a distance metric value can be obtained for them through a preset weight and a calculation formula.

[0133] In some implementations, pixel points with a distance metric less than a metric threshold value can be determined to be in the same scene area. The metric threshold value is a pre-set standard value for determining whether two pixel points belong to the same scene area. If the distance metric of two pixel points is less than the metric threshold value, it means that the comprehensive difference between them in terms of prior information, color and spatial position is small, and they are determined to be in the same scene area; otherwise, if the distance metric is greater than the metric threshold value, they are considered not to be in the same scene area. For example, if the metric threshold value is set to 8, the distance metric value of the previously mentioned pixel points E and F is 7 after calculation, which is less than the metric threshold value, so E and F are determined to be in the same scene area; while the distance metric value of pixel points A and B is assumed to be 20, which is greater than the metric threshold value, so they are not in the same scene area. In this way, the spatial area of the power transmission line can be scene segmented to obtain multiple scene areas with spatial continuity, providing a basis for subsequent analysis and management of the power transmission line.

[0134] In other implementations, a superpixel segmentation algorithm based on Simple Linear Iterative Clustering (SLIC) can be used to divide the power transmission line corridor into scene units with spatial continuity. Since the power transmission line is widely distributed, the monitoring area is large, and the environment around the line is complex, the environmental conditions in different areas differ greatly, which has a great influence on the line operation state and the monitoring process.

[0135] Specifically, the distance metric defines the pixel and the pixel similarity between two pixels, The smaller the two pixels are more similar, and the greater the probability of being divided into the same region. The present disclosure proposes a calculation method for power transmission line corridor region division, and the specific calculation formula is as follows:

[0136] wherein, and are adjustable weight parameters.

[0137] wherein, is the distance between the RGB color representations of two pixels, and the more similar the colors are the smaller the distance is; is the spatial distance between two pixels; is the distance between two pixels based on the tower coordinates and the line direction prior information, and the calculation method is as follows:

[0138] Firstly, it is determined whether two pixel points belong to the same power transmission line section. The nearest line section of pixel point i is , and the vertical distance of i to the line is ; the nearest line section of pixel point j is , and the vertical distance of j to the line is . If , it is considered that the line sections to which the two pixel points belong are different, and the penalty value is set, that is, the closer the two pixel points are to the line to which they belong, the stronger the correlation between the pixel points and different line sections, and the weaker the correlation between the two pixel points; if , it is considered that the line sections to which the two pixel points belong are the same, and the penalty value is set. Therefore .

[0139] Secondly, it is determined whether two pixel points belong to the same tower. The nearest tower of pixel point i is , and the distance is . The nearest tower of pixel point j is , and the distance is . If , the two pixel points belong to different towers, and the penalty value about the tower is set, that is, the closer the two pixel points are to the tower to which they belong, the stronger the correlation between the pixel points and different towers, and the weaker the correlation between the two pixel points. If , the two pixel points belong to the same tower, and the penalty value is set. Therefore .

[0140] The penalties of the line section difference and the tower difference are added to obtain the final prior information distance:

[0141] .

[0142] By the above-mentioned calculation method of pixel similarity and the SLIC superpixel segmentation algorithm, the transmission line corridor in the monitoring range is divided into multiple regions with spatial continuity. Each region presents its own characteristics due to different geographical elements, environmental conditions and transmission line directions.

[0143] The spatial neighborhood knowledge base is used to describe the formation of a region with obvious characteristics due to the interaction of environmental elements in adjacent spatial units. It can represent the environmental characteristics of the internal space. Therefore, a spatial neighborhood knowledge base with regional characteristics is first constructed to reflect the characteristics of each monitoring region. In order to realize the labeled management of different regions of the transmission line corridor, a pre-defined label system and software analysis method are used to add labels to each region. The labels are divided into several categories such as terrain, vegetation cover, meteorological environment, etc., and each category has multiple different labels corresponding to different regional characteristics. Through the analysis of GIS data, remote sensing image data and other information, the extracted regional characteristics are matched with the pre-defined label system to automatically add corresponding labels to different regions. For example, if the analysis result shows that the average elevation of a certain region is high and the slope is large, the "mountainous area" label is added; if the analysis result shows that the humidity of a certain region is large, the "high humidity" label is added. Through the above method, the transmission line corridor in the monitoring range is divided into different regions, each region has its own characteristics, and corresponding labels are added according to the regional characteristics, and finally the construction of the spatial neighborhood knowledge base is completed.

[0144] In step S320, the original multi-modal data of the transmission line is obtained; wherein the original multi-modal data includes sensor data and unmanned aerial vehicle data with position information of the transmission line.

[0145] Among them, the original multi-modal data refers to data sets from different data acquisition devices or with different characteristics, such as including sensor data with position information and unmanned aerial vehicle data. Sensor data can be obtained by sensors installed on the transmission line or related facilities, which can monitor various physical quantities (such as temperature, humidity, current, voltage, etc.), and transmit these data together with position information to the detection system. Unmanned aerial vehicle data refers to data obtained by unmanned aerial vehicles (unmanned aerial vehicles) flying near or above the transmission line. Unmanned aerial vehicles can carry various sensors and cameras to take pictures, videos or obtain other related data of the transmission line, providing intuitive information about the appearance, structure or operating state of the transmission line.

[0146] Specifically, a variety of sensors are deployed on the power transmission line to continuously collect data, including wind deviation sensors, tower inclination sensors, tension sensors, ice-coating sensors, temperature sensors, galloping sensors, current and voltage sensors, etc. Some types of sensors are only deployed on the tower, such as tower inclination sensors; some types of sensors are only deployed on the power transmission line, such as galloping sensors; and some sensors can be deployed on both the tower and the power transmission line, such as ice-coating sensors. Once the sensors are deployed, their positions remain unchanged, so the position information of the sensors can be obtained directly when they are deployed, and all the collected data also has the position information of the sensors. The unmanned aerial vehicle is equipped with visible light cameras, laser radars and other devices to conduct regular inspection of the entire power transmission line. During the inspection, the unmanned aerial vehicle follows the pre-set flight route to take pictures and scan at fixed points, ensuring that the information of each tower and power transmission line on the inspection route can be obtained. At the same time, the unmanned aerial vehicle is equipped with a GPS module, which can obtain GPS information in real time during the inspection. Therefore, each image and radar data collected by the unmanned aerial vehicle has GPS information of the shooting position. According to the method of the embodiments of the present disclosure, sensor data, images and radar data collected by the unmanned aerial vehicle are used as input. The sensor data contains the position information of the sensors, and the images and radar data collected by the unmanned aerial vehicle contain the position information of the collection time, so that all the input data have accurate position information.

[0147] In step S330, target multi-modal data of a target region in a target time period is selected from the original multi-modal data based on a preset distance condition; wherein the target region includes a first scene region where a target position in the scene region is located and a second scene region adjacent to the first scene region, and the target time period includes a target time.

[0148] The preset distance condition can include a spatial distance between positions corresponding to the original multi-modal data being less than a preset distance. In the process of monitoring the state of the power transmission line, due to the wide distribution of the data acquisition terminal, the large amount of data, and the regional correlation of the operating state of the power transmission line, multi-modal data with strong correlation with the state of the monitoring point should be selected as the data source for the state perception of the power transmission line, so as to more accurately reflect the real state of the line. By combining the methods of adjacent region screening and distance threshold screening, the data source with strong correlation with the state of the monitoring point is screened out. First, based on the spatial topological relationship of each region in the constructed spatial adjacent knowledge base, all regions directly adjacent to the region containing the monitoring point are determined, and the region containing the monitoring point and all directly adjacent regions are taken as candidate regions. Select the data in the candidate region whose acquisition time is within t' time before the current monitoring time t, and calculate the spatial distance between the data and the monitoring point. If the distance is less than the set distance threshold, the data is taken as the monitoring data source; if the distance is greater than the threshold, the data is not considered in the process of monitoring the point. The spatial topological relationship of the region reflects the spatial adjacency between the regions, and the adjacent regions usually have high state correlation. The introduction of the distance threshold further constrains the spatial range of the data source, avoiding the introduction of data far away from the monitoring point, thereby ensuring the spatial locality and correlation of the data source. By comprehensively considering the adjacency relationship of the region and the distance constraint condition of the data source, the data source closely related to the monitoring point in space can be effectively screened out, providing high-quality data support for subsequent monitoring of the state of the power transmission line.

[0149] Reference is made to Figure 4 , Figure 4 A schematic diagram of a power transmission line detection method based on multi-modal data fusion according to an embodiment of the present disclosure is shown. Figure 4 In the method, the data source can be first screened based on the spatial position prior, then the graph structure is constructed based on the position information and the data relationship, then the multi-modal data feature is extracted based on deep learning, and finally the cross-modal feature fusion and perception enhancement judgment are performed based on the graph attention network and the conditional random field.

[0150] The sensor has fixed position information when it is deployed, and the unmanned aerial vehicle device obtains the current position information according to the GPS module during data acquisition, so all the collected data have specific position coordinates. The coordinate information of the data includes latitude, longitude and height information. Since the position of the sensor is fixed and unchanged, the coordinates of the sensor data are denoted as Since the unmanned aerial vehicle continuously moves during the inspection process, the data acquisition position continuously changes, and the coordinates of the unmanned aerial vehicle data are denoted as , that is, the data collected by the unmanned aerial vehicle at the position at time t. It is assumed that the state perception system monitors the position The line operation status at time t can calculate the distance between all data sources and the monitoring point, and the formula is as follows:

[0151] ,

[0152] Wherein, respectively represent the coordinate position of the monitoring point, respectively represent the coordinate position of the data.

[0153] Based on the above formula, the distance between all data and the monitoring point is calculated, and the data of is selected as the data source of state perception. Wherein is the distance threshold parameter. After screening the adjacent data, the data of all collection times in the time is taken , so the final data source of state perception is the data collected in the time adjacent to the detection position, wherein is the time threshold parameter. Because the collection frequency of different sensors is different, and the radar scanning target of the unmanned aerial vehicle changes continuously, the shape of each data is different, and the neural network will be used later to extract the data of different shapes into uniform size features, so as to further perform feature fusion and state perception. Based on the preset spatial distance condition, the sensor data and unmanned aerial vehicle data in the first scene area covering the target position and its adjacent second scene area (i.e. target area) in the specific target time period (i.e. target multi-modal data) are selected from the original multi-modal data, so as to realize the accurate focusing analysis of the power transmission line state in a specific area and time period.

[0154] In step S340, the graph structure representing the relationship between the target multi-modal data is obtained based on the target multi-modal data and the corresponding position information.

[0155] Wherein, the graph structure associates the sensor data and unmanned aerial vehicle data with position information according to their corresponding position information. Each multi-modal data is taken as a node, and each node contains information from different modalities (sensor data, unmanned aerial vehicle data). The edges can be constructed according to the position proximity between nodes. The nodes with closer distance are more likely to have mutual influence, so the edges are connected.

[0156] Traditional data processing methods can only process sensor data or UAV data separately, and it is difficult to discover the potential association between different modal data. The method based on graph structure can consider multiple modal data comprehensively, and through the relationship between nodes and edges, the internal relationship between different positions and different types of data can be mined. For example, the association between the abnormal sensor data at a certain position and the appearance change of the line captured by the nearby UAV is discovered. Through the feature fusion guided by the graph structure, the global fusion feature contains more rich information, which can more accurately represent the running state of the transmission line. Compared with simple feature splicing or addition, the fusion feature based on graph structure has stronger expression ability and discrimination, provides more comprehensive information for the running state detection of the transmission line, reduces the error and uncertainty brought by a single data source, thereby improves the detection accuracy, and has better generalization ability.

[0157] In some embodiments, obtaining a graph structure representing the relationship between the target multi-modal data and the corresponding position information based on the target multi-modal data, comprises:

[0158] determining a node of the graph structure based on each of the target multi-modal data;

[0159] in response to the spatial distance between the nodes being less than a preset threshold, constructing an edge between the nodes, and determining an edge weight of the edge based on the data type of the nodes and the spatial distance;

[0160] traversing the nodes and the edges based on a preset weight rule to update the node weight of the nodes and / or the edge weight of the edges.

[0161] In some embodiments, each data point in the target multi-modal data is mapped to a node of the graph structure, and the node carries corresponding position information and data type. If the spatial distance between two nodes is less than a preset threshold, an edge is constructed between them, and the weight of the edge is calculated according to whether the data types of the nodes are the same and the spatial distance. Finally, the graph structure is iteratively optimized through a preset weight rule to dynamically adjust the node weight and / or the edge weight, thereby constructing a spatio-temporal graph network capable of representing the complex relationship between multi-modal data.

[0162] Specifically, after screening the data sources, the topology graph of the data sources is constructed in combination with the information of the data itself and the regional label information to represent the relationship between the data. The present disclosure proposes a data source topology graph construction method based on spatial proximity knowledge base predefined rules, aiming to more effectively fuse the semantic information in the power transmission line monitoring data. Each node in the topology graph represents a data source, the edge between two nodes represents the association between the two data sources, and the weight of the edge represents the strength of the association between the data sources. Each node contains the following attributes: device ID; data type; collection time; location information, i.e., GPS coordinates; region ID, used to identify the region where the node is located; regional label, such as "mountainous area", "high humidity", etc.; and node weight , which is used to adjust the importance of the node feature in the subsequent graph attention network, and the initial value is set to 1.0.

[0163] First, a set of rules is predefined, which describes the influence of data types, regional labels, etc. on edge weights and node weights through these rules. These rules are stored in a rule base, and each rule contains the following elements: rule generation condition; object of action, i.e., acting on node weight or acting on edge weight; weight adjustment method and weight adjustment amplitude. For example, "the regional label contains 'forest' and the data type is 'can light image', then the node weight is multiplied by 0.7." and "the regional label of data 1 and the regional label of data 2 are the same and the number is greater than 2, then the edge weight is multiplied by 1.1." and so on.

[0164] In constructing the topology graph, first, each data is taken as a node, if the spatial distance between two nodes is less than a threshold value, then an edge is established between the two nodes, i.e., there is a correlation between the two data; if the distance is greater than the threshold value, then there is no edge between the two nodes, i.e., there is no correlation between the two data. The node weight of all nodes is initialized to 1.0, and the initial edge weight is calculated according to the spatial distance and data type of the two nodes.

[0165] In some embodiments, determining the edge weight of the edge based on the data type of the node and the spatial distance includes:

[0166] , wherein, is the data type relationship of node i and node j ; ; is the spatial distance between node i and node j ; ; is a parameter for controlling the weight decay speed.

[0167] wherein, x i , y i , zi Coordinates of node i, x j , y j , z j Coordinates of node j. Then, each rule in the rule library is traversed to check whether a node or an edge satisfies the rule condition. If the condition is satisfied, the weight of the corresponding node or edge is updated according to the rule. Through this rule-based topology graph construction method, semantic information can be flexibly integrated into the topology graph, and the representation ability of the topology graph for data can be improved.

[0168] In step S350, feature extraction is performed on the target multi-modal data to obtain target multi-modal features.

[0169] After the data source screening and the graph structure construction are completed, different neural networks can be used to extract features from the multi-modal data. According to the characteristics of sensor data, unmanned aerial vehicle image data, and radar point cloud data, a special feature extraction network can be designed, and feature standardization and dimension alignment are performed to ensure that the features of different modalities have the same shape and size, thereby providing a consistent feature space for subsequent cross-modal fusion.

[0170] A DNN composed of stacked autoencoders can be used to extract features from sensor data. Since different types of sensors have different sampling frequencies, for example, a dance sensor performs second-level sampling to monitor the dance of a power transmission line in a short time, while a tower tilt sensor performs hour-level sampling to monitor the slow changes of a tower over a long period of time, and different sensor data have different change rules and characteristics, different DNN networks are designed for different sensors to extract features to adapt to different sampling frequencies and data characteristics. Referring to Figure 5 , Figure 5 FIG. 7 shows a schematic diagram of a feature extraction network for sensor data according to an embodiment of the present disclosure. Figure 5 In the network shown in FIG. 7, for a sensor with a sampling frequency of , the DNN network can be set as follows: the neural network is composed of 3 layers of encoders and 3 layers of decoders (the decoder is used for unsupervised pre-training, and only the encoder part is used in the actual feature extraction process), the shape of the input data is , where is the length of the time window (in seconds), is the number of channels of the sensor (for example, the data collected by a vibration sensor has 3 channels, corresponding to vibration data in directions); and the output is a feature vector of dimensions. According to the sampling frequency f, a specific network structure is designed to extract different sensor data into feature vectors of the same size.

[0171] Referring to Figure 6 , Figure 6A schematic diagram of a feature extraction network of image data according to an embodiment of the present disclosure is shown. Due to the advantages of convolutional neural network (CNN) in image processing, CNN is used to extract features from the RGB image data collected by the UAV. The images collected by the UAV are usually of the same size, assuming the image size is , where is the image width, i.e. the number of horizontal pixels; is the image height, i.e. the number of vertical pixels; is the number of channels, which is 3 in an RGB image. A CNN network is designed for images of this size to extract image features. The CNN network consists of 4 convolutional blocks and 1 global average pooling layer, with an input shape of image and an output of a dimensional feature vector. Through the above CNN network, effective features can be extracted from high-resolution images taken by the UAV.

[0172] Referring to Figure 7 , Figure 7 A schematic diagram of a feature extraction network of radar data according to an embodiment of the present disclosure is shown. Since the number of point clouds in radar point cloud data is usually variable, depending on factors such as scanning range, object density, radar resolution and environmental conditions. Dynamic graph convolutional network (DGCNN) can effectively process radar data with different numbers of point clouds, without the need to train the network separately for each input size, while capturing the local geometric structure and global features of the point cloud, so dynamic graph convolutional network is used to extract features from radar point cloud data. Assuming that the point cloud data collected by the radar contains N points, the shape of the input data is N x 3, and the output is a dimensional feature vector. Through the above dynamic graph convolutional network, effective features can be extracted from three-dimensional point cloud data collected by the radar.

[0173] In step S340, the target multi-modal features are fused based on the graph structure to obtain global fusion features.

[0174] Where, after completing the feature extraction of multi-modal data, the feature vectors of sensor data, UAV image data and UAV radar point cloud data are obtained, respectively , and . In order to consider the mutual relationship between data while realizing cross-modal feature fusion, graph attention network (GAT) is used for feature fusion, and conditional random field is further used for optimization, so that the fusion result has global consistency, and finally Softmax is used for state judgment.

[0175] In some embodiments, the target multi-modal features are fused based on the graph structure to obtain global fusion features, comprising:

[0176] determining an attention coefficient of the node based on the node weight, the edge weight and the target multi-modal feature of the node;

[0177] performing feature fusion on the target multi-modal feature of the node based on the attention coefficient to obtain a node fusion feature;

[0178] performing global optimization on the node fusion feature to obtain the global fusion feature.

[0179] Specifically, in the graph attention network, based on the constructed graph structure, each node aggregates the information of its neighbor nodes through the attention mechanism. For node i and its neighbor node j, the attention coefficient includes:

[0180] ,

[0181] ,

[0182] ,

[0183] wherein, is the attention coefficient of node i and neighbor node j, exp is an exponential function, is the neighbor node set of node i, j and k are the serial numbers of the neighbor nodes, is an activation function, is an attention vector, W is a weight matrix, is the weighting of the node weight and the target multi-modal feature of node i, is the weighting of the node weight and the target multi-modal feature of neighbor node j, h k weighted is the weighting of the node weight and the target multi-modal feature of neighbor node k, is a feature concatenation operation, is the edge weight between node i and neighbor node j, w ik is the edge weight between node i and neighbor node k.

[0184] wherein, through the attention coefficient , the updated node fusion feature of node i can be calculated is:

[0185] wherein, is a nonlinear activation function.

[0186] In order to capture higher-order relationships between nodes, multiple layers of GAT are stacked, and the output feature of each layer is taken as the input of the next layer, and finally the fusion feature of each node is obtained .

[0187] After completing the cross-modal feature fusion based on the graph attention network, the fusion features of each node are obtained . In order to further optimize the fusion results and make them have global consistency, a conditional random field (CRF) is introduced, which models the dependency between nodes to ensure the continuity of the fusion features in space and time, and provides more reliable input for subsequent state judgment.

[0188] The core of the conditional random field is to model the joint probability distribution of node labels y through an energy function . The energy function includes two parts: node terms and transition terms. The node term represents the contribution of each node's own features to the label, and the calculation formula is: where is the label probability distribution calculated by the fully connected layer and the Softmax function, and the calculation formula is: where and are learnable parameters, and K is the number of label categories. The transition term represents the dependency between nodes, encouraging spatially adjacent nodes to have similar labels. The calculation formula is: where is the label compatibility function, which is usually defined as:

[0189] , is the feature similarity function, which usually uses a Gaussian kernel:

[0190] where is an adjustable parameter that controls the similarity decay.

[0191] Combining the node term and the transition term, the energy function can be defined as: . In order to minimize the energy function , the mean field approximation is used for inference, and the specific steps are as follows: initialize the label distribution of each node , perform message passing between each node i and its neighbor node j, and the message is combined with the node term and the message passing result. Update the label distribution of node i: where is the normalization constant. Repeat the message passing and label distribution update until convergence or reach the maximum number of iterations. After CRF optimization, the label distribution of each node not only considers its own features, but also incorporates the context information of neighboring nodes. In order to facilitate subsequent state judgment, the global average pooling is performed on all nodes' fusion features to obtain the global fusion feature , and the fusion formula is as follows:

[0192] .

[0193] Finally, the global consistency optimization is performed on the global fusion features to obtain the final state judgment result. As the output of global consistency optimization, it is used for subsequent state judgment.

[0194] In step S350, the running state of the target position of the power transmission line at the target time is determined based on the global fusion features.

[0195] Wherein, after the feature fusion based on the graph attention network and the global consistency optimization based on the conditional random field are completed, the global fusion features are obtained . Next, these information is used to realize accurate judgment of the power transmission line state through a classification task. The state judgment network can include three parts of an input layer, a fully connected layer and an output layer. The input layer is the global feature , and the global feature is further extracted in the fully connected layer: , wherein , . Finally, the Softmax function is used to output the probability distribution of each category of running state: , wherein , , C is the number of fault categories. The running state category with the highest probability can be the running state judgment result of the current time and the current position.

[0196] Referring to Figure 8 , Figure 8 shows the accuracy curve of the power transmission line detection method based on multi-modal data fusion according to the embodiments of the present disclosure. When the node adjacency threshold in the graph structure is 25 meters, the state perception accuracy and the calculation efficiency when the data screening radius based on the spatial position prior is set to 0.1 kilometers to 1.5 kilometers. It can be found that when the node adjacency threshold is 25 meters, the data screening radius is greater than 0.4 kilometers, the power transmission line state perception can be completed with high accuracy; when the data screening radius is less than 0.7 kilometers, the power transmission line state perception calculation can be completed with high efficiency.

[0197] Referring to Figure 9 , Figure 9 shows the accuracy curve of the power transmission line detection method based on multi-modal data fusion according to the embodiments of the present disclosure. When the data screening radius is 0.7 kilometers, the state perception accuracy and the calculation efficiency when the node adjacency threshold in the graph structure is set to 10 meters to 50 meters. It can be found that when the data screening radius is 0.7 kilometers, the node adjacency threshold is set to 20 meters to 30 meters, the power transmission line fault detection can be completed with high accuracy; when the node adjacency threshold is less than 40 meters, the power transmission line state perception calculation can be completed with high efficiency.

[0198] It can be seen that, according to the method of the embodiment of the present disclosure, based on the data position information and the mutual dependence relationship, the model combining the graph attention network and the conditional random field is adopted to dynamically perceive the operation state of the power transmission line, thereby providing strong state perception support for fault detection, and thus the power transmission line fault diagnosis with high accuracy is carried out. The method proposed in the present application can effectively solve the problems of weak perception ability and low accuracy of the method relying only on single modal data in power transmission line state perception, and the heterogeneity problem of multi-source data, and at the same time, in the case of partial sensor data missing, insufficient unmanned aerial vehicle inspection coverage and interference of bad weather, the accurate evaluation of the line state can be completed. The multi-modal data can be screened and fused according to the position information, and finally the state perception enhancement judgment of the power transmission line is completed. In the power transmission line perception enhancement judgment method, a multi-modal data source screening method based on spatial position prior is proposed for the data source problem, a multi-modal data feature extraction method based on multiple deep neural networks is proposed for the heterogeneity problem of different modal data, and a cross-modal feature fusion and perception enhancement judgment method based on graph attention network and conditional random field is proposed for the multi-modal feature fusion problem. Simulation experiments show that the method proposed in the present application can still complete accurate evaluation of the line state in the case of partial sensor data missing, insufficient unmanned aerial vehicle inspection coverage and interference of bad weather, and has good power transmission line state perception effect.

[0199] In summary, the method of the embodiment of the present disclosure realizes the area division of the power transmission line corridor through the pixel point distance measurement calculation method for the power transmission line scene, filters the data with strong correlation with the monitoring points through the multi-modal data source screening method based on the spatial proximity knowledge base, and provides high-quality data sources for the power transmission line state perception. The graph structure construction method based on position information and data relationship reflects the spatial relationship, time relationship and attribute relationship between data through the graph structure. The multi-modal data feature extraction method based on deep learning designs a specific neural network for different modal data to extract features with the same size and shape from multi-modal data, thereby solving the heterogeneity problem of multi-modal data. The multi-modal feature fusion and perception enhancement judgment method based on graph attention network and conditional random field efficiently fuses the features of multi-modal data based on various relationships between data and performs global consistency optimization, thereby providing accurate results for power transmission line state perception and fault judgment. In the case of partial sensor data missing, insufficient unmanned aerial vehicle inspection coverage and interference of bad weather, the accurate evaluation of the line state can still be completed, and the method has good power transmission line state perception effect.

[0200] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server, etc. The method of the embodiments can also be applied to a distributed scenario, and be completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiments of the present disclosure, and the multiple devices can interact with each other to complete the method.

[0201] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still achieve desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0202] Based on the same technical concept, the present disclosure also provides a power transmission line detection device based on multi-modal data fusion corresponding to the method of any of the above embodiments, referring to Figure 10 , the power transmission line detection device based on multi-modal data fusion, the device comprises:

[0203] a region division module for performing scene segmentation on the spatial region of the power transmission line based on geographic information and remote sensing image data to obtain a plurality of scene regions with spatial continuity;

[0204] a data acquisition module for acquiring original multi-modal data of the power transmission line; wherein the original multi-modal data includes sensor data and unmanned aerial vehicle data with location information of the power transmission line;

[0205] a data screening module for screening target multi-modal data of a target region in a target time period from the original multi-modal data based on a preset distance condition; wherein the target region includes a first scene region where a target position in the scene region is located and a second scene region adjacent to the first scene region, and the target time period includes a target time;

[0206] a graph structure module for obtaining a graph structure representing the relationship between the target multi-modal data based on the target multi-modal data and the corresponding location information;

[0207] a feature extraction module for performing feature extraction on the target multi-modal data to obtain target multi-modal features;

[0208] a feature fusion module for performing feature fusion on the target multi-modal features based on the graph structure to obtain global fusion features;

[0209] a state detection module configured to determine an operating state of the target position of the power transmission line at the target time based on the global fusion feature.

[0210] In some embodiments, a graph structure representing relationships between the target multi-modal data is obtained based on the target multi-modal data and corresponding position information, including:

[0211] determining nodes of the graph structure based on each of the target multi-modal data;

[0212] in response to a spatial distance between the nodes being less than a preset threshold, constructing an edge between the nodes, and determining an edge weight of the edge based on a data type of the nodes and the spatial distance;

[0213] traversing the nodes and the edges based on a preset weight rule to update a node weight of the nodes and / or the edge weight of the edges.

[0214] In some embodiments, the determining of the edge weight of the edge based on the data type of the nodes and the spatial distance includes:

[0215] wherein w ij is an edge weight of a node i and a neighbor node j , is a data type relationship of a node i and a neighbor node j , ; is a spatial distance between a node i and a neighbor node j , and p is a parameter for controlling a weight decay speed.

[0216] In some embodiments, the target multi-modal feature is fused based on the graph structure to obtain a global fusion feature, including:

[0217] determining an attention coefficient of the node based on the node weight, the edge weight, and the target multi-modal feature;

[0218] fusing the target multi-modal feature of the node based on the attention coefficient to obtain a node fusion feature;

[0219] optimizing the node fusion feature globally to obtain the global fusion feature;

[0220] wherein the attention coefficient includes:

[0221] ,

[0222] ,

[0223] ,

[0224] wherein, is an attention coefficient of node i and neighbor node j, exp is an exponential function, is a neighbor node set of node i, j and k are serial numbers of neighbor nodes, is an activation function, is an attention vector, W is a weight matrix, is a node weight of node i and a weighting of a target multi-modal feature, is a node weight of neighbor node j and a weighting of a target multi-modal feature, k weighted is a node weight of neighbor node k and a weighting of a target multi-modal feature, is a feature concatenation operation, is an edge weight between node i and neighbor node j, w ik is an edge weight between node i and neighbor node k.

[0225] In some embodiments, the feature fusion of the target multi-modal feature of the node based on the attention coefficient obtains a node fusion feature, comprising:

[0226] ,

[0227] wherein, is a node fusion feature of node i, is a nonlinear activation function.

[0228] In some embodiments, the global optimization of the node fusion feature obtains the global fusion feature, comprising:

[0229] initializing a label distribution of the node;

[0230] updating the label distribution based on message passing between the node and neighbor nodes until a preset condition is met, to obtain the global fusion feature;

[0231] wherein, updating the label distribution based on message passing between the node and neighbor nodes until a preset condition is met, to obtain the global fusion feature further comprises:

[0232] the node and neighbor nodes perform message passing, wherein the message , of node i, is a label of neighbor node j, a transition term for node i and neighbor node j, a label distribution for node i;

[0233] updating the label distribution based on the message comprises:

[0234] wherein, an updated label distribution for node i, a node term for node i, Z i a normalization constant;

[0235] obtaining the global fusion feature based on the updated label distribution comprises:

[0236] wherein, a global fusion feature for node i.

[0237] In some embodiments, scene segmentation is performed on a spatial region of the power transmission line based on geographic information and remote sensing image data, to obtain a plurality of scene regions with spatial continuity, comprising:

[0238] determining a first penalty value based on a vertical distance from two pixel points in the remote sensing image data to a corresponding nearest power transmission line, and whether the two pixel points correspond to the same power transmission line section of the nearest power transmission line;

[0239] determining a second penalty value based on a vertical distance from two pixel points in the remote sensing image data to a corresponding nearest tower, and whether the two pixel points correspond to the same tower of the nearest tower;

[0240] determining a prior information distance based on a sum of the first penalty value and the second penalty value;

[0241] obtaining a distance metric based on the prior information distance, a color distance of the two pixel points, and a spatial distance, wherein the spatial distance is obtained based on geographic information of the two pixel points;

[0242] determining pixel points belonging to the same scene region based on the distance metric.

[0243] For the convenience of description, the above apparatus is described in various modules based on functions. Of course, the functions of each module can be implemented in one or more software and / or hardware when implementing the present disclosure.

[0244] The apparatus of the above embodiments is used to implement the corresponding power transmission line detection method based on multi-modal data fusion in any of the preceding embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here again.

[0245] Corresponding to any of the above embodiment methods based on the same technical concept, the disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the power line detection method based on multi-modal data fusion as described in any of the above embodiments.

[0246] The computer-readable medium of the embodiments includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0247] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the power line detection method based on multi-modal data fusion as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which are not described here.

[0248] Those skilled in the art should understand that the above discussion of any of the embodiments is only exemplary and is not intended to imply that the scope (including claims) of the disclosure is limited to these examples; under the idea of the disclosure, the above embodiments or technical features in different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of different aspects of the embodiments of the disclosure as described above. In order to be brief, they are not provided in details.

[0249] Additionally, to simplify the description and discussion, and so as not to obscure the understanding of the embodiments of the disclosure, the well-known power / ground connections of the integrated circuits (ICs) and other components can or can not be shown in the provided figures. Furthermore, the apparatus can be shown in block diagram form in order to simplify and advance the description of such embodiments and also to highlight the fact that the details regarding how the apparatus is implemented, e.g., in terms of its working details, are highly dependent on the platform within which the embodiments of the disclosure are to be implemented (i.e., these details should be well within the understanding of one of ordinary skill in the art). In situations where detailed circuitry is set forth in order to describe the exemplary embodiments of the disclosure, it should be understood that the disclosure can be practiced with the full understanding and

[0250] Although the present disclosure has been described in connection with certain embodiments, numerous alternatives, modifications, and variations can become apparent to those skilled in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.

[0251] Embodiments of the present disclosure are intended to cover any alternatives, modifications, and variations of the present disclosure falling within the scope of the appended claims. Accordingly, any and all such alternatives, modifications and variations should be included in the scope of the present disclosure. Thus, any omission, modification, substitution, improvement, or the like that is made within the spirit and principle of the embodiments of the present disclosure should be included in the scope of the present disclosure.

Claims

1. A transmission line detection method based on multimodal data fusion, characterized in that, include: Based on geographic information and remote sensing image data, the spatial area of ​​the transmission line is segmented to obtain multiple scene areas with spatial continuity. Acquire raw multimodal data of the transmission line; wherein, the raw multimodal data includes sensor data and UAV data containing the location information of the transmission line; Target multimodal data of the target region within a target time period are obtained by filtering the original multimodal data based on preset distance conditions; wherein, the target region includes a first scene region where the target location is located and a second scene region adjacent to the first scene region, and the target time period includes the target time. A graph structure representing the relationship between the target multimodal data is obtained based on the target multimodal data and the corresponding location information; Feature extraction is performed on the target multimodal data to obtain target multimodal features; Based on the graph structure, feature fusion is performed on the target multimodal features to obtain global fused features, including: The attention coefficient of the node is determined based on the node weight, edge weight, and the target multimodal features; Based on the attention coefficient, feature fusion is performed on the target multimodal features of the node to obtain node fusion features; The global fusion feature is obtained by globally optimizing the node fusion feature; The attention coefficient includes: , , , in, Let be the attention coefficients between node i and its neighbor node j, and exp be an exponential function. Let j and k be the set of neighboring nodes of node i, where j and k are the indices of the neighboring nodes. For activation function, Let W be the attention vector and W be the weight matrix. The weighted sum of the node weights of node i and the target multimodal features. h is a weighted sum of the node weights of neighbor node j and the target multimodal features. k weighted This is a weighted sum of the node weights of neighbor node k and the target multimodal features. For feature splicing operations, w represents the edge weight between node i and its neighbor node j. ik Let be the edge weight between node i and its neighbor node k, where k represents the neighbor node of node i; The edge weights are: , where w ij For nodes i and neighboring nodes j edge weights, For nodes i and neighboring nodes j Data type relationships, ; For nodes i and neighboring nodes j Spatial distance between them Parameters used to control the rate of weight decay; The operating status of the target location of the transmission line at the target time is determined based on the global fusion features.

2. The method according to claim 1, characterized in that, Based on the target multimodal data and corresponding location information, a graph structure representing the relationship between the target multimodal data is obtained, including: The nodes of the graph structure are determined based on each of the target multimodal data; In response to the spatial distance between the nodes being less than a preset threshold, an edge is constructed between the nodes, and the edge weight is determined based on the data type of the nodes and the spatial distance. The nodes and edges are traversed based on preset weight rules to update the node weights of the nodes and / or the edge weights of the edges.

3. The method according to claim 1, characterized in that, Based on the attention coefficient, feature fusion is performed on the target multimodal features of the node to obtain node fused features, including: , in, Let i be the node fusion feature. It is a non-linear activation function.

4. The method according to claim 3, characterized in that, The global fusion feature is obtained by globally optimizing the node fusion feature, including: Initialize the label distribution of the nodes; The label distribution is updated based on message passing between the node and its neighboring nodes until a preset condition is met, thus obtaining the global fusion feature; The process of updating the label distribution based on message passing between the node and its neighboring nodes until a preset condition is met, to obtain the global fusion feature, further includes: The node exchanges messages with its neighboring nodes, where the messages between node i and neighboring node j are... , Let i be the label of node i. Let j be the label of the neighbor node j. Let be the transition term between node i and its neighbor node j. Let i be the label distribution of node i; Updating the label distribution based on the message includes: ,in, Let i be the updated label distribution. Z is the node item of node i. i This is a normalization constant; The global fusion feature is obtained based on the updated label distribution, including: ',in, Let i be the global fusion feature of node i.

5. The method according to claim 1, characterized in that, Based on geographic information and remote sensing image data, the spatial area of ​​the transmission line is segmented to obtain multiple spatially continuous scene areas, including: Based on the vertical distance from two pixels in the remote sensing image data to the corresponding nearest transmission line, and whether the nearest transmission line corresponding to the two pixels belongs to the same transmission line network segment, a first penalty value is determined. A second penalty value is determined based on the vertical distance from two pixels in the remote sensing image data to the corresponding nearest tower, and whether the two pixels to the corresponding nearest tower belong to the same tower. The prior information distance is determined based on the sum of the first penalty value and the second penalty value; A distance metric is obtained based on the prior information distance, the color distance between the two pixels, and the spatial distance; wherein, the spatial distance is obtained based on the geographic information of the two pixels. Pixels belonging to the same scene region are determined based on the distance metric.

6. A transmission line detection device based on multimodal data fusion, characterized in that, include: The region segmentation module is used to segment the spatial region of the transmission line based on geographic information and remote sensing image data, and obtain multiple scene regions with spatial continuity. The data acquisition module is used to acquire the raw multimodal data of the transmission line; wherein, the raw multimodal data includes sensor data and UAV data containing the location information of the transmission line; The data filtering module is used to filter target multimodal data of a target region within a target time period from the original multimodal data based on preset distance conditions; wherein, the target region includes a first scene region where the target location is located and a second scene region adjacent to the first scene region, and the target time period includes a target time. The graph structure module is used to obtain a graph structure representing the relationship between the target multimodal data based on the target multimodal data and the corresponding location information; The feature extraction module is used to extract features from the target multimodal data to obtain target multimodal features; The feature fusion module is used to perform feature fusion on the target multimodal features based on the graph structure to obtain global fused features; A status detection module is used to determine the operating status of the target location of the transmission line at the target time based on the global fusion features; The method of fusing the target multimodal features based on the graph structure to obtain global fused features includes: The attention coefficient of the node is determined based on the node weight, edge weight, and the target multimodal features; Based on the attention coefficient, feature fusion is performed on the target multimodal features of the node to obtain node fusion features; The global fusion feature is obtained by globally optimizing the node fusion feature; The attention coefficient includes: , , , in, Let be the attention coefficients between node i and its neighbor node j, and exp be an exponential function. Let j and k be the set of neighboring nodes of node i, where j and k are the indices of the neighboring nodes. For activation function, Let W be the attention vector and W be the weight matrix. The weighted sum of the node weights of node i and the target multimodal features. h is a weighted sum of the node weights of neighbor node j and the target multimodal features. k weighted This is a weighted sum of the node weights of neighbor node k and the target multimodal features. For feature splicing operations, w represents the edge weight between node i and its neighbor node j. ik Let be the edge weight between node i and its neighbor node k, where node k is a neighbor node of node i; The edge weights are: , where w ij For nodes i and neighboring nodes j edge weights, For nodes i and neighboring nodes j Data type relationships, ; For nodes i and neighboring nodes j Spatial distance between them Parameters used to control the rate of weight decay.

7. The apparatus according to claim 6, characterized in that, Based on the target multimodal data and corresponding location information, a graph structure representing the relationship between the target multimodal data is obtained, including: The nodes of the graph structure are determined based on each of the target multimodal data; In response to the spatial distance between the nodes being less than a preset threshold, an edge is constructed between the nodes, and the edge weight is determined based on the data type of the nodes and the spatial distance. The nodes and edges are traversed based on preset weight rules to update the node weights of the nodes and / or the edge weights of the edges.

8. The apparatus according to claim 7, characterized in that, Based on the attention coefficient, feature fusion is performed on the target multimodal features of the node to obtain node fused features, including: , in, Let i be the node fusion feature. It is a non-linear activation function.

9. The apparatus according to claim 8, characterized in that, The global fusion feature is obtained by globally optimizing the node fusion feature, including: Initialize the label distribution of the nodes; The label distribution is updated based on message passing between the node and its neighboring nodes until a preset condition is met, thus obtaining the global fusion feature; The process of updating the label distribution based on message passing between the node and its neighboring nodes until a preset condition is met, to obtain the global fusion feature, further includes: The node exchanges messages with its neighboring nodes, where the messages between node i and neighboring node j are... , Let i be the label of node i. Let j be the label of the neighbor node j. Let be the transition term between node i and its neighbor node j. Let i be the label distribution of node i; Updating the label distribution based on the message includes: ,in, Let i be the updated label distribution. Z is the node item of node i. i This is a normalization constant; The global fusion feature is obtained based on the updated label distribution, including: ',in, Let i be the global fusion feature of node i.

10. The apparatus according to claim 6, characterized in that, Based on geographic information and remote sensing image data, the spatial area of ​​the transmission line is segmented to obtain multiple spatially continuous scene areas, including: Based on the vertical distance from two pixels in the remote sensing image data to the corresponding nearest transmission line, and whether the nearest transmission line corresponding to the two pixels belongs to the same transmission line network segment, a first penalty value is determined. A second penalty value is determined based on the vertical distance from two pixels in the remote sensing image data to the corresponding nearest tower, and whether the two pixels to the corresponding nearest tower belong to the same tower. The prior information distance is determined based on the sum of the first penalty value and the second penalty value; A distance metric is obtained based on the prior information distance, the color distance between the two pixels, and the spatial distance; wherein, the spatial distance is obtained based on the geographic information of the two pixels. Pixels belonging to the same scene region are determined based on the distance metric.

11. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in any one of claims 1 to 5.

12. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-modal knowledge graph method based on power grid dispatching

    CN117171358A

  • Multi-mode power transmission line state monitoring method and system based on thunder-shooting fusion

    CN120601621A