Method for positioning accurate position of middle and low voltage power distribution equipment based on multi-modal perception

Through multimodal perception and positioning technology, combined with binocular cameras and RTK positioning, and using methods such as DGCNN, KAF, CLIP and GDCF, the problems of inaccurate positioning and incomplete identification of medium and low voltage distribution equipment have been solved, achieving centimeter-level equipment positioning and efficient identification, and improving the accuracy and automation of layout.

CN120668102APending Publication Date: 2025-09-19GUANGDONG UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510623354.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing methods for collecting the positions of medium and low voltage distribution equipment have problems such as insufficient positioning accuracy, inaccurate equipment identification, easy occlusion, and insufficient generalization ability of single modal recognition methods in small sample environments, which affect the accuracy and efficiency of layout.

Method used

A multimodal perception-based method is adopted, combined with binocular cameras and RTK high-precision positioning. Through MPR multimodal recognition and GCP geometric solution, feature extraction and fusion are performed using technologies such as DGCNN, KAF, CLIP and GDCF to achieve high-precision positioning and identification of equipment.

Benefits of technology

It achieves centimeter-level device positioning accuracy and recognition accuracy in complex environments, improves the efficiency and intelligence level of layout drawing, and reduces labor costs and error risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120668102A_ABST
    Figure CN120668102A_ABST
Patent Text Reader

Abstract

The invention provides a medium and low voltage power distribution equipment accurate position positioning method based on multi-mode perception, and belongs to the technical field of medium and low voltage equipment management. S1, the shooting instrument is started and positioned; s2, shooting a target device; s3, data preprocessing; s4, adopting MPR (Maximum Power Ratio) multi-mode identification; s5, resolving the position of the GCP; and S6, storing the actual position of the target equipment. According to the invention, through multi-modal data acquisition, an innovative feature fusion mechanism and an accurate key point extraction and positioning method, the efficiency and precision of drawing the edge layout of the medium and low voltage power distribution equipment are significantly improved. Compared with a traditional manual acquisition and identification mode, the method not only reduces the labor cost and error risk, but also improves the automation and intelligence level of data acquisition, and has higher engineering application value and popularization potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception, belonging to the technical field of medium and low voltage equipment management. Background Art

[0002] With the development of power networks, the mapping of medium and low voltage distribution equipment has become a crucial component of power system maintenance and management. A distribution equipment layout diagram (also known as a distribution network diagram or distribution system diagram) depicts the path of electricity from the substation to the end user. It illustrates all major components of the power distribution network and how they are interconnected. These components may include substations, switchyards, cables, overhead lines, circuit breakers, disconnectors, load switches, capacitor banks, and more. A layout diagram typically identifies the location of each piece of equipment and how they are physically connected. When constructing new facilities or expanding existing ones, a layout diagram can help engineers and designers determine the optimal routing of lines and equipment locations to ensure the most efficient power transmission. When a power system fault occurs, such as a short circuit, ground fault, or overload, a layout diagram can quickly help technicians identify the affected area, expediting fault location and repair. A distribution equipment layout diagram is an integral part of power grid operations and is crucial for ensuring safe, reliable, and efficient operation.

[0003] Currently, the location data collection for medium and low voltage distribution equipment along the layout is done manually by field workers. These workers record the equipment's location data based on their own positioning and then submit the collected data to engineers at the equipment center. The engineers then manually draw a layout map for the medium and low voltage equipment. This current method of collecting location data for medium and low voltage distribution equipment has the following drawbacks:

[0004] 1. Positioning accuracy challenges. Currently, workers typically use handheld GPS devices or mobile terminals to record locations. However, due to limitations in device accuracy, signal interference, and manual errors, it's difficult to accurately calibrate the actual location of distribution equipment. In areas with dense buildings, obstructed by trees, or with strong electromagnetic interference, GPS signal accuracy decreases, resulting in significant deviations in location data. This error further impacts the accuracy of equipment layout, which in turn affects subsequent operations, maintenance, and scheduling decisions.

[0005] 2. Power distribution equipment is diverse. This equipment includes transformers, ring main units, switch stations, distribution boxes, and pole-mounted circuit breakers, all with varying appearances, sizes, and installation methods. Currently, equipment identification relies primarily on the experience of on-site personnel, lacking standardized automated identification methods. Different personnel may have misunderstandings during the identification process, leading to inaccurate equipment type labeling. Furthermore, some equipment may have unclear nameplate information due to aging, modification, or repainting, further complicating identification.

[0006] 3. Power distribution equipment is easily obscured. Because power distribution equipment is often installed in environments such as roads, residential areas, green belts, or industrial parks, it is easily obscured by obstacles such as trees, buildings, and fences, making it difficult for workers to directly observe and measure their location. For example, some box transformers or distribution cabinets may be obscured by walls or billboards, making it impossible to directly measure their exact location. Workers may even need to take detours or use additional tools to obtain data. This not only increases the workload for data collection but can also lead to omissions or misjudgments, compromising the completeness and accuracy of the layout.

[0007] Currently, there is no complete and mature method for accurately capturing the location of medium- and low-voltage power distribution equipment. Existing technologies primarily fall into two categories: visible light image target detection and point cloud data modeling. Visible light image target detection solutions can more accurately identify power distribution equipment, reducing manual labeling errors caused by the wide variety of equipment, and improving recognition efficiency. However, due to the lack of depth information in two-dimensional images, this method struggles to directly obtain the spatial location of the equipment. Furthermore, in complex environments, such as when equipment is obscured by trees, buildings, or other structures, target detection accuracy can be significantly reduced, compromising complete device identification and positioning. Point cloud data modeling solutions, which utilize the three-dimensional structure of the target for analysis, offer advantages in handling occlusion issues and directly obtain the spatial location of the equipment after successful recognition. However, compared to image data, the acquisition process for point cloud data is more complex, limited by factors such as acquisition equipment, data storage, and computational costs. Currently, the scarcity of existing point cloud data on medium- and low-voltage power distribution equipment limits the promotion and optimization of point cloud modeling methods in practical applications. Summary of the Invention

[0008] Given the limitations of existing technologies for collecting the locations of medium and low voltage distribution equipment, this paper proposes a method for accurately collecting the locations of medium and low voltage distribution equipment based on multimodal perception and positioning. This method aims to improve the accuracy and efficiency of mapping distribution equipment along a layout. Through multimodal fusion and high-precision calculation methods, it can achieve higher recognition accuracy even with small sample sizes. This solution is divided into a recognition phase (MPR) and a relocalization phase (GCP). The first step, the MPR method (Multi-modal Perception for Recognition), uses a binocular camera to simultaneously collect visible light images and point cloud data to construct a multi-source dataset. It then utilizes a multimodal fusion method of multi-angle images and point clouds to accurately identify distribution equipment, extract key points of the target equipment, and obtain spatial distance information relative to the camera. The MPR solution effectively alleviates the recognition difficulties caused by the wide variety of distribution equipment and environmental occlusion, improving the integrity and accuracy of equipment information. The second step, the GCP method (Geometric Computation-based Positioning), uses high-precision satellite positioning technology based on an RTK system integrated with a 5G network to calculate the precise location of the camera. Then, based on the spatial distance information of the target device relative to the shooting instrument identified by MPR, multiple sensors are used to obtain data, and then a geometric calculation model is used to perform coordinate conversion and precise deduction, and finally the actual geographical location of the target device is solved to achieve centimeter-level positioning accuracy. This solution effectively makes up for the limitation that traditional RTK can only provide the coordinates of the shooting device and cannot directly obtain the location of the target device. Overall, the present invention combines multimodal intelligent perception with high-precision positioning technology, which not only enhances the ability to identify distribution equipment in complex environments, but also significantly improves the accuracy of location information, and ultimately improves the level of intelligence in the layout drawing of medium and low voltage distribution equipment.

[0009] The specific technical solutions are:

[0010] The method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception includes the following steps:

[0011] S1. Power on and locate the camera: After the camera is powered on, it performs 5G positioning and RTK secondary positioning calibration to obtain the current precise positioning of the instrument. The camera uses a binocular camera with a built-in edge computing development board, which can analyze and process the collected point cloud data and image data in real time. Both the MPR algorithm and the GCP algorithm are deployed in the camera, realizing the full process of automated detection and positioning of medium and low voltage distribution equipment at the edge.

[0012] S2. Photograph the target device: Use a camera to photograph the target device and obtain visible light image data and point cloud data. Each visible light image and each set of point cloud data is accompanied by a timestamp and the precise positioning of the camera at the moment of capture. The visible light image data and point cloud data captured at the same moment are aligned to ensure they correspond to the same power distribution device.

[0013] S3. Data preprocessing: Preprocess the visible light image data. Through the calibration and spatial registration method of the binocular camera, the three-dimensional points in the point cloud are mapped to the image pixels to ensure the spatial consistency of the point cloud and image data.

[0014] S4. Use MPR multimodal recognition: Use the MPR method to identify and detect multi-source data, accurately identify the target device, and finally use the key three-dimensional coordinate information in the point cloud data to calculate the relative distance S between the shooting instrument and the distribution equipment.

[0015] S5. GCP Position Calculation: Based on the geometric positioning (GCP) method and combined with the RTK positioning data from the camera, the relative distance S is converted to coordinates to calculate the actual horizontal distance S' of the target device. Finally, the actual geographic coordinates of the target device are determined through geometric calculation, ensuring centimeter-level accuracy.

[0016] S6. Save the actual location of the target device. The calculated actual location information of the target device is stored in the photographing device in a standard format of category-serial number-location information.

[0017] The multimodal perception and recognition MPR method in S4 specifically includes the following steps:

[0018] (1) Preparation of training dataset

[0019] The dataset partitioning method follows the basic framework of Few-Shot Learning (FSL) and adopts an N-way K-shot scheme to ensure that the model can efficiently distinguish categories and learn target key points under limited sample conditions. The dataset is divided into a training set, a validation set, and a testing set, and is optimized in combination with a support set and a query set to improve generalization.

[0020] The training set, used for model feature learning, covers multiple categories of power distribution equipment, with each category containing only K samples (K-shots). Assuming the number of categories is N, the total number of samples in the training set is N × K. In each training round, the training data is further divided into a support set and a query set. The support set is used for feature learning, while the query set is used for model optimization.

[0021] The validation set is used to monitor the model's generalization performance during training. It uses data from categories not used in training to ensure that the model does not overfit to samples from a specific category. The test set is used for final performance evaluation and has no overlap with the training categories to ensure that the model can generalize to new categories of power distribution equipment.

[0022] Data partitioning uses a stratified sampling method to ensure the balance of data across different categories. The proposed training and test dataset partitioning scheme ensures that the MPR+GCP solution can stably learn the key features of target devices even in a small sample environment, and achieve high-precision target recognition and positioning during inspection tasks.

[0023] (2) A multi-view depth rendering method is used, combined with a minimum depth selection strategy and an adaptive extended projection method to optimize the projection accuracy of 3D point clouds on 2D views. The generalization ability of the model is enhanced by randomizing the view angle and shooting distance, reducing the overfitting problem. This method not only improves the geometric consistency between 2D visual features and 3D point cloud features, but also effectively addresses the problem of missing projections caused by point cloud sparsity, providing more stable input for multimodal fusion in the MPR method.

[0024] In traditional point cloud projection methods, a single-view depth map often cannot fully reflect the global geometric features of an object, especially in complex environments, where some key points may be lost due to occlusion or projection distortion. This invention adopts a continuous three-view (XY, YZ, XZ view) depth rendering strategy to ensure that the geometric information of the object can be fully presented at different angles, and improves the projection quality through the optimal depth selection mechanism. Specifically, for a 3D point (x, y, z), the process of projecting it to a 2D depth map is as follows:

[0025] (XY plane projection)

[0026] (YZ plane projection)

[0027] (XZ plane projection)

[0028] For each perspective, the minimum depth selection strategy is adopted, that is, the depth value closest to the camera is selected from multiple projection points:

[0029]

[0030] This method can ensure the correctness of the occlusion relationship while avoiding depth blur and improving rendering accuracy. However, due to the sparsity of the 3D point cloud, direct projection may result in some pixel areas having no corresponding depth values, thereby affecting the integrity of the depth map. To this end, the present invention further proposes an adaptive extended projection method, which uses a neighborhood extension strategy to fill in blank areas to make the depth map more coherent. For pixels (X, Y) with missing depth values, a neighborhood extension radius R is defined, and the matching area M(X, Y, R) is determined by the following conditions:

[0031] M(X,Y,R)={(x,y,z)∈P∣XR / 2≤X' <X+R / 2,Y-R / 2≤Y'<Y+R / 2}

[0032] Finally, the depth value of the blank pixel is completed by the minimum depth in the neighborhood:

[0033] D(X,Y)=min{z∣(x,y,z)∈M(X,Y,R)}

[0034] Among them, when R=1, only the depth value of the pixel itself is used without expansion; when R>1, the neighboring points are used for completion, thereby enhancing the integrity of the depth map and avoiding the problem of traditional interpolation methods losing object boundary information. In addition, in order to prevent the model from overfitting to a fixed viewing angle, the present invention further introduces a randomized shooting distance mechanism. During each depth rendering, the position of the virtual camera is randomly perturbed to simulate target projections at different scales and improve the generalization ability of the model. Assume that the initial position of the camera C0=(x c ,y c ,z c ), the actual camera position at each rendering is:

[0035] C=(x c +Δx,y c +Δy,z c +Δz)

[0036] Here, Δx, Δy, and Δz follow a uniform distribution U(-r, r), where r is the preset random perturbation range. This approach ensures that the GCP does not overfit certain features due to a fixed view angle during depth rendering, while also improving the model's adaptability to different shooting distances.

[0037] The multi-view depth rendering method of the present invention ensures the projection quality through the minimum depth selection strategy, adopts the adaptive extension projection method to fill the blank areas caused by the sparse point cloud, and combines the randomized shooting distance strategy to enhance the generalization ability of the model, ultimately improving the recognition accuracy of the MPR method and the calculated geographic positioning stability, providing a more efficient and reliable technical solution for the precise location acquisition of medium and low voltage distribution equipment under small sample conditions.

[0038] (3) The DGCNN network is used to extract features from 3D point clouds, providing robust deep feature support for the MPR recognition process.

[0039] DGCNN (Dynamic Graph Convolutional Neural Network) is a deep learning model for 3D point cloud data that enhances the expressiveness of point cloud features through dynamic graph construction and EdgeConv operations. Compared to traditional point cloud processing methods (such as PointNet and PointNet++), DGCNN uses a dynamic k-nearest neighbor graph (k-NN graph) to recalculate the adjacency relationship between points in each layer, thereby achieving more effective local feature extraction. Its core EdgeConv operation not only retains the global information of the point cloud, but also uses neighborhood relationships to learn the geometric structure between points, improving the model's sensitivity to target morphology and its ability to model complex structures.

[0040] In the present invention, DGCNN is mainly used to extract features from 3D point clouds, providing strong and robust deep feature support for the MPR recognition process. Specifically, DGCNN is used to generate high-dimensional point cloud features, and the key points of the target device are further screened in the KAF (key point perception fusion) module. These key points are not only used for classification and recognition, but also serve as input for subsequent GCP position solution to improve the accuracy of device position calculation. The deep geometric information extracted by DGCNN can effectively enhance the semantic representation ability of the point cloud, and improve the recognition accuracy and generalization ability of the MPR scheme under small sample conditions, so that the present invention can more efficiently complete the detection and positioning of target equipment in power inspection tasks.

[0041] (4) The keypoint-aware fusion (KAF) module combines keypoint screening, local feature enhancement, and cross-instance matching to build an efficient point cloud feature extraction and fusion solution.

[0042] This paper proposes a Keypoint-Aware Fusion (KAF) module specifically designed to enhance keypoint extraction capabilities and improve the recognition and localization accuracy of 3D objects in small-sample point cloud classification tasks. Inspired by salient part fusion and cross-instance fusion enhancement, the KAF module combines keypoint screening, local feature enhancement, and cross-instance matching to construct an efficient point cloud feature extraction and fusion solution.

[0043] In point cloud classification and target recognition tasks, some key points are more representative of the target category, and not all points contribute equally to the recognition task. To this end, KAF uses the key point selection mechanism to select the most discriminative key points from the 3D point cloud features extracted by DGCNN in the support set and query set. Specifically, let the input point cloud be P = {p1,p2,...,p N}, where each point p i With d-dimensional feature representation F(p i ), then calculate the saliency score of each point:

[0044]

[0045] Among them, F g It is a global feature, calculated by Global Max-Pooling:

[0046]

[0047] Finally, take the top k with the highest significance scores s points as the key point set P k Since point cloud data usually has local noise and occlusion problems, relying solely on the features of a single point may lead to misjudgment. Therefore, KAF adopts a neighborhood feature aggregation strategy to enhance the local features of key points and extract their k nearest neighbor information for local feature enhancement. For each key point p j , construct its neighborhood feature N j As a set of k nearest neighbor points:

[0048] N j ={p j1 ,p j2 ,...,p jk},p ji ∈k-NN(p j )

[0049] Then, the global feature F g Connect to neighborhood features to enhance the keypoint’s ability to represent the target category:

[0050] F e,j =g([F g ,N j ])

[0051] Here, g(·) represents a 1×1 convolutional mapping operation, which ensures the consistency of information at different scales. This allows local features to not only retain their own information but also combine global features to improve the discrimination capability.

[0052] At the same time, in small sample point cloud tasks, there may be a problem of feature distribution mismatch between the support set and the query set. The second stage of KAF optimizes the feature matching of key points through a cross-instance fusion mechanism. Given the global prototype feature matrix of the support set (containing d-dimensional prototype features of N categories) and the query set global feature matrix (Including N q d-dimensional features of query instances), first calculate the cross-instance relationship graph through the scaled dot product attention mechanism:

[0053]

[0054] in is a scaling factor (default d = 256), which is used to control the numerical range of the matrix product to avoid the gradient vanishing problem. It represents the semantic association strength between the i-th class prototype and the j-th query instance, and its value range is [0,1].

[0055] In the dual-branch feature fusion design, the channel-level fusion branch is used for each prototype feature. Filter out the 5 most relevant query features by cosine similarity Construct the interaction tensor:

[0056]

[0057] Then, channel interaction is achieved through two levels of convolution. The first layer of 1×1 convolution maps the channel dimension from d to the hidden layer h r =64: The second layer of 1×1 convolution restores to the original dimension d: Finally, feature update is achieved through attention weighting:

[0058]

[0059] Where ⊙ represents element-wise multiplication, which explicitly enhances the local geometric features related to the target category. The instance-level fusion branch propagates features through the global relationship matrix:

[0060]

[0061] This branch will query the features in M pq A linear combination of weights is performed so that each prototype absorbs contextual information across instances. For example, for the i-th class prototype, its corresponding instance-level feature is:

[0062]

[0063] In the feature aggregation stage, the dual-branch outputs are fused through learnable weight parameters α and β:

[0064]

[0065] The innovation of the KAF module lies in its unique key point screening strategy. By calculating the cosine similarity between point cloud features and global features, it automatically identifies the most discriminative key points and constructs a key point neighborhood to enhance the ability to express local information. Compared with the traditional global point cloud feature extraction method, this method can more accurately extract geometric information related to the target device category, avoid the interference of redundant features, and improve the discriminative ability of features. In addition, KAF adopts a local feature enhancement strategy to construct dynamic neighborhood features for selected key points to retain local structural information, and at the same time combines global features for feature fusion, making the geometric information of the point cloud more complete and reducing the impact of environmental noise on target features. This method is particularly suitable for complex scenarios in power inspection tasks, such as occlusion, lighting changes, and point cloud sparseness, ensuring that the system can still maintain high-precision recognition in real inspection environments.

[0066] In the matching process of the support set and the query set, KAF further introduces a cross-instance fusion mechanism. By calculating the similarity matrix between the support set samples and the query samples, it adaptively optimizes the feature expression, so that the query samples can be more accurately matched to the most relevant support samples, thereby improving the effect of small-sample learning. Compared with traditional small-sample point cloud classification methods, this mechanism can effectively reduce classification errors caused by insufficient support set samples and improve the model's generalization ability on new categories. In addition, KAF combines 3D key point saliency screening, local feature enhancement, and cross-instance feature matching, which not only improves the classification accuracy of 3D targets, but also ensures the stability of key point information, making the 3D key point coordinates finally passed to GCP for geographic location calculation more accurate and reliable.

[0067] After adopting the KAF module, the present invention can achieve more efficient feature extraction in the 3D point cloud processing of the MPR scheme and ensure high robustness in a small sample environment. The high-quality 3D key point information generated by KAF is not only used for the classification of target devices, but also provides accurate input data for GCP, ensuring that GCP can efficiently calculate the geographic coordinates of the target device with the assistance of RTK data, thereby improving the positioning accuracy in the inspection task. In addition, KAF adopts adaptive significance scoring and cross-instance optimization strategies, so that this patent can adapt to different inspection environments in actual applications, and maintain stable recognition performance even in the case of data scarcity. At the same time, this module does not require additional manual labeling costs, so that the overall computational overhead of the system is kept at a low level, further enhancing the engineering application value of the MPR+GCP solution in power grid inspection tasks.

[0068] (5) CLIP network is used as a 2D visual feature extraction module to enhance the image perception capability of the MPR scheme for distribution equipment.

[0069] This paper uses the CLIP (Contrastive Language-Image Pretraining) network as a 2D visual feature extraction module to enhance the image perception capability of the MPR solution for power distribution equipment. Proposed by OpenAI, CLIP is based on a contrastive learning method and is trained on large-scale image-text pairs, enabling the model to have strong generalization performance in unsupervised or small-sample environments. Compared to traditional single-vision models such as CNN or ViT (Vision Transformer), CLIP excels in cross-modal learning and can learn universal visual representations through text-image alignment, allowing the model to adapt to new category recognition tasks without additional supervision.

[0070] In our MPR+GCP framework, CLIP is primarily used to extract global features from 2D images and provide additional semantic context through textual information. Specifically, CLIP uses a pre-trained image encoder to visually encode device images. By extracting these features, it generates a comprehensive prototype that includes the features of the 2D image prototype, providing a richer and more recognizable feature description of power distribution equipment.

[0071] During the multimodal fusion phase, the 2D prototype features generated by CLIP are aligned and fused with the keypoint features of the 3D point cloud. The visual features provided by CLIP, using the 3D keypoint information extracted by the KAF module, complement the prototype features, enhancing the model's ability to recognize different device morphologies and detailed features. Especially when faced with objects with similar morphologies but different detailed features, the image information provided by CLIP helps distinguish device categories, thereby improving recognition accuracy and stability.

[0072] The introduction of CLIP not only optimizes the utilization of 2D visual information, but also enables the present invention to effectively improve the generalization ability of 3D target recognition in small sample environments. By linking with the KAF key point perception module, the global features extracted by CLIP can assist in the screening and optimization of 3D key points, making the features ultimately passed to the GCP for positioning calculations more stable and reliable. In addition, CLIP adopts a Transformer structure, which gives it stronger contextual modeling capabilities. In complex inspection environments, even if there is a certain degree of occlusion or interference in the image information, the recognition accuracy of the model can be ensured.

[0073] (6) The gated dual-path feature fusion (GDCF) mechanism is used to fuse 3D key point features (output of the KAF module) and 2D visual features (output of the CLIP network).

[0074] This paper designs a gated dual-path cross-modal feature fusion mechanism (GDCF) to fuse 3D keypoint features (output of the KAF module) and 2D visual features (output of the CLIP network). This mechanism combines global perspective aggregation and gated fusion to ensure efficient interaction between 3D structural information and 2D visual features in a unified feature space and enhance the distinguishability of features. The final fused features can support small sample target classification and optimize the positioning calculation of the GCP module.

[0075] In the GDCF mechanism, the 3D key point feature F is first P and 2D visual features F I Perform global feature aggregation, where F P and F I Both contain feature matrices of their respective support sets and query sets, Figure 6 The support set operation is displayed, and the query set is similar. Since 3D key point features are often local, while 2D visual features have stronger global perception capabilities, global perspective aggregation is used to integrate multi-view 2D visual features into a global visual representation through a learnable linear transformation:

[0076]

[0077] Among them, f1 and f2 are two-layer linear mappings, N is the number of perspectives, and concat(·) represents the splicing operation on the channel dimension. In this way, a stable 2D visual global feature G can be obtained. I , ensuring that information is consistent across different perspectives.

[0078] For 3D keypoint features, due to their strong geometric information, it is necessary to ensure that different keypoint features can be correctly aggregated. To this end, a learnable transformation layer is introduced after the KAF output to align it with the global feature space of 2D visual features:

[0079] G P =W P F P +b P

[0080] Among them, W P and b P It is a learnable parameter used to adjust the representation ability of 3D key point features to meet the needs of cross-modal fusion.

[0081] In the fusion stage, a gated fusion mechanism is used to dynamically adjust the contribution ratio of 3D key point features and 2D visual features through learnable gate weights σ:

[0082] G=σ·G I +(1-σ)·G P

[0083] Among them, σ is calculated by the learnable Sigmoid transformation:

[0084] σ=Sigmoid(W G G P +b G )

[0085] In this way, the model can automatically learn how to adjust the fusion ratio of 3D keypoint information and 2D visual information during training, ensuring the optimal feature combination for different task scenarios. Ultimately, the fused global feature G is passed to the prototype network for object classification, while the optimized 3D keypoint information is input into the GCP module to improve the precise positioning capability of distribution equipment.

[0086] Global perspective aggregation ensures the stability and consistency of 2D visual features and avoids the modal bias problem caused by insufficient single-perspective information. Since the 2D visual features extracted by CLIP may be affected by different shooting angles, resulting in uneven information distribution, GDCF fuses multi-perspective information into a stable global feature through global feature splicing and linear transformation, ensuring that the visual features do not lose key morphological information during the fusion stage, while improving the model's ability to recognize target devices at different angles. In addition, due to the modal differences in data distribution between 3D key point features and 2D visual features, GDCF uses a learnable transformation layer to ensure that 3D key point features can be aligned with 2D visual features to enhance the mutual matching ability of different modal features. This transformation layer can adapt to different categories of power distribution equipment, so that the 3D key point features can be better mapped to the global feature space generated by CLIP after conversion, thereby improving the fusion effect.

[0087] During the fusion stage, GDCF adopts a gated fusion mechanism, which dynamically adjusts the contribution ratio of 3D key point features and 2D visual features through adaptive learning of the gating weight σ to ensure the optimal fusion of different modal information. Compared with the fixed weighting strategy, gated fusion can automatically adjust the fusion ratio of features according to the specific task scenario. For example, in a scenario where the geometric information is strong but the texture information is weak, GDCF will tend to give higher weights to 3D key point features to ensure that the model can accurately identify the morphological characteristics of the target device; in the case where the 3D key points are sparse or the point cloud quality is poor, GDCF will enhance the contribution of 2D visual features and use global texture information to make up for the shortcomings of 3D key point features, thereby improving the final classification and positioning stability. This gating mechanism not only optimizes the 3D-2D fusion strategy, but also enables the multimodal perception solution of the present invention to have stronger environmental adaptability, and can maintain high-precision recognition and positioning effects in complex inspection tasks.

[0088] After the present invention adopts the GDCF mechanism, compared with the traditional 3D key point independent classification method, the fused features can effectively improve the recognition accuracy of the MPR scheme in a small sample environment, and ensure that the key point information calculated by GCP is more stable, making the centimeter-level positioning of the target device more accurate and reliable. The GDCF mechanism can not only enhance the cross-modal feature expression capability under small sample conditions, so that the system can maintain a high classification performance even when data is limited, but also can make full use of 2D visual information for compensation when 3D point cloud data is missing or of low quality, thereby improving the robustness and environmental adaptability of target detection in inspection tasks. In addition, because GDCF adopts a learnable dynamic weight strategy, the model can automatically adapt to changes in data quality in different inspection environments without manually adjusting feature fusion parameters, thereby reducing the workload of model debugging, making the present invention have higher engineering application value in actual power grid inspection tasks.

[0089] (7) The innovative prototype network superimposes the key point calculation function on the traditional prototype network. The trained traditional prototype network can only complete the classification task, while the innovative prototype network can realize key point solution, finely process the key point features, calculate the relative distance S between the key point and the camera, and pass the optimized data to the GEP for position solution.

[0090] In the MPR+GEP framework of this invention, the innovative Prototype Network not only completes the efficient classification of target devices in a small sample environment, but also plays a core role in optimizing and extracting 3D key point information, ensuring the provision of high-quality key point data for the GEP (Geometric Estimation-based Positioning) module, thereby achieving accurate centimeter-level positioning.

[0091] In the fusion stage, the GDCF mechanism has deeply integrated the 3D key point features extracted by the KAF module with the 2D visual features generated by the CLIP network to obtain the fused high-dimensional features F final The prototype network first receives these fused features and calculates the prototype vector of each category based on the support set samples. The prototype of each category not only represents the overall feature center of the category, but also retains the geometric information of the fused 3D key points. After the key point solution module, the key points representing the sample can be found. Specifically, assuming that the support set sample feature set of category c is S c ={x1,x2,...,x k}, then the category prototype vector P c The calculation is as follows:

[0092]

[0093] The prototype vector P c It not only represents the category center in the feature space, but also has an optimized 3D key point structure. The prototype network combines these prototype vectors with the query sample F q Perform distance measurement and complete category classification:

[0094] d(F q ,P c )=‖F q -P c ‖2

[0095] This paper introduces a prototype reverse decoding module (Prototype-to-Keypoints, P2K) to directly reconstruct the three-dimensional keypoint coordinates of the target device from the category prototype vector generated by the prototype network. This module maps the fused prototype features into the spatial coordinates of three geometric keypoints by constructing a learnable nonlinear decoding network. Let the prototype vector be The number of key points is 3, each key point is The reverse decoding process can be uniformly expressed by the following formula:

[0096] [k1,k2,k3]=Reshape 3×3 (φ(p c ))

[0097] φ is a mapping function composed of a multi-layer perceptron (MLP), Reshape 3×3 Indicates reshaping the 9-dimensional vector into 3 3-dimensional coordinate points;

[0098] After obtaining the coordinates of the three key points, the system further calculates the Euclidean distance of each key point with the three-dimensional space coordinates of the current shooting instrument:

[0099] d i =‖k i -c‖,i=1,2,3

[0100] Finally, the three distances are averaged as the estimated spatial distance of the device relative to the shooting instrument:

[0101]

[0102] The optimized keypoint data and relative distance S are encapsulated and passed to the GEP module. After receiving this data, GEP combines it with RTK high-precision positioning data to further calculate the actual geographic coordinates of the target device. The prototype network ensures the accuracy and stability of the keypoint information transmitted to GEP, greatly improving the final positioning accuracy and ensuring centimeter-level positioning of the target device.

[0103] In this paper, the innovative prototype network not only completes the small sample classification task, but also provides reliable and high-quality keypoint data for GEP through optimized 3D keypoint extraction and relative distance S calculation. This design greatly improves the recognition and positioning capabilities of the MPR+GEP framework in small sample environments, ensuring that the system can still efficiently and accurately complete target recognition and positioning in power inspection tasks even in data-scarce or complex environments.

[0104] S5's GCP method (Geometric Computation-based Positioning) includes the following steps:

[0105] (1) Determine the relative position of the shooting instrument

[0106] The electronic compass module can automatically obtain the offset angle ω of the current shooting instrument and determine the accurate position of the medium and low voltage distribution equipment relative to the shooting instrument. The relative distance S is obtained using the MPA method.

[0107] (2) Obtaining the actual horizontal distance

[0108] The gyroscope module can obtain the pitch angle θ of the shooting instrument relative to the key point of the power distribution equipment, and obtain the actual distance S' according to the trigonometric function relationship.

[0109] (3) Superposition calculation

[0110] The actual distance S' and offset angle ω are superimposed on the RTK positioning information to obtain the precise coordinates (long2, lat2) of the actual distribution equipment. The specific formula is as follows. Assume that the known positioning information of the camera is (long1, lat1) in degrees.

[0111] long2=long1+S'×sinω / [ARC×cos(lat1)×2π / 360]

[0112] lat2=lat1+S'cosω / (ARC*2π / 360)

[0113] ARC is the average radius of the Earth, which is 6,371,393 meters.

[0114] (4) Photographing the display terminal of the instrument.

[0115] The technical effects of the present invention are as follows:

[0116] First, the existing method for collecting the location of distribution equipment mainly relies on manual on-site registration, combined with handheld GPS or mobile devices for location marking. However, this method is limited by GPS positioning accuracy, manual measurement errors and complex environmental interference, resulting in large errors in the collection results, which in turn affects the drawing accuracy along the layout. To address this problem, the present invention innovatively proposes a collection method based on multimodal perception and positioning, using a binocular camera combined with RTK high-precision positioning, and cooperating with a self-developed GCP geometric solution method, which effectively improves the collection accuracy and can achieve centimeter-level positioning effects, which is significantly better than traditional GPS positioning accuracy.

[0117] Secondly, existing methods rely primarily on manual experience and judgment for device identification. Faced with a wide variety of distribution equipment with diverse shapes, manual identification is not only inefficient but also susceptible to subjective errors. This is especially true when the equipment is obscured or the nameplate information is blurred, resulting in low recognition accuracy. This invention uses the MPR multimodal recognition method to fuse 2D images and 3D point cloud information, utilizing the KAF key point perception fusion and GDCF gated dual-path fusion mechanism to ensure a more efficient and accurate device identification process. This method can effectively cope with complex environments and occlusions, reducing the probability of misidentification and missed identification.

[0118] Third, to address the existing inability to effectively determine the spatial distance between the target device and the camera, this invention uses a prototype network to simultaneously extract optimized 3D keypoint information during classification and calculate the relative distance between the camera and the target device. Combined with the GCP module, this method utilizes geometric calculation methods and RTK data for coordinate conversion, ultimately accurately determining the device's actual geographic location. This overcomes the limitation of traditional RTK, which only provides the coordinates of the camera but cannot determine the device's actual location.

[0119] Furthermore, existing single-modality recognition methods based on images or point clouds suffer from insufficient generalization and overfitting in small sample environments. This invention, through innovative dataset preparation methods, employs multi-view depth rendering and a randomized shooting distance strategy to enhance the model's generalization capabilities. This ensures high recognition accuracy and stability even when sample data is scarce, significantly improving the model's adaptability and practicality.

[0120] In summary, this invention significantly improves the efficiency and accuracy of mapping medium and low voltage power distribution equipment along the layout through multimodal data collection, an innovative feature fusion mechanism, and precise key point extraction and location methods. Compared to traditional manual data collection and identification methods, this invention not only reduces labor costs and error risks, but also enhances the automation and intelligence of data collection, offering greater engineering application value and potential for widespread adoption. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] Figure 1 It is the overall flow chart of the present invention;

[0122] Figure 2 This is the overall flow chart of the MPR training method of the present invention;

[0123] Figure 3 This is a flow chart of the first stage of the KAF module process of the present invention;

[0124] Figure 4 This is a flow chart of the second stage of the KAF module process of the present invention;

[0125] Figure 5 This is a flow chart of the GDCF module of the present invention;

[0126] Figure 6 A schematic diagram of the prototype network of the present invention;

[0127] Figure 7 The electronic compass of the present invention is used in FIG.

[0128] Figure 8 This is a diagram showing the use of the gyroscope of the present invention;

[0129] Figure 9 This is a real-time interface diagram of the photographing instrument of the present invention;

[0130] Figure 10 The results of the embodiment are compared and shown in FIG. DETAILED DESCRIPTION

[0131] like Figure 1 As shown, the method for accurately locating the position of medium and low voltage distribution equipment based on multimodal perception includes the following steps:

[0132] S1. Turn on the camera and position it:

[0133] S2. Shooting target device:

[0134] S3. Data preprocessing:

[0135] S4, using MPR multimodal recognition:

[0136] S5. GCP position calculation:

[0137] S6. Save the actual location of the target device.

[0138] S4's MPR method (Multi-modal Perception for Recognition), such as Figure 2 As shown, the specific steps include:

[0139] (1) Preparation of training dataset

[0140] (2) A multi-view depth rendering method is adopted, combined with a minimum depth selection strategy and an adaptive extended projection method, to optimize the projection accuracy of the 3D point cloud on the 2D view, and the generalization ability of the model is enhanced by randomizing the view angle and shooting distance to reduce the overfitting problem.

[0141] (3) DGCNN extracts features from 3D point clouds and provides robust deep feature support for the MPR recognition process.

[0142] (4) The KAF module combines key point screening, local feature enhancement, and cross-instance matching to build an efficient point cloud feature extraction and fusion solution. Figure 3 and Figure 4 .

[0143] (5) CLIP network is used as a 2D visual feature extraction module to enhance the image perception capability of the MPR scheme for distribution equipment.

[0144] (6) The gated dual-path feature fusion (GDCF) mechanism fuses 3D keypoint features (output of the KAF module) and 2D visual features (output of the CLIP network). Figure 5 .

[0145] (7) Figure 6 ,While classifying features, the prototype network finely processes and filters key point features, ,calculates the relative distance S between the key point and the shooting instrument, and ,passes the optimized data to GEP for position solution.

[0146] S5's GCP method (Geometric Computation-based Positioning) includes the following steps:

[0147] (1) Determine the relative position of the shooting instrument

[0148] The electronic compass module can automatically obtain the offset angle ω of the current shooting instrument and determine the accurate position of the medium and low voltage distribution equipment relative to the shooting instrument, such as Figure 7 As shown in Figure 2. The relative distance S is obtained using the MPA method.

[0149] (2) Obtaining the actual horizontal distance

[0150] The gyroscope module can obtain the pitch angle θ of the shooting instrument relative to the key point of the power distribution equipment, and obtain the actual distance S' according to the trigonometric function relationship, such as Figure 8 shown.

[0151] (3) Superposition calculation

[0152] The actual distance S' and offset angle ω are superimposed on the RTK positioning information to obtain the precise coordinates (long2, lat2) of the actual distribution equipment. The specific formula is as follows. Assume that the known positioning information of the camera is (long1, lat1) in degrees.

[0153] long2=long1+S'×sinω / [ARC×cos(lat1)×2π / 360]

[0154] lat2=lat1+S'cosω / (ARC*2π / 360)

[0155] ARC is the average radius of the Earth, which is 6,371,393 meters.

[0156] (4) Shooting instrument display terminal, such as Figure 9 .

[0157] In order to verify the effectiveness of the multimodal fusion recognition scheme in this invention, Figure 10 The paper demonstrates a comparison of recognition effects under different data input conditions, comparing the recognition results of using 3D point cloud data alone, using 2D image data alone, and the multimodal fusion scheme proposed in this invention, and clearly demonstrates the differences between the schemes in recognition accuracy and anti-environmental interference capabilities.

[0158] First, in the recognition scheme that only uses 3D point cloud data, although the geometric shape of the object can be extracted, due to the limited capture of object texture and detail information by point cloud data, it is easy to make misjudgments among targets with similar shapes. Figure 10The paper demonstrates an example of utility pole recognition. The results show that when relying solely on 3D point cloud data, some trees are mistakenly identified as utility poles due to their similar shape. This misidentification stems primarily from the lack of detailed information in point cloud data against complex backgrounds, making it impossible to effectively distinguish objects with similar shapes but different categories, resulting in reduced recognition accuracy.

[0159] Secondly, in the recognition scheme that only uses 2D image data, although the image data can provide rich texture and color information, it is easy to introduce redundant interference information in complex backgrounds. Figure 10 The recognition results in Figure 2 show that when relying solely on 2D image data, the ground area at the base of the utility pole is mistakenly identified as part of the pole. This demonstrates that single 2D visual information is easily affected by lighting, shadows, and background complexity when processing environmental noise, resulting in unclear recognition boundaries and reduced recognition accuracy and stability.

[0160] In contrast, the multimodal fusion recognition solution proposed in this invention deeply fuses 3D key point features with 2D visual global features through the GDCF mechanism, achieving the complementary advantages of the two modalities. In the actual recognition process, 3D point cloud data provides accurate geometric structure information, ensuring that the model can accurately capture the morphological characteristics of the target object, while 2D image data provides rich texture and visual details, effectively making up for the shortcomings of 3D point clouds in texture expression. The recognition results after fusion show that the model can accurately identify utility poles with clear recognition boundaries, and there is no misidentification of the ground or background environment, which significantly improves the recognition accuracy and robustness.

[0161] Example 1: Multimodal location acquisition application for medium and low voltage distribution equipment

[0162] In this embodiment, the system uses binocular cameras to collect visible light images and point cloud data, and combines it with the RTK positioning system of the 5G network to achieve accurate position calibration of the power distribution equipment. The specific implementation steps are as follows:

[0163] Equipment Configuration: The system is deployed on a power inspection vehicle equipped with a binocular camera and an RTK positioning system. When the vehicle is powered on, the RTK positioning system automatically activates and calibrates the camera position.

[0164] Data Collection: In this embodiment, a Few-Shot Learning (FSL) strategy is employed. By training from a limited number of labeled samples, the system is able to efficiently distinguish and identify power distribution equipment using a small number of samples. In the training dataset, the number of power distribution equipment categories is set to N, with only K samples provided for each category (N-way K-shot). This ensures that the model can efficiently learn features and identify equipment using a small number of samples.

[0165] Data processing and positioning: In the MPR method, a few-sample classification framework is adopted, and the multimodal fusion recognition process is optimized by dividing the support set (Support Set) and the query set (Query Set). During the training process, the system adopts cross-perspective learning and multi-perspective deep rendering strategies to improve the generalization ability of the model under small sample conditions. By randomizing the shooting angle and shooting distance to enhance the diversity of training data, the model can still learn highly recognizable features on limited samples, successfully coping with the problems of complex environment and a wide variety of devices. Finally, the actual geographic location of the target device is calculated through the geometric calculation method (GCP).

[0166] Application Results: In practical applications, this method can reduce human error, improve the accuracy of position calibration, and ensure the accuracy of device location maps. Compared with traditional manual registration methods, this system can maintain high accuracy even in complex environments and accurately measure the relative position and geographic coordinates of devices.

[0167] Example 2: Device positioning in complex urban environments

[0168] This embodiment relates to an application for accurate location acquisition of power distribution equipment in an urban environment where dense buildings, trees, and other obstructions affect the positioning of equipment.

[0169] Environmental Configuration: The system is deployed on an inspection drone, combining binocular cameras and LIDAR sensors for data collection. The drone obtains precise positioning information via a real-time 5G network.

[0170] Data Collection and Fusion: In complex urban environments, a drone flies near a target device, capturing clear 2D images with a binocular camera while simultaneously acquiring 3D point cloud data around the device using a LIDAR sensor. The system employs an MPR algorithm for data fusion, leveraging multimodal data from point clouds and images to identify and locate the target device.

[0171] Positioning and Analysis: The relative position of the target device is calculated using a GCP algorithm and combined with real-time RTK data to calculate geographic coordinates. Due to the high precision of the LIDAR sensor and its multimodal fusion method, the system can effectively overcome obstructions such as buildings and trees and accurately identify obscured distribution equipment.

[0172] Application effect: This embodiment demonstrates the precise positioning of power distribution equipment in complex environments through multimodal technology. Especially when the equipment is obscured or located in densely populated urban areas, the system can provide centimeter-level positioning information, greatly improving inspection efficiency and positioning accuracy.

[0173] Example 3: Intelligent identification and positioning of power distribution equipment based on offshore fixed sensor arrays

[0174] Offshore Platform Configuration: This embodiment utilizes fixed offshore platforms or buoy arrays for intelligent identification and location of power distribution equipment. Multiple sensors, including binocular cameras, LiDAR (Light Detection and Ranging) sensors, and RTK base stations, are installed on offshore platforms. These sensors are fixedly mounted on offshore buoys, offshore wind turbine platforms, offshore substations, or other offshore structures, enabling real-time monitoring of power distribution equipment across a wide area.

[0175] Small-sample dataset construction: Using a small-sample learning approach, we construct a training dataset using an N-way K-shot approach, providing only a small number of labeled samples (K) for each device category. This small-sample learning framework enables efficient training in offshore environments using only a limited number of labeled samples, ensuring accurate identification of different offshore power distribution equipment.

[0176] Multimodal Data Fusion and Small Sample Optimization: A fixed offshore sensor array simultaneously collects 2D images and 3D point cloud data of target devices. Using the MPR multimodal perception method, combined with the DGCNN and CLIP networks, efficient learning is achieved in a small sample environment. The KAF module perceives and fuses key point information of the device, ensuring accurate recognition with limited sample data.

[0177] Positioning and Geometric Solution: The RTK base station system based on the offshore platform provides precise positioning data. Combining the fixed position of the offshore platform with the relative distance information acquired by the sensors, the device's positioning is solved using GCP geometric calculation methods, ultimately calculating the device's precise geographic coordinates. Multimodal learning using point cloud and image data ensures centimeter-level positioning accuracy.

[0178] Application Results: In offshore applications, the system, through a fixed sensor array, enables real-time monitoring and location acquisition of offshore power distribution equipment around the clock. Thanks to the introduction of small-sample learning, the system is able to identify and locate equipment even with minimal data annotation. Particularly in environments such as offshore wind farms and offshore substations, the system can overcome interference from harsh sea conditions on equipment identification and positioning, providing accurate equipment monitoring results.

Claims

1. A method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception, characterized in that: The following steps are involved: S1. Power on and locate the camera: After the camera is powered on, it performs 5G positioning and RTK secondary positioning calibration to obtain the current precise positioning of the instrument. The camera is a binocular camera with a built-in edge computing development board, which can analyze and process the collected point cloud data and image data in real time. The MPR algorithm and GCP algorithm are both deployed in the camera, realizing the full process of automated detection and positioning of medium and low voltage distribution equipment at the edge. S2. Photograph the target device: Use a camera to photograph the target device and obtain visible light image data and point cloud data of the device. Each visible light image and each set of point cloud data will be accompanied by a timestamp and the precise positioning of the camera at the moment of capture. The visible light image data and point cloud data captured at the same time will be aligned to ensure that they correspond to the same power distribution equipment. S3. Data preprocessing: Preprocess the visible light image data. Through the calibration and spatial registration method of the binocular camera, the 3D points in the point cloud are mapped to the image pixels to ensure the spatial consistency of the point cloud and image data. S4. Use MPR multimodal recognition: Use the MPR method to identify and detect multi-source data, accurately identify the target device, and finally use the key three-dimensional coordinate information in the point cloud data to calculate the relative distance S between the shooting instrument and the distribution equipment; S5, GCP Position Solution: Based on the geometric calculation positioning GCP method and combined with the RTK positioning data of the shooting instrument, the relative distance S is converted to coordinates to calculate the actual horizontal distance S' of the target device. Finally, the actual geographic coordinates of the target device are solved through geometric solution methods to ensure centimeter-level accuracy. S6. Save the actual location of the target device; The calculated actual location information of the target device is stored in the shooting instrument in the standard format of category-serial number-location information.

2. The method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception according to claim 1 is characterized in that: The multimodal perception and recognition (MPR) method in S4 specifically includes the following steps: (1) Preparation of training dataset; The dataset is divided into training, validation, and test sets, following the basic framework of FSL (Few-Shot Learning) and using an N-way K-shot approach. The dataset is optimized using a support set and a query set. (2) A multi-view depth rendering method is adopted, combined with a minimum depth selection strategy and an adaptive extended projection method to optimize the projection accuracy of 3D point clouds on 2D views, and the generalization ability of the model is enhanced by randomizing the view angle and shooting distance to reduce the overfitting problem; (3) DGCNN network extracts features from 3D point clouds; Use DGCNN to generate high-dimensional point cloud features, and further filter the key points of the target device in the key point perception fusion (KAF) module; (4) The key point perception fusion KAF module combines key point screening, local feature enhancement, and cross-instance matching to build a point cloud feature extraction and fusion solution; (5) CLIP network is used as a 2D visual feature extraction module to enhance the image perception capability of the MPR scheme for distribution equipment; (6) A gated dual-path feature fusion mechanism (GDCF) is used to fuse 3D key point features and 2D visual features. (7) While classifying features, the prototype network finely processes and filters key point features, calculates the relative distance S between the key point and the camera, and passes the optimized data to GEP for position solution.

3. The method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception according to claim 2 is characterized in that: The training set in (1) is used for model learning features, covering multiple categories of power distribution equipment, and each category contains only K samples. If the number of categories is N, then the total number of samples in the training set is N × K. In each training round, the training data is further divided into a support set and a query set, where the support set is used for feature learning and the query set is used for model optimization. The validation set is used to monitor the generalization performance of the model during training. It uses data from categories that were not used in training for evaluation to ensure that the model does not overfit to samples of a specific category. The test set is used for final performance evaluation and has no overlap with the training categories to ensure that the model can generalize to new categories of power distribution equipment. The data were divided using a stratified sampling method.

4. The method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception according to claim 2 is characterized in that: (2) The specific method is: adopting a continuous three-view depth rendering strategy, namely XY, YZ, and XZ views, and improving the projection quality through the optimal depth selection mechanism; specifically, for a 3D point (x, y, z), the process of projecting it to a 2D depth map is as follows: For each perspective, the minimum depth selection strategy is adopted, that is, the depth value closest to the camera is selected from multiple projection points: We further adopt an adaptive extended projection method and use a neighborhood expansion strategy to fill the blank areas, making the depth map more coherent. For pixels (X, Y) with missing depth values, we define a neighborhood expansion radius R, and the matching area M (X, Y, R) is determined by the following conditions: M(X,Y,R)={(x,y,z)∈P∣XR / 2≤X' <X+R / 2,Y-R / 2≤Y'<Y+R / 2} Finally, the depth value of the blank pixel is completed by the minimum depth in the neighborhood: D(X,Y)=min{z∣(x,y,z)∈M(X,Y,R)} When R=1, only the depth value of the pixel itself is used without expansion; when R>1, neighboring points are used for completion, thereby enhancing the integrity of the depth map and avoiding the problem of losing object boundary information in traditional interpolation methods; In addition, in order to prevent the model from overfitting to a fixed viewing angle, a randomized shooting distance mechanism is introduced. During each depth rendering, the position of the virtual camera is randomly perturbed to simulate target projections at different scales and improve the generalization ability of the model. Assume that the initial camera position C0 = (x c ,y c ,z c ), the actual camera position at each rendering is: C=(x c +Δx,y c +Δy,z c +Δz) Among them, Δx, Δy, and Δz obey the uniform distribution U(-r, r), and r is the preset random perturbation range.

5. The method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception according to claim 2 is characterized in that: (4) The specific method is: KAF uses a key point screening mechanism to select the most discriminative key points from the 3D point cloud features extracted by DGCNN in the support set and query set; Specifically, let the input point cloud be P = {p1,p2,...,p N }, where each point p i With d-dimensional feature representation F(p i ), then calculate the significance score of each point: Among them, F g It is a global feature, calculated by Global Max-Pooling: Finally, take the top k with the highest significance scores s points as the key point set P k ; KAF adopts the neighborhood feature aggregation strategy to enhance the local features of key points and take its k nearest neighbor information for local feature enhancement; for each key point p j , construct its neighborhood feature N j As a set of k nearest neighbor points: N j ={p j1 ,p j2 ,...,p jk },p ji ∈k-NN(p j ) Then, the global feature F g Connect to neighborhood features to enhance the keypoint’s ability to represent the target category: F e,j =g([F g ,N j ]) Here, g(·) represents a 1×1 convolutional mapping operation, which ensures the consistency of information at different scales, so that local features not only retain their own information but also combine global features to improve the discrimination ability. The second stage of KAF optimizes the feature matching of key points through the cross-instance fusion mechanism; given the global prototype feature matrix of the support set (containing d-dimensional prototype features of N categories) and the query set global feature matrix (Including N q d-dimensional features of query instances), first calculate the cross-instance relationship graph through the scaled dot product attention mechanism: in is the scaling factor; each element of the relationship matrix Indicates the semantic association strength between the i-th class prototype and the j-th query instance, with a value range of [0,1]; In the dual-branch feature fusion design, the channel-level fusion branch is used for each prototype feature. Filter out the 5 most relevant query features by cosine similarity Construct the interaction tensor: Channel interaction is then achieved through two levels of convolution; the first layer of 1×1 convolution maps the channel dimension from d to the hidden layer The second layer of 1×1 convolution restores to the original dimension d: Finally, feature update is achieved through attention weighting: Where ⊙ represents element-wise multiplication, which explicitly enhances the local geometric features related to the target category; the instance-level fusion branch realizes feature propagation through the global relationship matrix: This branch will query the features in M pq Perform a linear combination of weights so that each prototype absorbs contextual information across instances; for example, for the i-th class prototype, its corresponding instance-level feature is: In the feature aggregation stage, the dual-branch outputs are fused through learnable weight parameters α and β:

6. The method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception according to claim 2 is characterized in that: (5) The specific method is as follows: In the MPR+GCP framework, CLIP is used to extract global features of 2D images and provide additional semantic context support through text information; specifically, CLIP uses a pre-trained image encoder to visually encode the device image, and by extracting these features, it generates a comprehensive prototype containing the prototype features of the 2D image.

7. The method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception according to claim 2 is characterized in that: (6) The specific method is: In the GDCF mechanism, the 3D key point feature F is first P and 2D visual features F I Perform global feature aggregation, where F P and F I They all contain feature matrices of their respective support sets and query sets; they use global perspective aggregation to integrate multi-view 2D visual features into a global visual representation through a learnable linear transformation: Among them, f1 and f2 are two-layer linear mappings, N is the number of perspectives, and concat(·) represents the splicing operation on the channel dimension; this can obtain a stable 2D visual global feature G I , ensuring that information is consistent across different perspectives; For 3D key point features, a learnable transformation layer is introduced after the KAF output to align the global feature space of the 2D visual features: G P =W P F P +b P Among them, W P and b P It is a learnable parameter used to adjust the representation ability of 3D key point features to meet the needs of cross-modal fusion; In the fusion stage, a gated fusion mechanism is used to dynamically adjust the contribution ratio of 3D key point features and 2D visual features through learnable gate weights σ: G=σ·G I +(1-σ)·G P Among them, σ is calculated by the learnable Sigmoid transformation: σ=Sigmoid(W G G P +b G ) Finally, the fused global feature G is passed to the prototype network for target classification, and the optimized 3D key point information is input into the GCP module.

8. The method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception according to claim 2 is characterized in that: (7) The specific method is: In the fusion stage, the GDCF mechanism has deeply integrated the 3D key point features extracted by the KAF module with the 2D visual features generated by the CLIP network to obtain the fused high-dimensional features F final The prototype network first receives these fused features and calculates the prototype vector of each category based on the support set samples. The prototype of each category not only represents the overall feature center of the category, but also retains the geometric information of the fused 3D key points. After passing through the key point solution module, the key points representing the sample are found. The prototype reverse decoding module P2K is introduced to directly reconstruct the three-dimensional key point coordinates of the target device from the category prototype vector generated by the prototype network. This module maps the fused prototype features into the spatial coordinates of three geometric key points by constructing a learnable nonlinear decoding network. Let the prototype vector be The number of key points is 3, each key point is The reverse decoding process is uniformly expressed by the following formula: [k1,k2,k3]=Reshape 3×3 (φ(p c )) φ is a mapping function composed of a multi-layer perceptron MLP, Reshape 3×3 Indicates reshaping the 9-dimensional vector into 3 3-dimensional coordinate points; After obtaining the coordinates of the three key points, the system further calculates the Euclidean distance of each key point with the three-dimensional space coordinates of the current shooting instrument: d i =‖k i -c‖,i=1,2,3 Finally, the three distances are averaged as the estimated spatial distance of the device relative to the shooting instrument: The optimized key point data and relative distance S are encapsulated and passed to the GEP module; after receiving this data, the GEP further calculates the actual geographic coordinates of the target device in combination with the RTK high-precision positioning data.

9. The method for accurately locating the position of medium and low voltage power distribution equipment based on multimodal perception according to claim 2 is characterized in that: S5's method for locating GCPs based on geometric solution includes the following steps: (1) Determine the relative position of the shooting instrument; The electronic compass module can automatically obtain the offset angle ω of the current shooting instrument and determine the accurate position of the medium and low voltage distribution equipment relative to the shooting instrument. The relative distance S is obtained using the MPA method. (2) Obtain the actual horizontal distance; The gyroscope module can obtain the pitch angle θ of the shooting instrument relative to the key point of the power distribution equipment, and obtain the actual distance S' based on the trigonometric function relationship; (3) superposition calculation; The actual distance S' and the offset angle ω are superimposed on the RTK positioning information to obtain the precise coordinates (long2, lat2) of the actual distribution equipment. Assume that the known positioning information of the camera is (long1, lat1), in degrees: long2=long1+S'×sinω / [ARC×cos(lat1)×2π / 360] lat2=lat1+S'cosω / (ARC*2π / 360) ARC is the average radius of the Earth, which is 6,371,393 meters; (4) Photographing the display terminal of the instrument.

Citation Information

Cited By

  • Three-dimensional positioning method and system for bimodal target

    CN122043356A

  • A three-dimensional localization method and system for dual-modal targets

    CN122043356B